# Uncontrolled merge process on few indices

**URL:** <https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103>\
**Category:** Elasticsearch\
**Created:** [December 10, 2018, 6:41am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103 "2018-12-10T06:41:26Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![shshnk28](https://avatars.discourse-cdn.com/v4/letter/s/d07c76/32.png) [@shshnk28](https://discuss.elastic.co/u/shshnk28)\
**Post date:** [December 10, 2018, 6:41am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103/1 "2018-12-10T06:41:26Z")

</div>

Hi ,  
We have a 7 node cluster running on ES version 5.2.1. We have observed that ES internal merge process sometimes goes berserk for couple of indices, runs for more than 3 days and consumes the server resources (like CPU) till they max out. I am interested in knowing that what might be causing such uncontrolled merge. Both of these indices have certain other types as their child.  
Here is the output of hot threads api.

[https://pastebin.com/A0ZcF3P3](https://pastebin.com/A0ZcF3P3)

Here are the index stats

[https://pastebin.com/KB2hT0Yv](https://pastebin.com/KB2hT0Yv)

Are there some pointers on where I should be looking ? Thanks

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 10, 2018, 7:05am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103/2 "2018-12-10T07:05:35Z")

</div>

What is the specification of the cluster? What kind of storage are you using? How frequently do you refresh the index?

---

<div class="post-metadata">

**Author:** ![shshnk28](https://avatars.discourse-cdn.com/v4/letter/s/d07c76/32.png) [@shshnk28](https://discuss.elastic.co/u/shshnk28)\
**Post date:** [December 10, 2018, 7:24am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103/3 "2018-12-10T07:24:16Z")

</div>

Hi, Thanks for your response. Here is the output of cluster stats command.  
{  
"\_nodes" : {  
"total" : 8,  
"successful" : 8,  
"failed" : 0  
},  
"cluster\_name" : "SearchCluster",  
"timestamp" : 1544426303031,  
"status" : "green",  
"indices" : {  
"count" : 30,  
"shards" : {  
"total" : 68,  
"primaries" : 68,  
"replication" : 0.0,  
"index" : {  
"shards" : {  
"min" : 1,  
"max" : 12,  
"avg" : 2.2666666666666666  
},  
"primaries" : {  
"min" : 1,  
"max" : 12,  
"avg" : 2.2666666666666666  
},  
"replication" : {  
"min" : 0.0,  
"max" : 0.0,  
"avg" : 0.0  
}  
}  
},  
"docs" : {  
"count" : 8034382162,  
"deleted" : 446758535  
},  
"store" : {  
"size" : "1.4tb",  
"size\_in\_bytes" : 1548871164733,  
"throttle\_time" : "0s",  
"throttle\_time\_in\_millis" : 0  
},  
"fielddata" : {  
"memory\_size" : "1.1gb",  
"memory\_size\_in\_bytes" : 1231750200,  
"evictions" : 0  
},  
"query\_cache" : {  
"memory\_size" : "10.7gb",  
"memory\_size\_in\_bytes" : 11498162014,  
"total\_count" : 8145076647,  
"hit\_count" : 1328346575,  
"miss\_count" : 6816730072,  
"cache\_size" : 170525,  
"cache\_count" : 122852859,  
"evictions" : 122682334  
},  
"completion" : {  
"size" : "0b",  
"size\_in\_bytes" : 0  
},  
"segments" : {  
"count" : 1691,  
"memory" : "3.1gb",  
"memory\_in\_bytes" : 3360026726,  
"terms\_memory" : "2.7gb",  
"terms\_memory\_in\_bytes" : 2899292132,  
"stored\_fields\_memory" : "128.1mb",  
"stored\_fields\_memory\_in\_bytes" : 134352368,  
"term\_vectors\_memory" : "0b",  
"term\_vectors\_memory\_in\_bytes" : 0,  
"norms\_memory" : "5mb",  
"norms\_memory\_in\_bytes" : 5327360,  
"points\_memory" : "54.9mb",  
"points\_memory\_in\_bytes" : 57603434,  
"doc\_values\_memory" : "251.2mb",  
"doc\_values\_memory\_in\_bytes" : 263451432,  
"index\_writer\_memory" : "294.1mb",  
"index\_writer\_memory\_in\_bytes" : 308432515,  
"version\_map\_memory" : "1.7mb",  
"version\_map\_memory\_in\_bytes" : 1839883,  
"fixed\_bit\_set" : "1.5gb",  
"fixed\_bit\_set\_memory\_in\_bytes" : 1650422272,  
"max\_unsafe\_auto\_id\_timestamp" : -1,  
"file\_sizes" : { }  
}  
},  
"nodes" : {  
"count" : {  
"total" : 8,  
"data" : 7,  
"coordinating\_only" : 1,  
"master" : 7,  
"ingest" : 7  
},  
"versions" : [  
"5.2.1"  
],  
"os" : {  
"available\_processors" : 88,  
"allocated\_processors" : 88,  
"names" : [  
{  
"name" : "Linux",  
"count" : 8  
}  
],  
"mem" : {  
"total" : "239.9gb",  
"total\_in\_bytes" : 257681920000,  
"free" : "4.1gb",  
"free\_in\_bytes" : 4407902208,  
"used" : "235.8gb",  
"used\_in\_bytes" : 253274017792,  
"free\_percent" : 2,  
"used\_percent" : 98  
}  
},  
"process" : {  
"cpu" : {  
"percent" : 121  
},  
"open\_file\_descriptors" : {  
"min" : 441,  
"max" : 1125,  
"avg" : 629  
}  
},  
"jvm" : {  
"max\_uptime" : "96.7d",  
"max\_uptime\_in\_millis" : 8360773114,  
"versions" : [  
{  
"version" : "1.8.0\_161",  
"vm\_name" : "OpenJDK 64-Bit Server VM",  
"vm\_version" : "25.161-b14",  
"vm\_vendor" : "Oracle Corporation",  
"count" : 1  
}  
],  
"mem" : {  
"heap\_used" : "64.7gb",  
"heap\_used\_in\_bytes" : 69473495008,  
"heap\_max" : "145.3gb",  
"heap\_max\_in\_bytes" : 156103606272  
},  
"threads" : 1073  
},  
"fs" : {  
"total" : "4.1tb",  
"total\_in\_bytes" : 4596416278528,  
"free" : "2.7tb",  
"free\_in\_bytes" : 3020499034112,  
"available" : "2.5tb",  
"available\_in\_bytes" : 2797994000384  
},  
"plugins" : ,  
"network\_types" : {  
"transport\_types" : {  
"netty4" : 8  
},  
"http\_types" : {  
"netty4" : 8  
}  
}  
}  
}

1. We are using EBS on AWS(SSDs) for the disks. Checked the disk stats using iowait. They do not show any kind of saturation.

2. As for the refresh interval we have not changed the default ones that comes with ES. On one of the troublesome indices we receive ~50 bulk indexing request per minute which would translate to ~100 documents per minute.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 10, 2018, 7:27am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103/4 "2018-12-10T07:27:02Z")

</div>

What kind of storage are you using? How frequently do you refresh the index?

---

<div class="post-metadata">

**Author:** ![shshnk28](https://avatars.discourse-cdn.com/v4/letter/s/d07c76/32.png) [@shshnk28](https://discuss.elastic.co/u/shshnk28)\
**Post date:** [December 10, 2018, 7:30am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103/5 "2018-12-10T07:30:48Z")

</div>

Hi, I have updated my answer above.

Additionally, Continuously hitting hot threads endpoint reveals these stack traces as common. Could this be the culprit .

org.apache.lucene.util.LongValues.get(LongValues.java:45)  
org.apache.lucene.index.SingletonSortedNumericDocValues.setDocument(SingletonSortedNumericDocValues.java:52)  
org.apache.lucene.codecs.DocValuesConsumer$SortedNumericDocValuesSub.nextDoc(DocValuesConsumer.java:449)  
org.apache.lucene.index.DocIDMerger$SequentialDocIDMerger.next(DocIDMerger.java:100)  
org.apache.lucene.codecs.DocValuesConsumer$3$1.setNext(DocValuesConsumer.java:511)  
org.apache.lucene.codecs.DocValuesConsumer$3$1.hasNext(DocValuesConsumer.java:491)  
org.apache.lucene.codecs.DocValuesConsumer.isSingleValued(DocValuesConsumer.java:998)  
org.apache.lucene.codecs.lucene54.Lucene54DocValuesConsumer.addSortedNumericField(Lucene54DocValuesConsumer.java:586)  
org.apache.lucene.codecs.DocValuesConsumer.mergeSortedNumericField(DocValuesConsumer.java:470)  
org.apache.lucene.codecs.DocValuesConsumer.merge(DocValuesConsumer.java:243)

\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*8

org.apache.lucene.codecs.DocValuesConsumer$10$1.next(DocValuesConsumer.java:1024)  
org.apache.lucene.codecs.DocValuesConsumer$10$1.next(DocValuesConsumer.java:1015)  
java.util.Spliterators$IteratorSpliterator.tryAdvance(Spliterators.java:1812)  
java.util.stream.StreamSpliterators$WrappingSpliterator.lambda$initPartialTraversalState$0(StreamSpliterators.java:294)  
java.util.stream.StreamSpliterators$WrappingSpliterator$$Lambda$1481/1465281058.getAsBoolean(Unknown Source)  
java.util.stream.StreamSpliterators$AbstractWrappingSpliterator.fillBuffer(StreamSpliterators.java:206)  
java.util.stream.StreamSpliterators$AbstractWrappingSpliterator.doAdvance(StreamSpliterators.java:169)  
java.util.stream.StreamSpliterators$WrappingSpliterator.tryAdvance(StreamSpliterators.java:300)  
java.util.Spliterators$1Adapter.hasNext(Spliterators.java:681)  
org.apache.lucene.codecs.lucene54.Lucene54DocValuesConsumer.addNumericField(Lucene54DocValuesConsumer.java:105)  
org.apache.lucene.codecs.lucene54.Lucene54DocValuesConsumer.addNumericField(Lucene54DocValuesConsumer.java:296)  
org.apache.lucene.codecs.lucene54.Lucene54DocValuesConsumer.addNumericField(Lucene54DocValuesConsumer.java:89)  
org.apache.lucene.codecs.lucene54.Lucene54DocValuesConsumer.addSortedNumericField(Lucene54DocValuesConsumer.java:589)

---

<div class="post-metadata">

**Author:** ![shshnk28](https://avatars.discourse-cdn.com/v4/letter/s/d07c76/32.png) [@shshnk28](https://discuss.elastic.co/u/shshnk28)\
**Post date:** [December 11, 2018, 10:30am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103/6 "2018-12-11T10:30:56Z")

</div>

@Christian_Dahlqvist request you to have a look . is there any additional info I can provide,

---

<div class="post-metadata">

**Author:** ![shshnk28](https://avatars.discourse-cdn.com/v4/letter/s/d07c76/32.png) [@shshnk28](https://discuss.elastic.co/u/shshnk28)\
**Post date:** [December 17, 2018, 4:20am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103/7 "2018-12-17T04:20:18Z")

</div>

@dadoonet as an active member if you take a look on this it would be helpful for me.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 17, 2018, 6:11am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103/8 "2018-12-17T06:11:17Z")

</div>

Are you performing a lot of updates? Do you have any non-default node or index settings?

---

<div class="post-metadata">

**Author:** ![shshnk28](https://avatars.discourse-cdn.com/v4/letter/s/d07c76/32.png) [@shshnk28](https://discuss.elastic.co/u/shshnk28)\
**Post date:** [December 17, 2018, 9:25am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103/9 "2018-12-17T09:25:51Z")

</div>

We are not performing a lot of updates. As I mentioned in above thread we are indexing ~100 documents per minute.  
There is nothing out of ordinary in our index settings . I am putting it here  
{  
"index.name.primary" : {  
"settings" : {  
"index" : {  
"routing" : {  
"allocation" : {  
"include" : {  
"type" : "primary"  
}  
}  
},  
"mapping" : {  
"ignore\_malformed" : "true"  
},  
"refresh\_interval" : "1s",  
"number\_of\_shards" : "1",  
"provided\_name" : "index.name.primary",  
"max\_result\_window" : "1000000",  
"mapper" : {  
"dynamic" : "false"  
},  
"creation\_date" : "1505657680807",  
"number\_of\_replicas" : "0",  
"uuid" : "LMCLskY4S4iPooQKY2pIzQ",  
"version" : {  
"created" : "5020199"  
}  
}  
}  
}  
}

Node setting are also default.  
If you check the above thread then threre are two methods in `hot_threads` api which catch my attention those being `mergeSortedNumericField` and `addSortedNumericField`. These are present in mostly all the stack traces. Can these lead us to the issue?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 17, 2018, 9:37am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103/10 "2018-12-17T09:37:47Z")

</div>

What type of storage do you have?

---

<div class="post-metadata">

**Author:** ![shshnk28](https://avatars.discourse-cdn.com/v4/letter/s/d07c76/32.png) [@shshnk28](https://discuss.elastic.co/u/shshnk28)\
**Post date:** [December 17, 2018, 11:19am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103/11 "2018-12-17T11:19:09Z")

</div>

We are using general purpose (gp2) volumes on AWS EBS

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 14, 2019, 11:19am UTC](https://discuss.elastic.co/t/uncontrolled-merge-process-on-few-indices/160103/12 "2019-01-14T11:19:10Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
