# Throttle merges

**URL:** <https://discuss.elastic.co/t/throttle-merges/31780>\
**Category:** Elasticsearch\
**Created:** [October 7, 2015, 2:32pm UTC](https://discuss.elastic.co/t/throttle-merges/31780 "2015-10-07T14:32:29Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sergey\_Novikov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sergey_novikov/32/831_2.png) [@Sergey\_Novikov](https://discuss.elastic.co/u/Sergey_Novikov)\
**Post date:** [October 7, 2015, 2:32pm UTC](https://discuss.elastic.co/t/throttle-merges/31780/1 "2015-10-07T14:32:29Z")

</div>

I'm having issues with heavy CPU usage and I/O after indexing. Long after indexing has finished, load average is still high and CPU is loaded (50-100%), this makes search operations slow.

I've set `index.merge.scheduler.max_thread_count` to 1, but still see many threads per node (`GET _nodes/hot_threads`):

```auto
 100.2% (500.7ms out of 500ms) cpu usage by thread 'elasticsearch[testing-195146d5][[testindex][9]: Lucene Merge Thread #49]'
   3/10 snapshots sharing following 14 elements
     org.apache.lucene.codecs.DocValuesConsumer$3$1.hasNext(DocValuesConsumer.java:316)
     org.apache.lucene.codecs.DocValuesConsumer$10$1.hasNext(DocValuesConsumer.java:855)
     ...
     org.apache.lucene.index.ConcurrentMergeScheduler$MergeThread.run(ConcurrentMergeScheduler.java:486)

64.7% (323.3ms out of 500ms) cpu usage by thread 'elasticsearch[testing-195146d5][[testindex][5]: Lucene Merge Thread #35]'
 10/10 snapshots sharing following 9 elements
   org.apache.lucene.codecs.perfield.PerFieldDocValuesFormat$FieldsWriter.addSortedNumericField(PerFieldDocValuesFormat.java:122)
   org.apache.lucene.codecs.DocValuesConsumer.mergeSortedNumericField(DocValuesConsumer.java:301)
   org.apache.lucene.index.SegmentMerger.mergeDocValues(SegmentMerger.java:223)
   org.apache.lucene.index.SegmentMerger.merge(SegmentMerger.java:122)
   org.apache.lucene.index.IndexWriter.mergeMiddle(IndexWriter.java:4223)
   org.apache.lucene.index.IndexWriter.merge(IndexWriter.java:3811)
   org.apache.lucene.index.ConcurrentMergeScheduler.doMerge(ConcurrentMergeScheduler.java:409)
   org.apache.lucene.index.TrackingConcurrentMergeScheduler.doMerge(TrackingConcurrentMergeScheduler.java:107)
   org.apache.lucene.index.ConcurrentMergeScheduler$MergeThread.run(ConcurrentMergeScheduler.java:486)

52.6% (262.9ms out of 500ms) cpu usage by thread 'elasticsearch[testing-195146d5][[testindex][1]: Lucene Merge Thread #46]'
 2/10 snapshots sharing following 17 elements
   org.apache.lucene.index.SingletonSortedNumericDocValues.setDocument(SingletonSortedNumericDocValues.java:52)
   org.apache.lucene.codecs.DocValuesConsumer$3$1.setNext(DocValuesConsumer.java:353)
   org.apache.lucene.codecs.DocValuesConsumer$3$1.hasNext(DocValuesConsumer.java:316)
...
   org.apache.lucene.index.ConcurrentMergeScheduler$MergeThread.run(ConcurrentMergeScheduler.java:486)

99.1% (495.6ms out of 500ms) cpu usage by thread 'elasticsearch[testing-195146d5][[testindex][0]: Lucene Merge Thread #47]'
 2/10 snapshots sharing following 13 elements
   org.apache.lucene.codecs.DocValuesConsumer$3$1.setNext(DocValuesConsumer.java:359)
   org.apache.lucene.codecs.DocValuesConsumer$3$1.hasNext(DocValuesConsumer.java:316)
       

```

Index section in my configuration looks like this:

```auto
index:
  auto_expand_replicas: 0-all
  merge:
    scheduler:
      max_thread_count: 1
...
indices:
  store:
    throttle:
      max_bytes_per_sec: 10mb
      type: merge

```

there are messages in log:

```auto
now throttling indexing: numMergesInFlight=4, maxNumMerges=3

```

By the way, merge problems started after we moved some of the data to `doc_values`.

Do I misunderstand the setting? Looks like it ensures only 1 merge thread will be running.  
What's the difference with `index.merge.policy.max_merge_at_once`, should I set it instead?

What I want is to make merge less aggressive.

---

<div class="post-metadata">

**Author:** ![Srinath\_C](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/srinath_c/32/48806_2.png) [@Srinath\_C](https://discuss.elastic.co/u/Srinath_C)\
**Post date:** [October 7, 2015, 3:50pm UTC](https://discuss.elastic.co/t/throttle-merges/31780/2 "2015-10-07T15:50:21Z")

</div>

See the discussion in [Index throttling issue](https://discuss.elastic.co/t/index-throttling-issue/31440/6).  
We faced a similar issue after using doc\_values and we resolved it by removing merge policy and index throttle settings.  
We are still facing increased disk utilization though.

---

<div class="post-metadata">

**Author:** ![Sergey\_Novikov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sergey_novikov/32/831_2.png) [@Sergey\_Novikov](https://discuss.elastic.co/u/Sergey_Novikov)\
**Post date:** [October 7, 2015, 3:56pm UTC](https://discuss.elastic.co/t/throttle-merges/31780/3 "2015-10-07T15:56:34Z")

</div>

hm, but it seems exactly the opposite of what I'm trying to achieve? I.e. I need to throttle merges, otherwise they take all the resources and make any other operations slow.

(with `doc_values` disk utilisation is increased for us as well, but no heap problems anymore)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 7, 2015, 9:10pm UTC](https://discuss.elastic.co/t/throttle-merges/31780/4 "2015-10-07T21:10:06Z")

</div>

There's always a cost, so you just need to balance those accordingly.

---

<div class="post-metadata">

**Author:** ![Sergey\_Novikov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sergey_novikov/32/831_2.png) [@Sergey\_Novikov](https://discuss.elastic.co/u/Sergey_Novikov)\
**Post date:** [October 26, 2015, 11:07am UTC](https://discuss.elastic.co/t/throttle-merges/31780/5 "2015-10-26T11:07:54Z")

</div>

After some experiments, it turns out `index.merge.scheduler.max_thread_count = 1` doesn't mean 1 thread (more like 4 in my case), but still reduces load. I ran 50000 partial updates in 50 threads

without throttling:

 ![](https://us1.discourse-cdn.com/elastic/original/2X/8/8115d5d45eabc2427b8e82a23e3d49a525ff05c2.png)  
 ![](https://us1.discourse-cdn.com/elastic/original/2X/f/f6fa132c8986b182a7024d426ccfa83694e4c134.png)

and with `max_thread_count = 1`

 ![](https://us1.discourse-cdn.com/elastic/original/2X/0/0226f3de8c253749f539305ef0def8375907435a.png)

![](https://us1.discourse-cdn.com/elastic/original/2X/1/1435f10b4188e3294ef60f0606227242f832c797.png)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:42pm UTC](https://discuss.elastic.co/t/throttle-merges/31780/6 "2017-07-05T23:42:42Z")

</div>


