# Too many Lucene merge threads while indexing

**URL:** <https://discuss.elastic.co/t/too-many-lucene-merge-threads-while-indexing/186822>\
**Category:** Elasticsearch\
**Created:** [June 21, 2019, 7:44am UTC](https://discuss.elastic.co/t/too-many-lucene-merge-threads-while-indexing/186822 "2019-06-21T07:44:41Z")\
**Posts on this page:** 1\
**Showing post:** 6

<div class="post-metadata">

**Author:** ![martinr\_ubi](https://avatars.discourse-cdn.com/v4/letter/m/b5e925/32.png) [@martinr\_ubi](https://discuss.elastic.co/u/martinr_ubi)\
**Post date:** [June 21, 2019, 10:26am UTC](https://discuss.elastic.co/t/too-many-lucene-merge-threads-while-indexing/186822/6 "2019-06-21T10:26:16Z")

</div>

--EDIT Was writing while you guys continued, sorry for stuff now out of context--

Node specs is hardware specs, what's the hardware profile of your nodes. CPU/core count, RAM size, JVM HEAP setting, etc.  
Are the disks SSD?  
How many fields do you have in each of your documents?  
Do the the fields names vary per document? (Different fields in different documents or all document have the same schema)  
How many total field in the indices, visible in your index mapping or by looking at the index pattern in Kibana.

I've had problems in the past with merge threads when I was suffering from field count explosion, I'm not guru enough to know if there is a true technical relationship there but I remember that's what I was seeing. Way too many fields in the same index lead to apparent lucene merge threads problems and what ES calls index throttling.

> [@What is the meaning of throttle\_time\_in\_millis?](https://discuss.elastic.co/t/what-is-the-meaning-of-throttle-time-in-millis/16128):
>
> I created an index in a 4 nodes Elasticsearch cluster. I added about 3.5 M documents using the java Elasticsearch API. When asking for the stats i get a very high number in throttle\_time\_in\_millis as follows: { "\_shards": { "total": 10, "successful": 10, "failed": 0 }, "\_all": { "primaries": { "docs": { "count": 3855540, "deleted": 0 }, "store": { "size\_in\_bytes": 1203074796, "throttle\_time\_in\_millis": 980255 }, "indexing": { "index\_total": 3855540, "index\_time\_in\_millis": …

 ![indexthrottling](https://us1.discourse-cdn.com/elastic/original/3X/5/4/54805a9399226bed17482aa02261e96544797f6a.png)  
I don't remember if the threads were idle though, and if I remember correctly there was CPU contention in my case.

2000/sec ok but are you receiving back pressure while you're indexing? What's to say that your bottleneck is not in the client(s) and that they are simply not even trying to push faster?  
If you are not finding any contention on your ES nodes, CPU RAM DISK IO, what makes you think that adding client(s) also trying to send bulk requests would not result in more doc/sec?

You could also try to provide more technical details like dumps of the info you're referencing. I don't actually know what to think of your:

> On profiling i saw that there are about 400 lucene merge threads for each index doing nothing. Just either thread among them is active and trying to merge segments.  
> The write threads seem to steadily write the data.

Because it's a narration instead of being a properly formatted output of a command in something like a gist. When looking for technical help, you have a better chance if you post technical questions with technical data, still not a guarantee 🙂 . Not sharing hardware spec is also an example, a pentium 166 with 2 spinning disk from 1994 in raid-0 is what I picture you're running since you didn't say. This goes deep, document example, index mapping, field count, what you tried, results obtained... "I added clients and couldn't past 2000doc/s, I added shards and ..., I added nodes and ... "

Hoping this helps,

---

_[View the full topic](https://discuss.elastic.co/t/too-many-lucene-merge-threads-while-indexing/186822)._
