# ES creating thousands of segments with 1 document each

**URL:** <https://discuss.elastic.co/t/es-creating-thousands-of-segments-with-1-document-each/31849>\
**Category:** Elasticsearch\
**Created:** [October 8, 2015, 1:52pm UTC](https://discuss.elastic.co/t/es-creating-thousands-of-segments-with-1-document-each/31849 "2015-10-08T13:52:31Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Herve\_Bry](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/herve_bry/32/12706_2.png) [@Herve\_Bry](https://discuss.elastic.co/u/Herve_Bry)\
**Post date:** [October 8, 2015, 1:52pm UTC](https://discuss.elastic.co/t/es-creating-thousands-of-segments-with-1-document-each/31849/1 "2015-10-08T13:52:31Z")

</div>

Hello Community,

I am calling for help as we are struggling with a strange problem: often when indexing a bulk of data, thousands of tiny segments with only one doc in each are created. This brings the cluster to its knees by consuming all the CPU on the impacted nodes, often disrupting the service.

Here is the situation:

- We have 3 servers (16 cores/1.2TB SSD/128GB RAM each) with 3 instances of ES 1.7.1 each (24G RAM per instance)
- Our index holds 1.6 Billion records in 10 shards with 1 replica, totalizing 1.2TB on disk
- This makes about 700 segments in normal conditions
- It is queried at about 150-200 queries per second
- Every hour, a few millions of records are added or updated, in bulk mode, using 8 parallel connections. This takes between 5 and 20 min.

The problem arise every hour during the bulk indexing. During the first few minutes, a huge number of segments are created (I have seen up to 6000) that hold only one document each. After that, the ES instance holding those segments use so much CPU that it is almost unresponsive, which slows own the other instances on the same server and sometimes even disrupts the service.

After a few minutes of very heavy CPU usage, the segments are finally merged (the count goes down to a normal ~ 700) and everything goes back to normal.

Is this a bug ? It appears to me that a segment should rarely hold only one doc...  
Do you have advice on what settings to tune to avoid this problem ? We have already tried different refresh intervals (-1, 1s, 10s, 30s) and different merge throttling throughputs (from unlimited down to 10MBps) but the problem still occurs almost every time.

Thanks for your wisdom !

Hervé BRY

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 8, 2015, 2:07pm UTC](https://discuss.elastic.co/t/es-creating-thousands-of-segments-with-1-document-each/31849/2 "2015-10-08T14:07:41Z")

</div>

That sounds odd. Can you check the index settings? Have by any chance set _index.translog.flush\_threshold\_ops_ or _index.translog.flush\_threshold\_size_ to an inappropriate value?

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [October 8, 2015, 2:10pm UTC](https://discuss.elastic.co/t/es-creating-thousands-of-segments-with-1-document-each/31849/3 "2015-10-08T14:10:16Z")

</div>

In addition to @Christian_Dahlqvist's ideas, have a look at this:

> <https://github.com/elastic/elasticsearch/pull/13918>

It _might_ be what you are seeing. Maybe.

You can test this by pushing a single document into the index before you do the bulk index and then waiting 45 seconds or so and then doing the bulk load.

---

<div class="post-metadata">

**Author:** ![Herve\_Bry](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/herve_bry/32/12706_2.png) [@Herve\_Bry](https://discuss.elastic.co/u/Herve_Bry)\
**Post date:** [October 8, 2015, 2:30pm UTC](https://discuss.elastic.co/t/es-creating-thousands-of-segments-with-1-document-each/31849/4 "2015-10-08T14:30:57Z")

</div>

Thanks for your suggestions.

@Christian_Dahlqvist: there are no specific translog options set for the index. Here is the config we use :

> <https://gist.github.com/setaou/1871dc5a8a7a6f23c633>

@nik9000: We are going to try your suggestion. It indeed seems like it can be the cause of our problem. Is there any chance this PR might be merged in ES 1.x ?

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [October 8, 2015, 2:34pm UTC](https://discuss.elastic.co/t/es-creating-thousands-of-segments-with-1-document-each/31849/5 "2015-10-08T14:34:44Z")

</div>

> [@Herve\_Bry](#):
>
> @nik9000: We are going to try your suggestion. It indeed seems like it can be the cause of our problem. Is there any chance this PR might be merged in ES 1.x ?

I doubt it. Its pretty deep in the 2.0 line so it'd be quite difficult.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:45pm UTC](https://discuss.elastic.co/t/es-creating-thousands-of-segments-with-1-document-each/31849/6 "2017-07-05T23:45:53Z")

</div>


