# Elasticsearch version 5.4.3 .bulk insert with so many disk reads,but there is no merge operation the same time

**URL:** <https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390>\
**Category:** Elasticsearch\
**Created:** [April 23, 2020, 6:15am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390 "2020-04-23T06:15:13Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![wfeng888](https://avatars.discourse-cdn.com/v4/letter/w/9de053/32.png) [@wfeng888](https://discuss.elastic.co/u/wfeng888)\
**Post date:** [April 23, 2020, 6:15am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/1 "2020-04-23T06:15:13Z")

</div>

my cluster has one node，16 core cpu，64G total ram，28G jvm heap，index settings is:  
{  
"testshardread": {  
"settings": {  
"index": {  
"refresh\_interval": "-1",  
"indexing": {  
"slowlog": {  
"level": "info",  
"threshold": {  
"index": {  
"warn": "10s",  
"trace": "500ms",  
"debug": "2s",  
"info": "5s"  
}  
},  
"source": "1000"  
}  
},  
"number\_of\_shards": "8",  
"translog": {  
"flush\_threshold\_size": "10G",  
"sync\_interval": "60s",  
"durability": "async"  
},  
"provided\_name": "testshardread",  
"merge": {  
"scheduler": {  
"max\_thread\_count": "1"  
},  
"policy": {  
"max\_merged\_segment": "1G"  
}  
},  
"creation\_date": "1587559581106",  
"number\_of\_replicas": "0",  
"uuid": "SIRX8njWRM2h-\_LsMaiG3g",  
"version": {  
"created": "5040399"  
}  
}  
}  
}  
}  
and my question is why so many disk reads between 21:02 and 21:18 ,and 21:32 to21:46 。thands

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/4/1/411750d76f49cf5f13892e53e597a8ca5200e731.png)  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/0/1/0102b02ce84ef978d3eabc860e1e95a85cc94411.png)

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 23, 2020, 6:34am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/2 "2020-04-23T06:34:44Z")

</div>

Welcome!  
May be Lucene segments are merged?

---

<div class="post-metadata">

**Author:** ![wfeng888](https://avatars.discourse-cdn.com/v4/letter/w/9de053/32.png) [@wfeng888](https://discuss.elastic.co/u/wfeng888)\
**Post date:** [April 23, 2020, 6:59am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/3 "2020-04-23T06:59:41Z")

</div>

thans for reply. i think it's not . there has no meges operations during this times.i get the hot\_threads and index stats merges.total and merges.current real time , but i didn't find it.

---

<div class="post-metadata">

**Author:** ![wfeng888](https://avatars.discourse-cdn.com/v4/letter/w/9de053/32.png) [@wfeng888](https://discuss.elastic.co/u/wfeng888)\
**Post date:** [April 23, 2020, 7:04am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/4 "2020-04-23T07:04:10Z")

</div>

is there any chance that it's a disk read modify write problem?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 23, 2020, 7:25am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/5 "2020-04-23T07:25:53Z")

</div>

You said in the title that you are doing bulk operations at the same time. Which means lot of segments created and segment merged probably. But that's a guess.  
Do you have elastic monitor activated? That would help to understand.

I notice that you have an old version which is EOL now. You should really upgrade to 6.8 at least, better to go to 7.6.  
Is your cluster exposed to internet by any chance?

---

<div class="post-metadata">

**Author:** ![wfeng888](https://avatars.discourse-cdn.com/v4/letter/w/9de053/32.png) [@wfeng888](https://discuss.elastic.co/u/wfeng888)\
**Post date:** [April 23, 2020, 7:47am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/6 "2020-04-23T07:47:50Z")

</div>

Not so far and i will try.Thanks for your advice . About the version,eh, long story. Maybe we will upgrade at some time.  
My cluster running on intranet now.

---

<div class="post-metadata">

**Author:** ![wfeng888](https://avatars.discourse-cdn.com/v4/letter/w/9de053/32.png) [@wfeng888](https://discuss.elastic.co/u/wfeng888)\
**Post date:** [April 23, 2020, 11:42am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/7 "2020-04-23T11:42:07Z")

</div>

Hi,that are es monitor results.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/a/8/a80a3d1bf81662c65a63e318f9c0c3131a621626.png)  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/5/a/5a7577885fbc46d44ee60d8b522a8775014e7175.png)  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/6/9/69f1ed7bd3bf97b8b83ab1120f9a03a68b2e6a2c.png)  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/6/3/63b6d2d874b02928b35d11a3aa6447e16410d554.png)

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 23, 2020, 1:01pm UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/8 "2020-04-23T13:01:13Z")

</div>

We can see that you are indexing lot of documents but at some point the total number of segments is staying at around 120 segments after a ramp up.  
That clearly indicates to me that segments are merged in the background (which is expected). That implies lot of reads in addition to the writes.

Could you open one of the node in the monitoring app and show the segments. Should be something like this (7.6)

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/c/a/ca7929c7021fcc8f7b7c7d499e9d3a8a47606af8.jpeg)

---

<div class="post-metadata">

**Author:** ![wfeng888](https://avatars.discourse-cdn.com/v4/letter/w/9de053/32.png) [@wfeng888](https://discuss.elastic.co/u/wfeng888)\
**Post date:** [April 23, 2020, 4:00pm UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/9 "2020-04-23T16:00:02Z")

</div>

## Yes.I am doing a stress test by indexing a lot of documents as soon as possible . My cluster has only on node. here is the node segment count graph: ![image](https://us1.discourse-cdn.com/elastic/original/3X/b/7/b7f88c0b15262e95ea7e116e0d34545f6fc7d99b.png) here is hot\_treads info,( i'm sorry i cann't upload the entire log file). as you can see, there has few merge threads within 19:00 - 19:15 and 19:30 - 19:43 . These two period of time contribute the most disk reads: [root@localhost es\_monitor]# grep -E '2020-04-23-19|cpu usage by thread.\*Lucene Merge Thread' hot\_threads.log|grep -i -B1 merge|more 2020-04-23-19-07-36 - ::: {localhost.localdomain}{DYjsg8E8RkGXh-ceEPlFBw}{iM51ywJBTxC5ZK6h-VZ5VA}{10.45.154.78}{10.45.154.78:9300}{cname=localhost.localdomain} 48.3% (483ms out of 1s) cpu usage by thread 'elasticsearch[localhost.localdomain][[.monitoring-es-2-2020.04.23][0]: Lucene Merge Thread #60]'

## 2020-04-23-19-50-34 - ::: {localhost.localdomain}{DYjsg8E8RkGXh-ceEPlFBw}{iM51ywJBTxC5ZK6h-VZ5VA}{10.45.154.78}{10.45.154.78:9300}{cname=localhost.localdomain} 1.9% (19.3ms out of 1s) cpu usage by thread 'elasticsearch[localhost.localdomain][[testshardread][1]: Lucene Merge Thread #0]'

2020-04-23-19-50-44 - ::: {localhost.localdomain}{DYjsg8E8RkGXh-ceEPlFBw}{iM51ywJBTxC5ZK6h-VZ5VA}{10.45.154.78}{10.45.154.78:9300}{cname=localhost.localdomain}  
10.1% (100.7ms out of 1s) cpu usage by thread 'elasticsearch[localhost.localdomain][[testshardread][1]: Lucene Merge Thread #0]'  
2020-04-23-19-50-48 - ::: {localhost.localdomain}{DYjsg8E8RkGXh-ceEPlFBw}{iM51ywJBTxC5ZK6h-VZ5VA}{10.45.154.78}{10.45.154.78:9300}{cname=localhost.localdomain}  
10.6% (106ms out of 1s) cpu usage by thread 'elasticsearch[localhost.localdomain][[testshardread][1]: Lucene Merge Thread #0]'  
2020-04-23-19-50-50 - ::: {localhost.localdomain}{DYjsg8E8RkGXh-ceEPlFBw}{iM51ywJBTxC5ZK6h-VZ5VA}{10.45.154.78}{10.45.154.78:9300}{cname=localhost.localdomain}  
8.5% (85.2ms out of 1s) cpu usage by thread 'elasticsearch[localhost.localdomain][[testshardread][1]: Lucene Merge Thread #0]'  
2020-04-23-19-50-53 - ::: {localhost.localdomain}{DYjsg8E8RkGXh-ceEPlFBw}{iM51ywJBTxC5ZK6h-VZ5VA}{10.45.154.78}{10.45.154.78:9300}{cname=localhost.localdomain}  
9.2% (91.9ms out of 1s) cpu usage by thread 'elasticsearch[localhost.localdomain][[testshardread][1]: Lucene Merge Thread #0]'  
2020-04-23-19-50-57 - ::: {localhost.localdomain}{DYjsg8E8RkGXh-ceEPlFBw}{iM51ywJBTxC5ZK6h-VZ5VA}{10.45.154.78}{10.45.154.78:9300}{cname=localhost.localdomain}  
9.4% (94.4ms out of 1s) cpu usage by thread 'elasticsearch[localhost.localdomain][[testshardread][1]: Lucene Merge Thread #0]'  
2020-04-23-19-51-00 - ::: {localhost.localdomain}{DYjsg8E8RkGXh-ceEPlFBw}{iM51ywJBTxC5ZK6h-VZ5VA}{10.45.154.78}{10.45.154.78:9300}{cname=localhost.localdomain}  
8.7% (86.8ms out of 1s) cpu usage by thread 'elasticsearch[localhost.localdomain][[testshardread][1]: Lucene Merge Thread #0]'  
2020-04-23-19-51-02 - ::: {localhost.localdomain}{DYjsg8E8RkGXh-ceEPlFBw}{iM51ywJBTxC5ZK6h-VZ5VA}{10.45.154.78}{10.45.154.78:9300}{cname=localhost.localdomain}  
3.7% (37.2ms out of 1s) cpu usage by thread 'elasticsearch[localhost.localdomain][[testshardread][1]: Lucene Merge Thread #0]'  
2020-04-23-19-51-04 - ::: {localhost.localdomain}{DYjsg8E8RkGXh-ceEPlFBw}{iM51ywJBTxC5ZK6h-VZ5VA}{10.45.154.78}{10.45.154.78:9300}{cname=localhost.localdomain}  
4.6% (45.5ms out of 1s) cpu usage by thread 'elasticsearch[localhost.localdomain][[testshardread][1]: Lucene Merge Thread #0]'  
2020-04-23-19-51-07 - ::: {localhost.localdomain}{DYjsg8E8RkGXh-ceEPlFBw}{iM51ywJBTxC5ZK6h-VZ5VA}{10.45.154.78}{10.45.154.78:9300}{cname=localhost.localdomain}  
1.2% (11.9ms out of 1s) cpu usage by thread 'elasticsearch[localhost.localdomain][[testshardread][1]: Lucene Merge Thread #0]'

I'd love to believe that too much disk reads because of merge operation,but i can not find evidence to prove it.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 24, 2020, 2:53am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/10 "2020-04-24T02:53:01Z")

</div>

I have no other ideas. I don't think replication in your scenario does that. A way to eliminate this option would be to have 0 replicas and see.

@jpountz any other idea?

---

<div class="post-metadata">

**Author:** ![wfeng888](https://avatars.discourse-cdn.com/v4/letter/w/9de053/32.png) [@wfeng888](https://discuss.elastic.co/u/wfeng888)\
**Post date:** [April 24, 2020, 3:00am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/11 "2020-04-24T03:00:06Z")

</div>

it already has no replicas any more. Settings are here:  
{  
"testshardread": {  
"settings": {  
"index": {  
"refresh\_interval": "-1",  
"indexing": {  
"slowlog": {  
"level": "info",  
"threshold": {  
"index": {  
"warn": "10s",  
"trace": "500ms",  
"debug": "2s",  
"info": "5s"  
}  
},  
"source": "1000"  
}  
},  
"number\_of\_shards": "8",  
"translog": {  
"flush\_threshold\_size": "10G",  
"sync\_interval": "60s",  
"durability": "async"  
},  
"provided\_name": "testshardread",  
"merge": {  
"scheduler": {  
"max\_thread\_count": "1"  
},  
"policy": {  
"max\_merged\_segment": "1G"  
}  
},  
"creation\_date": "1587638733557",  
"number\_of\_replicas": "0",  
"uuid": "mGRhj\_S5Q8StAO1rUSh4fA",  
"version": {  
"created": "5040399"  
}  
}  
}  
}  
}

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [April 24, 2020, 2:21pm UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/12 "2020-04-24T14:21:55Z")

</div>

Can you share the entire content of the hot threads? Do I understand correctly that these hot threads were captured during this time when you observed a peak of reads?

---

<div class="post-metadata">

**Author:** ![wfeng888](https://avatars.discourse-cdn.com/v4/letter/w/9de053/32.png) [@wfeng888](https://discuss.elastic.co/u/wfeng888)\
**Post date:** [April 24, 2020, 3:50pm UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/13 "2020-04-24T15:50:07Z")

</div>

yes.it's about 35M size.Shoud i send it to your email address?

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [April 24, 2020, 8:07pm UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/14 "2020-04-24T20:07:48Z")

</div>

35MB of hot threads with a single node? Why is it so large, have you configured very large threadpools?

---

<div class="post-metadata">

**Author:** ![wfeng888](https://avatars.discourse-cdn.com/v4/letter/w/9de053/32.png) [@wfeng888](https://discuss.elastic.co/u/wfeng888)\
**Post date:** [April 26, 2020, 12:47am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/15 "2020-04-26T00:47:02Z")

</div>

No，it's not.I did a 1s sampling and output the results to a file like this:  
curl -X GET "10.45.154.78:9200/\_nodes/hot\_threads?interval=1s&threads=5" \> hot\_threads.log  
it's hot\_threads.log file that about 35MB.

---

<div class="post-metadata">

**Author:** ![wfeng888](https://avatars.discourse-cdn.com/v4/letter/w/9de053/32.png) [@wfeng888](https://discuss.elastic.co/u/wfeng888)\
**Post date:** [April 28, 2020, 1:17am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/16 "2020-04-28T01:17:57Z")

</div>

Hi,Any advice?

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [April 28, 2020, 7:36am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/17 "2020-04-28T07:36:51Z")

</div>

OK, please send it to adrien (at) [elastic.co](http://elastic.co).

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [April 28, 2020, 9:52am UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/18 "2020-04-28T09:52:21Z")

</div>

I'm seeing the following reads in the hot threads:

- Compound files: segments that are created by a refresh are created as a compound file (similar to a tar archive), which is implemented by first writing every file to disk as temporary files and then creating the archive from these files. These reads are usually small and from data that has just been read, so the vast majority of the time it doesn't actually go to the disk but only to the filesystem cache.
- Monitoring is periodically computing the size of files, which reads file attributes. For the record this is something that we cache more aggressively since version 6.4: [https://github.com/elastic/elasticsearch/pull/30581](https://github.com/elastic/elasticsearch/pull/30581).
- I am seeing merges in the hot threads. Merges are especially likely to trigger reads due to the fact that they read the incoming segments one time to verify checksums, and then a second time to actually merge the data.

---

<div class="post-metadata">

**Author:** ![wfeng888](https://avatars.discourse-cdn.com/v4/letter/w/9de053/32.png) [@wfeng888](https://discuss.elastic.co/u/wfeng888)\
**Post date:** [April 28, 2020, 2:40pm UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/19 "2020-04-28T14:40:37Z")

</div>

Nice, thanks a lot. I will study the first two carefully. Any Suggestions to solving the problem？

As to the third, there has merges actually,but it is not in the serious disk reading period.In fact, it happens in the very beginning place, and the others too later. So it doesn't really matter. Am i right?

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [April 28, 2020, 4:51pm UTC](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390/20 "2020-04-28T16:51:10Z")

</div>

First it's not clear to me there is a problem at all. Were the hot threads you shared with me computed during the period during which you are observing lots of reads?

[Next page](https://discuss.elastic.co/t/elasticsearch-version-5-4-3-bulk-insert-with-so-many-disk-reads-but-there-is-no-merge-operation-the-same-time/229390.md?page=2)
