# Bulk import response times

**URL:** <https://discuss.elastic.co/t/bulk-import-response-times/27918>\
**Category:** Elasticsearch\
**Created:** [August 24, 2015, 5:06am UTC](https://discuss.elastic.co/t/bulk-import-response-times/27918 "2015-08-24T05:06:39Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Srinath\_C](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/srinath_c/32/48806_2.png) [@Srinath\_C](https://discuss.elastic.co/u/Srinath_C)\
**Post date:** [August 24, 2015, 5:06am UTC](https://discuss.elastic.co/t/bulk-import-response-times/27918/1 "2015-08-24T05:06:39Z")

</div>

Hi,

We are observing increasing response times to bulk index requests with passage of time.

Setup: ES version 1.7.1, 3 master nodes, 5 data nodes (m3.xlarge - 4 core, 7.5g heap space, 15g ram, data written to 2 SSD instance store drives of 37g each).  
Test setup has 3 [NodeClient instances][1], making bulk requests of 5Mb and around ~2900 documents (1.7k each). The bulk inserts are sychronous (second request is made after completion for first)  
Data is being indexed into 3 indices each having 3 primary shards and async replication of 1.

Observations:  
After about 1 hour we see increased response times from each bulk request.  
CPU Utilization across the data nodes is \< 20%.

iostat (more or less the same across the nodes):  
Filesystem Size Used Avail Use% Mounted on  
/dev/xvdf 37G 1.9G 34G 6% /data  
/dev/xvdg 37G 1.9G 34G 6% /logs

Initially, we see response times of 1.5-2s but after about 1.5 hours the response times are now 6-8 seconds.

Settings: Mostly defaults, except for:  
index.merge.policy.type: tiered  
index.merge.scheduler.type: concurrent  
index.refresh\_interval: 20s  
index.translog.flush\_threshold\_ops: 50000  
index.translog.interval: 10s  
index.warmer.enabled: false  
indices.fielddata.cache.size: 10%  
indices.memory.index\_buffer\_size: 30%  
indices.store.throttle.type: none

Questions:

1. Does the amount of data already in the index affect the bulk ingestion rate?
2. How can we better utilize the resources?
3. With asynchronous call backs using

![](https://us1.discourse-cdn.com/elastic/original/2X/e/e28c370e8a607c732db201cc3c103c7c662cedb3.png)

 ![](https://us1.discourse-cdn.com/elastic/original/2X/0/03413e44d1905a0e59a45595fb85711c8f81d9ee.png)  
[1]: [https://gist.github.com/Srinathc/6a4c520f3e025aaea017#file-estest-java](https://gist.github.com/Srinathc/6a4c520f3e025aaea017#file-estest-java)

---

<div class="post-metadata">

**Author:** ![Srinath\_C](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/srinath_c/32/48806_2.png) [@Srinath\_C](https://discuss.elastic.co/u/Srinath_C)\
**Post date:** [August 26, 2015, 2:08am UTC](https://discuss.elastic.co/t/bulk-import-response-times/27918/2 "2015-08-26T02:08:45Z")

</div>

Attaching snapshots for merge activity at that point of time.

 ![](https://us1.discourse-cdn.com/elastic/original/2X/4/41e45523d014f264025dc69bd8a242f0e0031eb1.png)

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [August 26, 2015, 12:31pm UTC](https://discuss.elastic.co/t/bulk-import-response-times/27918/3 "2015-08-26T12:31:44Z")

</div>

If your index was inactive then the first bulks are faster because there is no merging activity, it's even better if the index was empty since the first merges are very cheap. However, as time goes, elasticsearch needs to make sure that merging can keep up with merging so you might indeed see indexing slowing down.

Since your cluster doesn't seem to be completely utilized, you could look into sending indexing requests from more parallel workers/threads.

---

<div class="post-metadata">

**Author:** ![Srinath\_C](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/srinath_c/32/48806_2.png) [@Srinath\_C](https://discuss.elastic.co/u/Srinath_C)\
**Post date:** [August 31, 2015, 3:04pm UTC](https://discuss.elastic.co/t/bulk-import-response-times/27918/4 "2015-08-31T15:04:32Z")

</div>

thanks @jpountz

Actually, the problem was with the "Open search contexts". Some of these search contexts were doing:

```
SearchRequestBuilder searchRequestBuilder = client.prepareSearch(aggregator.mapToIndex(aggregateRequest))
                    .setSearchType(SearchType.QUERY_THEN_FETCH)
                    .setTypes(aggregator.mapToType(tuple))
                    .setQuery(aggregator.createQuery(aggregateRequest))
                    .setFrom(0).setSize(Integer.MAX_VALUE);

```

The .setSize(Integer.MAX\_VALUE) was actually the culprit and setting it to a saner value fixed it.  
The response times were of the order of 100-250ms initially but after merge kicked in it went up to 700-900ms.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:53pm UTC](https://discuss.elastic.co/t/bulk-import-response-times/27918/5 "2017-07-05T23:53:03Z")

</div>


