# Bulk Indexing Rate

**URL:** <https://discuss.elastic.co/t/bulk-indexing-rate/124973>\
**Category:** Elasticsearch\
**Created:** [March 21, 2018, 11:48am UTC](https://discuss.elastic.co/t/bulk-indexing-rate/124973 "2018-03-21T11:48:43Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![wpearson4](https://avatars.discourse-cdn.com/v4/letter/w/cdc98d/32.png) [@wpearson4](https://discuss.elastic.co/u/wpearson4)\
**Post date:** [March 21, 2018, 11:48am UTC](https://discuss.elastic.co/t/bulk-indexing-rate/124973/1 "2018-03-21T11:48:43Z")

</div>

I am running a 4 node cluster 2 data nodes 1.5 TB SSD each, 64gb Ram Each and Hexacore Processor,  
1 Master node 16gb ram 4 core processor, 1 Client node 16gb ram 4 core processor. I am bulking index to http endpoint \_bulk. I am currently only able to index 200k documents every 111 seconds. I am bulk indexing directly to 1 of my data nodes. This seems awfully slow. If I try to increase threads I start running into

es\_rejected\_execution\_exception : rejected execution of org.elasticsearch.transport.TransportService$7@1736e37a on EsThreadPoolExecutor[name = ecluster01-dalc1-prod-data02/bulk, queue capacity = 200, org.elasticsearch.common.util.concurrent.EsThreadPoolExecutor@5e8e7774[Running, pool size = 24, active threads = 24, queued tasks = 200, completed tasks = 374618]]

What can I be doing wrong? A single document is ~10k. I am running 10k batches of documents in my post.

Thank in advance

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 21, 2018, 11:56am UTC](https://discuss.elastic.co/t/bulk-indexing-rate/124973/2 "2018-03-21T11:56:18Z")

</div>

It is generally recommended to keep the size of bulk requests around 5MB or so. Larger bulk sizes does not necessarily result in better throughput. 10k documents at 10kB each is obviously much larger than that (~100MB).

How many concurrent indexing threads do you use?

Have you tried to identify what is limiting throughput? Is it CPU, disk I/O and iowait, GC?

---

<div class="post-metadata">

**Author:** ![wpearson4](https://avatars.discourse-cdn.com/v4/letter/w/cdc98d/32.png) [@wpearson4](https://discuss.elastic.co/u/wpearson4)\
**Post date:** [March 21, 2018, 2:44pm UTC](https://discuss.elastic.co/t/bulk-indexing-rate/124973/3 "2018-03-21T14:44:54Z")

</div>

1. Should I multi-thread to a single node?
2. Currently, I am sending simultaneous requests to other nodes in my cluster. I am currently indexing to the data nodes directly. Would it be better to index to my client node and hit that node with multiple threads?
3. There seems to be a big performance hit incurred when dealing with the response from an index request. I assume that will be mitigated when I lower the request down to 5mb?
4. No, I have not identified the bottleneck. Can I retrieve some of the stats from ES directly?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 21, 2018, 2:48pm UTC](https://discuss.elastic.co/t/bulk-indexing-rate/124973/4 "2018-03-21T14:48:14Z")

</div>

> [@wpearson4](#):
>
> Should I multi-thread to a single node?

Reduce the bulk size and try multiple concurrent connections. Gradually increase the level of concurrency until you see no further improvement in throughput.

> [@wpearson4](#):
>
> Currently, I am sending simultaneous requests to other nodes in my cluster. I am currently indexing to the data nodes directly. Would it be better to index to my client node and hit that node with multiple threads?

You can continue sending indexing requests to the 2 data nodes.

> [@wpearson4](#):
>
> There seems to be a big performance hit incurred when dealing with the response from an index request. I assume that will be mitigated when I lower the request down to 5mb?

Probably.

> [@wpearson4](#):
>
> No, I have not identified the bottleneck. Can I retrieve some of the stats from ES directly?

I would recommend installing X-Pack monitoring to get a better idea about how your cluster is performing.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 18, 2018, 2:48pm UTC](https://discuss.elastic.co/t/bulk-indexing-rate/124973/5 "2018-04-18T14:48:27Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
