# Horizontal scaling of indexing

**URL:** <https://discuss.elastic.co/t/horizontal-scaling-of-indexing/31962>\
**Category:** Elasticsearch\
**Created:** [October 10, 2015, 10:20pm UTC](https://discuss.elastic.co/t/horizontal-scaling-of-indexing/31962 "2015-10-10T22:20:10Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![tinle](https://avatars.discourse-cdn.com/v4/letter/t/c77e96/32.png) [@tinle](https://discuss.elastic.co/u/tinle)\
**Post date:** [October 11, 2015, 12:09am UTC](https://discuss.elastic.co/t/horizontal-scaling-of-indexing/31962/2 "2015-10-11T00:09:17Z")

</div>

Glad to hear someone else seeing same problems we are seeing. We're using slightly different HW with similar results.

See my old post here:

> [@Indexing speed in ES v1.71 approx 18% slower than v1.4.5](https://discuss.elastic.co/t/indexing-speed-in-es-v1-71-approx-18-slower-than-v1-4-5/28323/13):
>
> Hmm my response (via email) was truncated for some reason ... trying again: I think (not certain) for logstash it's the number of workers you specify for the Elasticsearch output? I think the default is 1: [https://www.elastic.co/guide/en/logstash/current/plugins-outputs-elasticsearch.html](https://www.elastic.co/guide/en/logstash/current/plugins-outputs-elasticsearch.html) Or in your custom Go client, you would control how many threads (goroutines?) are sending bulk indexing requests concurrently. You can ask for node stats, then look under thread\_pool -\> bulk -\> active to se…

Once we got past the testing harness setup, we were able to reproduce the slow indexing performance internally.

Our HW is:

Virident PCIe SSD card (config for performance) 1.8TB  
64GB RAM  
2x12 core Xeon (HT on, or equiv of 48) (5 physical bare metal nodes x 2 sets for faster testing of various parameters combination)  
ES v1.7.2  
JDK 8u60 (also tested with JDK7u51, JDK8u40)  
Tested various maxheap from 16G to 31G.  
mlockall on  
max fd is 64K  
refresh interval is -1  
index.store.throttle.type: none  
index.store.throttle.max\_bytes\_per\_sec: 700mb  
index.translog.flush\_threshold\_size: 1gb  
indices.memory.index\_buffer\_size: 512mb  
5 shards so we get 1 per node  
no replica  
various doc size from 1k to 16K  
same data set on a RAMdisk so we always read same data via logstash file input  
tested with 1 LS instance, 5 instances, 20 instances, etc.  
Various bulk indexing sizes (100, 500, 1000, 5000, 10000, etc.).

Our conclusion is that I/O, CPU and memory are not the problem. We always hit a limit in how fast ES can index.

How are you ingesting data? Logstash? or your own client doing bulk insert? You can try increasing the number of instances feeding ES. We notice a slight increase in indexing speed, but it falls off after 10 concurrent LS instances into the 5 ES nodes.

---

_[View the full topic](https://discuss.elastic.co/t/horizontal-scaling-of-indexing/31962)._
