# Indexing speed: Config changes don't seem to have any effect

**URL:** <https://discuss.elastic.co/t/indexing-speed-config-changes-dont-seem-to-have-any-effect/50703>\
**Category:** Elasticsearch\
**Created:** [May 23, 2016, 11:41am UTC](https://discuss.elastic.co/t/indexing-speed-config-changes-dont-seem-to-have-any-effect/50703 "2016-05-23T11:41:20Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ac-inv](https://avatars.discourse-cdn.com/v4/letter/a/eb9ed0/32.png) [@ac-inv](https://discuss.elastic.co/u/ac-inv)\
**Post date:** [May 23, 2016, 11:41am UTC](https://discuss.elastic.co/t/indexing-speed-config-changes-dont-seem-to-have-any-effect/50703/1 "2016-05-23T11:41:20Z")

</div>

Hi,

We have an elasticsearch cluster deployed on AWS. We are using 1.7.5. The cluster contains 6 data nodes and 3 master nodes.  
The 6 data nodes are running on r3.xlarge instances (4 vCPU, 30.5G memory, 100G gp2 SSD).  
The master nodes are running on t2.medium instances (2 vCPU, 4G memory)

The elasticseach.yml file for the data nodes looks like this:

> bootstrap.mlockall: true  
> cluster.name: dev1  
> cluster.routing.allocation.node\_concurrent\_recoveries: 4  
> cluster.routing.allocation.awareness.attributes: aws\_availability\_zone  
> node.master: false  
> node.data: true  
> discovery.type: ec2  
> discovery.ec2.minimum\_master\_nodes: 2  
> discovery.ec2.tag.elasticsearch\_cluster: dev1  
> cloud.aws.region: us-west-2  
> cloud.node.auto\_attributes: true  
> path.data: /data  
> action.write\_consistency: one

> index.search.slowlog.threshold.query.warn: 10s  
> index.search.slowlog.threshold.query.info: 5s  
> index.search.slowlog.threshold.query.debug: 2s  
> index.search.slowlog.threshold.query.trace: 500ms

> index.search.slowlog.threshold.fetch.warn: 1s  
> index.search.slowlog.threshold.fetch.info: 800ms  
> index.search.slowlog.threshold.fetch.debug: 500ms  
> index.search.slowlog.threshold.fetch.trace: 200ms

> http.cors.enabled: true  
> http.cors.allow-origin: /https?://localhost(:[0-9]+)?/

> indices.breaker.total.limit: 90%  
> indices.breaker.fielddata.limit: 80%  
> indices.breaker.request.limit: 20%

> indices.fielddata.cache.size: 70%

> script.groovy.sandbox.enabled: true

There are multiple indexes to which the data gets indexed to. All indexes have the same mappings. We use 3 shards and 2 replicas per index.  
The document type that gets indexed (elements) has nested documents (properties). The number of properties per element can vary from tens to few hundreds. Each property is less than 1 KB. The element minus the properties can be few KBs to tens of KBs. We use bulk requests to index the data. The bulk indexing batch size is 10MB which on average contains 500 elements and 40k properties. These batches are indexed concurrently with 20 threads on different machines indexing 100 batches each.. The documents get indexed to different indexes. The bulk indexing requests are sent to an elastic load balancer which forwards the requests to the data nodes. The rate at which the data gets indexed is about 49K individual documents (elements + properties) per second. Given the resources available to the data nodes and the fact that the document sizes are small, I feel that this rate is slow. The refresh interval is set to the default of 1s.

I have tried a number of configuration changes that are recommended like changing the bulk indexing buffer size, testing with different indexing batch sizes, using client nodes etc but we seem to get the same indexing throughput of 49k docs per second. I haven't been able to see any effect of configuration changes on the throughput. The CPU usage is quite high on the data nodes during indexing (almost 99-100%). The indices store throttle limit is set to the default 20MB/sec. The recommended limit for SSD is 100MB/sec but we don't seem to be hitting that yet.

Any pointers on how to improve the indexing throughput would be greatly appreciated.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 23, 2016, 12:04pm UTC](https://discuss.elastic.co/t/indexing-speed-config-changes-dont-seem-to-have-any-effect/50703/2 "2016-05-23T12:04:43Z")

</div>

Based on your description, it sounds like performance is limited by the amount of available CPU rather than disk I/O. If this is the case, you should be able to increase the throughput by reducing the number of replicas. It may also be worthwhile experiment with an increased refresh interval.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:49pm UTC](https://discuss.elastic.co/t/indexing-speed-config-changes-dont-seem-to-have-any-effect/50703/3 "2017-07-05T22:49:48Z")

</div>


