# Reducing Memory Consuption while Bulk Indexing

**URL:** <https://discuss.elastic.co/t/reducing-memory-consuption-while-bulk-indexing/26447>\
**Category:** Elasticsearch\
**Created:** [July 28, 2015, 9:18pm UTC](https://discuss.elastic.co/t/reducing-memory-consuption-while-bulk-indexing/26447 "2015-07-28T21:18:52Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![klahnakoski](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klahnakoski/32/3679_2.png) [@klahnakoski](https://discuss.elastic.co/u/klahnakoski)\
**Post date:** [July 28, 2015, 9:18pm UTC](https://discuss.elastic.co/t/reducing-memory-consuption-while-bulk-indexing/26447/1 "2015-07-28T21:18:52Z")

</div>

I seems bulk indexing is using a lot of memory. It is not for long, but the day sees much more indexing demand than shown here. Is there a way I can reduce memory required during indexing, while still keeping the indexing rate high?

 ![](https://us1.discourse-cdn.com/elastic/original/2X/d/d1a4cd147bbe7f2d9ca6422390038c93a7e4a839.png)

Version 1.4.2

Maybe my \_routing is causing this?

> "mappings": { "test\_result": { "\_routing": { "path": "build.revision12","required": true}, ....

Here is my yml file:

> cluster.name: active-data  
> node.zone: primary  
> node.name: primary  
> node.master: true  
> node.data: true

> cluster.routing.allocation.awareness.force.zone.values: primary,spot  
> cluster.routing.allocation.awareness.attributes: zone  
> cluster.routing.allocation.cluster\_concurrent\_rebalance: 1  
> cluster.routing.allocation.balance.shard: 0.70  
> cluster.routing.allocation.balance.primary: 0.05

> bootstrap.mlockall: true  
> path.data: /data1, /data2, /data3  
> path.logs: /data1/logs  
> cloud:  
> aws:  
> region: us-west-2  
> protocol: https  
> ec2:  
> protocol: https  
> discovery.type: ec2  
> discovery.zen.ping.multicast.enabled: false  
> discovery.zen.minimum\_master\_nodes: 1

> index.number\_of\_shards: 1  
> index.number\_of\_replicas: 1  
> index.cache.field.type: soft  
> index.translog.interval: 60s  
> index.translog.flush\_threshold\_size: 1gb

> indices.memory.index\_buffer\_size: 20%  
> indices.fielddata.cache.expire: 20m  
> indices.recovery.concurrent\_streams: 1  
> indices.recovery.max\_bytes\_per\_sec: 1000mb

> http.compression: true  
> http.cors.allow-origin: "/.\*/"  
> http.cors.enabled: true  
> http.compression: true  
> http.max\_content\_length: 1000mb  
> http.timeout: 600

> threadpool.bulk.queue\_size: 3000  
> threadpool.index.queue\_size: 1000

Thanks

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [July 28, 2015, 9:57pm UTC](https://discuss.elastic.co/t/reducing-memory-consuption-while-bulk-indexing/26447/2 "2015-07-28T21:57:17Z")

</div>

Increasing your bulk threadpool will only be increasing that memory use.

How big are you bulk sizings? How many nodes in the cluster?

That aside I don't see a major problem, you aren't really approaching your heap maximum and things appear pretty good in terms of resource usage.

---

<div class="post-metadata">

**Author:** ![klahnakoski](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klahnakoski/32/3679_2.png) [@klahnakoski](https://discuss.elastic.co/u/klahnakoski)\
**Post date:** [July 29, 2015, 1:23am UTC](https://discuss.elastic.co/t/reducing-memory-consuption-while-bulk-indexing/26447/3 "2015-07-29T01:23:58Z")

</div>

There are 10 nodes in the cluster, but this node is the "primary" containing all shards. The 9 other nodes are recipients of bulk index requests (and containing some replica shards). Each of the 9 have a daemon inserting 5000 documents per bulk request. Each document is about 1K. It works out to a rate of about 400\*9 documents per second. Faster would be nicer, but this primary node is the bottleneck.

Here is an example of what it looks like just before ES dies:

 ![](https://us1.discourse-cdn.com/elastic/original/2X/8/8bc69414f8f2b2e3fb22656974c353cd522e8078.png)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [July 29, 2015, 1:28am UTC](https://discuss.elastic.co/t/reducing-memory-consuption-while-bulk-indexing/26447/4 "2015-07-29T01:28:07Z")

</div>

Why do you have all the primaries on a single node then?

---

<div class="post-metadata">

**Author:** ![klahnakoski](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klahnakoski/32/3679_2.png) [@klahnakoski](https://discuss.elastic.co/u/klahnakoski)\
**Post date:** [July 29, 2015, 11:27am UTC](https://discuss.elastic.co/t/reducing-memory-consuption-while-bulk-indexing/26447/5 "2015-07-29T11:27:05Z")

</div>

That is the only node in that zone right now, as per the \*.yml file.

If I understood why ES is consuming so much memory during bulk indexing, I may be able to mitigate the problem.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [July 29, 2015, 11:37am UTC](https://discuss.elastic.co/t/reducing-memory-consuption-while-bulk-indexing/26447/6 "2015-07-29T11:37:38Z")

</div>

It'll use what it needs.  
But not spreading the primaries, and thus the resource consumption, around is going to be part of your problem.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 29, 2015, 11:45am UTC](https://discuss.elastic.co/t/reducing-memory-consuption-while-bulk-indexing/26447/7 "2015-07-29T11:45:07Z")

</div>

If you are using 1 replica and two zones, where one of the zones only have a single node, this will be indexing all records, which will lead to much higher load than the other nodes. Elasticsearch is usually deployed in clusters where all nodes are equal so that load can be evenly distributed. What is it you are trying to achieve using this setup?

---

<div class="post-metadata">

**Author:** ![klahnakoski](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klahnakoski/32/3679_2.png) [@klahnakoski](https://discuss.elastic.co/u/klahnakoski)\
**Post date:** [July 29, 2015, 2:29pm UTC](https://discuss.elastic.co/t/reducing-memory-consuption-while-bulk-indexing/26447/8 "2015-07-29T14:29:05Z")

</div>

I am aware that a node with more shards will have more load. I am concerned with what I can do to reduce indexing load, or move the work to the other nodes. Maybe the fact this one node is master? Should I assign the master to be in the other zone? Maybe a slave with all shards will use less memory?

My setup is designed to achieve minimal price on AWS. The benefit to having one machine in a zone is it's 1/2 the cost of two machines, and 1/3 the cost of three machines.

---

<div class="post-metadata">

**Author:** ![klahnakoski](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klahnakoski/32/3679_2.png) [@klahnakoski](https://discuss.elastic.co/u/klahnakoski)\
**Post date:** [July 29, 2015, 5:57pm UTC](https://discuss.elastic.co/t/reducing-memory-consuption-while-bulk-indexing/26447/9 "2015-07-29T17:57:18Z")

</div>

I added a "coordinator node" in the same zone as the 9 nodes. This coordinator is the master (node.master: true), and has no shards (node.data: false). Here is the coordinator under heavy indexing load:

 ![](https://us1.discourse-cdn.com/elastic/original/2X/7/7696279c8f8353a206c076d7fc4b17327e19d166.png)

It is doing nothing, as expected.

The primary node is the same, it still has all shards, but was restarted with (node.master: false). Here is it under higher load:

 ![](https://us1.discourse-cdn.com/elastic/original/2X/5/5db451a36e370570c7fdbc535127dd00f4256ff8.png)

What's important to see is the rate of memory consumption is much smaller. Hopefully this leads to less OOM events that crash this node. I believe the growth in memory is proportional to the number of primary shards on a node. By moving the master to the other zone, it appears my primary node is assigned less primary shards, and so has lower memory consumption.

This is not perfect, eventually the primary node's shards are designated primary, and the memory consumption rate goes up accordingly. I hope I can control the algorithm that determines what shard is declared primary; preferring primaries be assigned to the zone with the most nodes.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:58pm UTC](https://discuss.elastic.co/t/reducing-memory-consuption-while-bulk-indexing/26447/10 "2017-07-05T23:58:17Z")

</div>


