# Bulk Indexing performance on AWS ES service

**URL:** <https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011>\
**Category:** Elasticsearch\
**Created:** [November 1, 2017, 9:55am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011 "2017-11-01T09:55:55Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![michaels](https://avatars.discourse-cdn.com/v4/letter/m/94ad74/32.png) [@michaels](https://discuss.elastic.co/u/michaels)\
**Post date:** [November 1, 2017, 9:55am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/1 "2017-11-01T09:55:55Z")

</div>

Hi

I am performing a massive bulk indexing daily and it seems that whatever I do my cluster nodes quickly use 100% CPU for hours which causes my searches to perform bad.

My configuration is as follows:

1. AWS ES cluster with 5 m4.2xlarge.elasticsearch instances, 500GB SSD, 3000 provisioned IOPS.
2. My index has 5 shards and is approx. ~200GB
3. My document type contains in addition to basic field, an array of nested document mapping.
4. Each day I need to index ~11000000 nested documents into ~700000 documents
5. Indexing is done daily using Bulk update requests, each with 3000 items which are executed over next 20 hours (request every ~6 minutes)
6. Each update request contains a script for adding new items into an existing array.

Is there anything that can be done to reduce CPU load?  
Is my indexing process correct?

thx  
Michael

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 1, 2017, 10:28am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/2 "2017-11-01T10:28:42Z")

</div>

It seems like you are updating the same documents repeatedly during this bulk indexing run. Each update of a nested document will result in multiple documents being indexed behind the scenes, which will result in a lot of merging activity. I suspect you would be better performance if you could 'aggregate' the updates per document prior to indexing so that you end up performing a single larger update per document instead of multiple small ones.

---

<div class="post-metadata">

**Author:** ![michaels](https://avatars.discourse-cdn.com/v4/letter/m/94ad74/32.png) [@michaels](https://discuss.elastic.co/u/michaels)\
**Post date:** [November 1, 2017, 10:33am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/3 "2017-11-01T10:33:50Z")

</div>

Thx for your reply

This is what I do. Each of my update requests refers to a single unique document. in total I have ~700000 requests where each contains a script to add new nested documents to an existing array.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 1, 2017, 10:44am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/4 "2017-11-01T10:44:02Z")

</div>

How large are the documents? How many nested objects/levels? How complex are the scripts performing the update? Which version of Elasticsearch are you using?

---

<div class="post-metadata">

**Author:** ![michaels](https://avatars.discourse-cdn.com/v4/letter/m/94ad74/32.png) [@michaels](https://discuss.elastic.co/u/michaels)\
**Post date:** [November 1, 2017, 10:48am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/5 "2017-11-01T10:48:01Z")

</div>

1. 1 nesting level
2. Nested documents are fairly small, they contain up to 24 bytes string field and a long field

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 1, 2017, 10:49am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/6 "2017-11-01T10:49:19Z")

</div>

Which version of Elasticsearch are you using? How many nested documents do you have on average per main document?

---

<div class="post-metadata">

**Author:** ![michaels](https://avatars.discourse-cdn.com/v4/letter/m/94ad74/32.png) [@michaels](https://discuss.elastic.co/u/michaels)\
**Post date:** [November 1, 2017, 10:53am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/7 "2017-11-01T10:53:14Z")

</div>

1. Version 5.3 (as provided by AWS)
2. Since I have ~11000000 nested items to add, on average I add ~16 items to each document daily. The total length of the nested document array can grow quit large for each document.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 1, 2017, 11:00am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/8 "2017-11-01T11:00:06Z")

</div>

What is the output from the [cat indices API](https://www.elastic.co/guide/en/elasticsearch/reference/5.6/cat-indices.html) for this index?

---

<div class="post-metadata">

**Author:** ![michaels](https://avatars.discourse-cdn.com/v4/letter/m/94ad74/32.png) [@michaels](https://discuss.elastic.co/u/michaels)\
**Post date:** [November 1, 2017, 11:02am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/9 "2017-11-01T11:02:17Z")

</div>

health status index uuid pri rep docs.count docs.deleted store.size pri.store.size  
green open myindex cVraQ7tpS86GthbXIbzqJQ 5 1 1739504457 1886305592 410.5gb 206.4gb

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 1, 2017, 11:06am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/10 "2017-11-01T11:06:32Z")

</div>

It looks like each document have an average of over 2.4k nested objects. Each update of a document will therefore result in that many documents being updated behind the scenes. Have you considered switching to a denormlized, flat data model or perhaps a parent-child relationship as these would result in inserts rather than very expensive updates?

---

<div class="post-metadata">

**Author:** ![michaels](https://avatars.discourse-cdn.com/v4/letter/m/94ad74/32.png) [@michaels](https://discuss.elastic.co/u/michaels)\
**Post date:** [November 1, 2017, 11:20am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/11 "2017-11-01T11:20:21Z")

</div>

I was not aware that adding new nested objects results in an expensive update of all old nested objects.  
The nested documents are time based and I need to perform aggregation queries on the main document given a specific time frame and values of nested documents.  
I will consider moving to parent-child relationship if I can achieve the same.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 1, 2017, 11:30am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/12 "2017-11-01T11:30:09Z")

</div>

This behaviour of nested documents is described in [Elasticsearch: the definitive guide](https://www.elastic.co/guide/en/elasticsearch/guide/2.x/nested-aggregation.html#_when_to_use_nested_objects). This section on data modelling is very useful, and even though it still references ES 2.x, most of it is as far as I know still valid.

---

<div class="post-metadata">

**Author:** ![michaels](https://avatars.discourse-cdn.com/v4/letter/m/94ad74/32.png) [@michaels](https://discuss.elastic.co/u/michaels)\
**Post date:** [November 1, 2017, 11:34am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/13 "2017-11-01T11:34:43Z")

</div>

Did not see that, definitely parent-child is more suitable for my use-case.  
Thx for your support.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 29, 2017, 11:34am UTC](https://discuss.elastic.co/t/bulk-indexing-performance-on-aws-es-service/106011/14 "2017-11-29T11:34:53Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
