# Investigate indexing bottleneck

**URL:** <https://discuss.elastic.co/t/investigate-indexing-bottleneck/86924>\
**Category:** Elasticsearch\
**Created:** [May 24, 2017, 8:54am UTC](https://discuss.elastic.co/t/investigate-indexing-bottleneck/86924 "2017-05-24T08:54:39Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![weibin.wu](https://avatars.discourse-cdn.com/v4/letter/w/4491bb/32.png) [@weibin.wu](https://discuss.elastic.co/u/weibin.wu)\
**Post date:** [May 24, 2017, 8:54am UTC](https://discuss.elastic.co/t/investigate-indexing-bottleneck/86924/1 "2017-05-24T08:54:39Z")

</div>

# We would like to index in a cluster as fast as possible, right now the cluster cap at 11000 doc/per second. All the documents we index look like this

{  
"\_index":"myteksi-changeling\_changeling\_models\_loglings",  
"\_type":"changeling/models/logling",  
"\_id":"91078043c6f9ab8816e5377cd817594b",  
"\_score":1,  
"\_source":{  
"id":"91078043c6f9ab8816e5377cd817594b",  
"klass":"candidate",  
"oid":"2681140321",  
"modified\_by":null,  
"modifications":"{"driver\_distance":[1.46391,3.842],"lock\_version":[0,1]}",  
"modified\_at":"2015-04-25T00:11:33+08:00",  
"modified\_fields":[  
"driver\_distance",  
"lock\_version"  
]  
}  
}

Our server that sends the documents is using about 50% CPU.  
Network bandwidth is not a worry because it is setup in AWS.

The CPU usage of the ES cluster is Avg 30%.  
We are using SSD. Write throughtput is about maximum 80MB /per second.

Any attribute I can look at to investigate the bottleneck. please advice thanks.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 24, 2017, 9:04am UTC](https://discuss.elastic.co/t/investigate-indexing-bottleneck/86924/2 "2017-05-24T09:04:46Z")

</div>

It would be useful if you could provide some additional details:

- Which version of Elasticsearch are you using?
- Which EC2 instance types are you using? How large is the cluster?
- How many indices/shards are you actively indexing into?
- What is the size of indices and shards being indexed into?
- Are indexed documents immutable or updated? If updated, how large portion of operations are updates?
- Are your mappings static or dynamic?
- Are you indexing in bulk? If so, what is your bulk size?
- How many parallel indexing threads do you have against the cluster?

---

<div class="post-metadata">

**Author:** ![weibin.wu](https://avatars.discourse-cdn.com/v4/letter/w/4491bb/32.png) [@weibin.wu](https://discuss.elastic.co/u/weibin.wu)\
**Post date:** [May 24, 2017, 10:59am UTC](https://discuss.elastic.co/t/investigate-indexing-bottleneck/86924/3 "2017-05-24T10:59:25Z")

</div>

Which version of Elasticsearch are you using?  
5.1

Which EC2 instance types are you using? How large is the cluster?  
3 x m4.xlarge

How many indices/shards are you actively indexing into?  
indices are separated by month.  
Each month index is about 140GB max with 5 shards.

What is the size of indices and shards being indexed into?  
140GB max with 5 shards

Are indexed documents immutable or updated? If updated, how large portion of operations are updates?  
The indexed document is totally new for the cluster. We insert with \_bulk of 20000 in a batch.

Are your mappings static or dynamic?  
mapping is static ,here is the mapping

{  
"aliases": {

```
},
"mappings": {
  "changeling/models/logling": {
    "properties": {
      "id": {
        "type": "string"
      },
      "klass": {
        "type": "string"
      },
      "modifications": {
        "type": "string"
      },
      "modified_at": {
        "type": "date",
        "format": "dateOptionalTime"
      },
      "modified_by": {
        "type": "string"
      },
      "modified_fields": {
        "type": "string",
        "analyzer": "keyword"
      },
      "oid": {
        "type": "string"
      }
    }
  }
},
"settings": {
  "index": {
    "number_of_replicas": "1",
    "number_of_shards": "5",
    "refresh_interval": "60s"
  }
},
"warmers": {
  
}

```

}

Are you indexing in bulk? If so, what is your bulk size?  
yes. 20000 in a batch  
How many parallel indexing threads do you have against the cluster?  
I used 3 processes in Linux, and allocate them to different CPU.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 24, 2017, 1:43pm UTC](https://discuss.elastic.co/t/investigate-indexing-bottleneck/86924/4 "2017-05-24T13:43:14Z")

</div>

Do you have monitoring installed? What does GC look like? How many queries per second are you seeing?

---

<div class="post-metadata">

**Author:** ![weibin.wu](https://avatars.discourse-cdn.com/v4/letter/w/4491bb/32.png) [@weibin.wu](https://discuss.elastic.co/u/weibin.wu)\
**Post date:** [May 25, 2017, 3:23am UTC](https://discuss.elastic.co/t/investigate-indexing-bottleneck/86924/5 "2017-05-25T03:23:52Z")

</div>

query per second as I mentioned, it was 10000 indexing per second. No query at that time.

Please see monitoring here

 ![](https://us1.discourse-cdn.com/elastic/original/3X/f/2/f2c7b3c5e1829a91e10960350b24053fd9ff55be.png)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 25, 2017, 6:53am UTC](https://discuss.elastic.co/t/investigate-indexing-bottleneck/86924/6 "2017-05-25T06:53:19Z")

</div>

Based on the sample record it looks like you are specifying the document ID at the application layer instead of letting Elasticsearch assign one. Is this correct? The way you assign a document id [can have an impact on indexing performance](http://blog.mikemccandless.com/2014/05/choosing-fast-unique-identifier-uuid.html) as Elasticsearch need to determine if it is an update or an new document. Are you by any chance seeing indexing throughput drop as the monthly index gets larger and then recover once you start a. new monthly index? If you are not updating documents, can you let Elasticsearch assign IDs and see if that makes a difference?

---

<div class="post-metadata">

**Author:** ![weibin.wu](https://avatars.discourse-cdn.com/v4/letter/w/4491bb/32.png) [@weibin.wu](https://discuss.elastic.co/u/weibin.wu)\
**Post date:** [May 25, 2017, 6:58am UTC](https://discuss.elastic.co/t/investigate-indexing-bottleneck/86924/7 "2017-05-25T06:58:51Z")

</div>

I am using \_bulk request to do the index like this  
{"index":{"\_index":"myteksi-changeling\_changeling\_models\_loglings","\_type":"changeling/models/logling","\_id":"4bdfa1cdab20cb352cf745db1fbc7cfd"}}  
{"id":"4bdfa1cdab20cb352cf745db1fbc7cfd","klass":"rating","oid":"27647823","modified\_by":null,"modifications":"{"current\_rating":["0.0","4.58765"]}","modified\_at":"2016-02-14T03:12:17+08:00","modified\_fields":["current\_rating"]}

If I want to let ES decide the \_id, should I remove \_id field like this?  
{"index":{"\_index":"myteksi-changeling\_changeling\_models\_loglings","\_type":"changeling/models/logling"}  
{"klass":"rating","oid":"27647823","modified\_by":null,"modifications":"{"current\_rating":["0.0","4.58765"]}","modified\_at":"2016-02-14T03:12:17+08:00","modified\_fields":["current\_rating"]}

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 25, 2017, 7:04am UTC](https://discuss.elastic.co/t/investigate-indexing-bottleneck/86924/8 "2017-05-25T07:04:56Z")

</div>

If you do not specify an id Elasticsearch will assign one, so that looks correct.

---

<div class="post-metadata">

**Author:** ![weibin.wu](https://avatars.discourse-cdn.com/v4/letter/w/4491bb/32.png) [@weibin.wu](https://discuss.elastic.co/u/weibin.wu)\
**Post date:** [May 25, 2017, 7:06am UTC](https://discuss.elastic.co/t/investigate-indexing-bottleneck/86924/9 "2017-05-25T07:06:12Z")

</div>

Okay thanks. I will have a test to see whether its faster.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 22, 2017, 7:06am UTC](https://discuss.elastic.co/t/investigate-indexing-bottleneck/86924/10 "2017-06-22T07:06:13Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
