# INDEX Performance

**URL:** <https://discuss.elastic.co/t/index-performance/136728>\
**Category:** Elasticsearch\
**Created:** [June 20, 2018, 4:00pm UTC](https://discuss.elastic.co/t/index-performance/136728 "2018-06-20T16:00:30Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![snigdharao](https://avatars.discourse-cdn.com/v4/letter/s/4af34b/32.png) [@snigdharao](https://discuss.elastic.co/u/snigdharao)\
**Post date:** [June 20, 2018, 4:00pm UTC](https://discuss.elastic.co/t/index-performance/136728/1 "2018-06-20T16:00:31Z")

</div>

We have been trying to move from solr to elastic search and want to compare the performance for indexing 100M records from database.  
Currently it takes 4 hours to index 52 Million records

Current confiuguration:  
default shards = 5, and also we have increased refresh time interval to 30s.  
Whats the best way to increase the performance , I am planning to increase shards to 7.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 20, 2018, 4:51pm UTC](https://discuss.elastic.co/t/index-performance/136728/2 "2018-06-20T16:51:00Z")

</div>

Increasing the number of shards will not necessarily improve performance. It could actually do the opposite.

Have you gone through [these guidelines](https://www.elastic.co/guide/en/elasticsearch/reference/6.3/tune-for-indexing-speed.html)? How many indexing threads are you using? What bulk size are you using?

---

<div class="post-metadata">

**Author:** ![snigdharao](https://avatars.discourse-cdn.com/v4/letter/s/4af34b/32.png) [@snigdharao](https://discuss.elastic.co/u/snigdharao)\
**Post date:** [June 20, 2018, 5:33pm UTC](https://discuss.elastic.co/t/index-performance/136728/3 "2018-06-20T17:33:59Z")

</div>

@Christian_Dahlqvist  
We are currently used 9 logstash threads to index the data, and i have set refresh interval to 30s .

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 20, 2018, 6:14pm UTC](https://discuss.elastic.co/t/index-performance/136728/4 "2018-06-20T18:14:33Z")

</div>

What is the specification of your Elasticsearch cluster? Which version are you running? What is the average size of your documents?

---

<div class="post-metadata">

**Author:** ![snigdharao](https://avatars.discourse-cdn.com/v4/letter/s/4af34b/32.png) [@snigdharao](https://discuss.elastic.co/u/snigdharao)\
**Post date:** [June 20, 2018, 6:31pm UTC](https://discuss.elastic.co/t/index-performance/136728/5 "2018-06-20T18:31:18Z")

</div>

the size of the VM is 4 cores , 24GB RAM .Cluster has one node by default 5 shards.  
Added bootstrap.memory\_lock: true by referring to the guide.

Version is 6.2.2  
We are basically trying to index 50 M records from database with 32 columns.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 20, 2018, 6:39pm UTC](https://discuss.elastic.co/t/index-performance/136728/7 "2018-06-20T18:39:15Z")

</div>

What type of storage do you have? What does disk I/O and iowait look like while you are indexing?

---

<div class="post-metadata">

**Author:** ![snigdharao](https://avatars.discourse-cdn.com/v4/letter/s/4af34b/32.png) [@snigdharao](https://discuss.elastic.co/u/snigdharao)\
**Post date:** [June 20, 2018, 6:58pm UTC](https://discuss.elastic.co/t/index-performance/136728/8 "2018-06-20T18:58:46Z")

</div>

Attached my indexing stats , please help me understand what is wrong  
 ![20%20PM](https://us1.discourse-cdn.com/elastic/original/3X/4/3/4321404b79e3d6f53cf9f103865a448f676ebfba.png) ![07%20PM](https://us1.discourse-cdn.com/elastic/original/3X/f/1/f19681eaf2df9d0edfa25abbf760dd1bd4c1d414.png) ![48%20PM](https://us1.discourse-cdn.com/elastic/original/3X/f/9/f94d3508388cab549d9d56c6bbe331d7c9934317.png)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 20, 2018, 7:03pm UTC](https://discuss.elastic.co/t/index-performance/136728/9 "2018-06-20T19:03:25Z")

</div>

It looks like you have quite a few deleted documents. Do a large portion of the documents you load result in an update?

---

<div class="post-metadata">

**Author:** ![snigdharao](https://avatars.discourse-cdn.com/v4/letter/s/4af34b/32.png) [@snigdharao](https://discuss.elastic.co/u/snigdharao)\
**Post date:** [June 20, 2018, 7:08pm UTC](https://discuss.elastic.co/t/index-performance/136728/10 "2018-06-20T19:08:21Z")

</div>

So we have changed configuration in logstash to remove duplicate documents but have mapped it by unique column and changed it from default \_id to unique id from our database.  
I dont know why we are seeing documents deleted

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 20, 2018, 7:16pm UTC](https://discuss.elastic.co/t/index-performance/136728/11 "2018-06-20T19:16:35Z")

</div>

An update would show up as a delete and an insert.

---

<div class="post-metadata">

**Author:** ![snigdharao](https://avatars.discourse-cdn.com/v4/letter/s/4af34b/32.png) [@snigdharao](https://discuss.elastic.co/u/snigdharao)\
**Post date:** [June 20, 2018, 7:18pm UTC](https://discuss.elastic.co/t/index-performance/136728/12 "2018-06-20T19:18:24Z")

</div>

We are not updating the documents and instead changed the logstash confirguration to include unique id has document id .Attaching my configuration , can you please let me if i need to change something ? ![34%20PM](https://us1.discourse-cdn.com/elastic/original/3X/e/c/ec56b14c03549cb20aeb662ea1cf0c89ad7084ff.png)

---

<div class="post-metadata">

**Author:** ![snigdharao](https://avatars.discourse-cdn.com/v4/letter/s/4af34b/32.png) [@snigdharao](https://discuss.elastic.co/u/snigdharao)\
**Post date:** [June 20, 2018, 8:17pm UTC](https://discuss.elastic.co/t/index-performance/136728/13 "2018-06-20T20:17:53Z")

</div>

![02%20PM](https://us1.discourse-cdn.com/elastic/original/3X/a/6/a67b2c53c633cb923be8bf58ce4430dcdfa8b7f8.png)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 20, 2018, 8:22pm UTC](https://discuss.elastic.co/t/index-performance/136728/14 "2018-06-20T20:22:14Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> What type of storage do you have? What does disk I/O and iowait look like while you are indexing?

Do you have any details about the underlying storage?

---

<div class="post-metadata">

**Author:** ![snigdharao](https://avatars.discourse-cdn.com/v4/letter/s/4af34b/32.png) [@snigdharao](https://discuss.elastic.co/u/snigdharao)\
**Post date:** [June 20, 2018, 9:05pm UTC](https://discuss.elastic.co/t/index-performance/136728/15 "2018-06-20T21:05:36Z")

</div>

No , i dont have any details about that.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 21, 2018, 5:18am UTC](https://discuss.elastic.co/t/index-performance/136728/16 "2018-06-21T05:18:47Z")

</div>

The type and performance of storage is often a limiting factor for Elasticsearch. Check what ‘iostat -x’ gives while indexing.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 19, 2018, 5:18am UTC](https://discuss.elastic.co/t/index-performance/136728/17 "2018-07-19T05:18:49Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
