# How to improve performance?

**URL:** <https://discuss.elastic.co/t/how-to-improve-performance/129685>\
**Category:** Elasticsearch\
**Created:** [April 26, 2018, 1:36pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685 "2018-04-26T13:36:40Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![flochon](https://avatars.discourse-cdn.com/v4/letter/f/e19b73/32.png) [@flochon](https://discuss.elastic.co/u/flochon)\
**Post date:** [April 26, 2018, 1:36pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/1 "2018-04-26T13:36:41Z")

</div>

Hello,

I use Elasticsearch last version and my cluster contains 3 nodes. I receive 100GB of data per day.  
Actually, I have 6GB of heap size.

My Elasticsearch has a lot of latence when I search data. I would like to know the best config to improve Elasticsearch.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 26, 2018, 1:38pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/2 "2018-04-26T13:38:57Z")

</div>

What is the specification of the hosts your cluster is deployed on? How much data do you have in the cluster? How many indices and shards is this data distributed across?

---

<div class="post-metadata">

**Author:** ![flochon](https://avatars.discourse-cdn.com/v4/letter/f/e19b73/32.png) [@flochon](https://discuss.elastic.co/u/flochon)\
**Post date:** [April 26, 2018, 1:45pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/3 "2018-04-26T13:45:26Z")

</div>

Each host has 4GB of RAM with 2GB for the heap size, 120GB of hard disk, 4 VCPU and Ubuntu 16,04.  
Actually, I have 172000000 documents and I have an index per day.  
The configuration of my index is 5 shards and 1 replica.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 26, 2018, 1:52pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/4 "2018-04-26T13:52:31Z")

</div>

> [@flochon](#):
>
> The configuration of my index is 5 shards and 1 replica.

That sounds excessive.

How many indices do you have in the cluster?

---

<div class="post-metadata">

**Author:** ![flochon](https://avatars.discourse-cdn.com/v4/letter/f/e19b73/32.png) [@flochon](https://discuss.elastic.co/u/flochon)\
**Post date:** [April 26, 2018, 1:55pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/5 "2018-04-26T13:55:57Z")

</div>

Actually, I have 43 indices.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 26, 2018, 1:59pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/6 "2018-04-26T13:59:22Z")

</div>

That means that you have 430 shards? That is a lot given the amount of data you have. You should probably look to reduce this significantly, e.g. through the [shrink index API](https://www.elastic.co/guide/en/elasticsearch/reference/6.2/indices-shrink-index.html). Also have a look ate [this blog post](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster) for some guidance on sharding.

---

<div class="post-metadata">

**Author:** ![flochon](https://avatars.discourse-cdn.com/v4/letter/f/e19b73/32.png) [@flochon](https://discuss.elastic.co/u/flochon)\
**Post date:** [April 26, 2018, 2:02pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/7 "2018-04-26T14:02:50Z")

</div>

I have 127 shards. In fact, I let the number of shards by default in the configuration.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 26, 2018, 2:04pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/8 "2018-04-26T14:04:44Z")

</div>

Have you identified what is limiting performance? What does CPU usage, memory usage, disk I/O (and iowait) look like?

---

<div class="post-metadata">

**Author:** ![flochon](https://avatars.discourse-cdn.com/v4/letter/f/e19b73/32.png) [@flochon](https://discuss.elastic.co/u/flochon)\
**Post date:** [April 26, 2018, 2:10pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/9 "2018-04-26T14:10:29Z")

</div>

My CPU usage is at 25% so I think isn't the problem.  
My memory usage is at 91%, it's problematic and I think improving it is a solution.  
The heap size is at 1GB on each node.  
My disk I/O is between 15 and 20 MB/s in the max.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 26, 2018, 2:30pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/10 "2018-04-26T14:30:08Z")

</div>

Do you see messages about long or frequent GC in the Elasticsearch logs?

---

<div class="post-metadata">

**Author:** ![flochon](https://avatars.discourse-cdn.com/v4/letter/f/e19b73/32.png) [@flochon](https://discuss.elastic.co/u/flochon)\
**Post date:** [April 26, 2018, 2:34pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/11 "2018-04-26T14:34:38Z")

</div>

No,I have not logs about that and the last log goes back to 16/04.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 26, 2018, 2:36pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/12 "2018-04-26T14:36:42Z")

</div>

> [@flochon](#):
>
> The heap size is at 1GB on each node.

I thought you said the heap size was 2GB per node.

> [@flochon](#):
>
> My memory usage is at 91%, it's problematic and I think improving it is a solution.

If this is constant you probably need to increase the size of the heap.

> [@flochon](#):
>
> My disk I/O is between 15 and 20 MB/s in the max.

How much `iowait` do you see?

---

<div class="post-metadata">

**Author:** ![flochon](https://avatars.discourse-cdn.com/v4/letter/f/e19b73/32.png) [@flochon](https://discuss.elastic.co/u/flochon)\
**Post date:** [April 26, 2018, 2:49pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/13 "2018-04-26T14:49:22Z")

</div>

Yes, you are right. It's 2GB per node and 1GB is used.  
Yes it's constant in the time.  
IOWAIT changes a lot, I see 45% as max value.

---

<div class="post-metadata">

**Author:** ![flochon](https://avatars.discourse-cdn.com/v4/letter/f/e19b73/32.png) [@flochon](https://discuss.elastic.co/u/flochon)\
**Post date:** [April 30, 2018, 1:32pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/14 "2018-04-30T13:32:41Z")

</div>

I create one index by day and on each index I have 47 millions of logs.  
What's the best number of shard to run this ?

Actually, I have 5 shards and 3 nodes.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 30, 2018, 1:40pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/15 "2018-04-30T13:40:56Z")

</div>

> [@flochon](#):
>
> What's the best number of shard to run this ?

That depends on the size. have a look at the blog post I linked to earlier for some guidance.

---

<div class="post-metadata">

**Author:** ![flochon](https://avatars.discourse-cdn.com/v4/letter/f/e19b73/32.png) [@flochon](https://discuss.elastic.co/u/flochon)\
**Post date:** [May 3, 2018, 12:11pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/16 "2018-05-03T12:11:02Z")

</div>

Is it possible with 275 millions of documents the search is long or with a good optimisation the search will be improved ?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 3, 2018, 12:28pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/17 "2018-05-03T12:28:53Z")

</div>

What type of data do you have? How are you modelling your data? What type of queries are you running?

---

<div class="post-metadata">

**Author:** ![flochon](https://avatars.discourse-cdn.com/v4/letter/f/e19b73/32.png) [@flochon](https://discuss.elastic.co/u/flochon)\
**Post date:** [May 3, 2018, 12:39pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/18 "2018-05-03T12:39:25Z")

</div>

I have logs.

All my logs pass in the filter of Logstash but I don't know if it's that you hear in "modelling"

It's in the discover, I change the filter of date and I take a more large scale and the search is very long.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 3, 2018, 12:50pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/19 "2018-05-03T12:50:14Z")

</div>

Make sure that you do not have a lot of small shards, as that can be inefficient and cause performance problems. Have a look at [this blog post](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster) (linked to it earlier) for some practical guidelines.

If you can provide the full output of the [cluster stats API](https://www.elastic.co/guide/en/elasticsearch/reference/current/cluster-stats.html), we will get a better view of the state of your cluster.

How many visualisations do you have in the dashboards that are slow? Are they slow for shorter time periods as well?

---

<div class="post-metadata">

**Author:** ![flochon](https://avatars.discourse-cdn.com/v4/letter/f/e19b73/32.png) [@flochon](https://discuss.elastic.co/u/flochon)\
**Post date:** [May 3, 2018, 12:52pm UTC](https://discuss.elastic.co/t/how-to-improve-performance/129685/20 "2018-05-03T12:52:31Z")

</div>

I have 6 visualisations and no if I take a less period with less logs the time is good.

[Next page](https://discuss.elastic.co/t/how-to-improve-performance/129685.md?page=2)
