# Elasticsearch performance tuning doubts

**URL:** <https://discuss.elastic.co/t/elasticsearch-performance-tuning-doubts/182299>\
**Category:** Elasticsearch\
**Created:** [May 22, 2019, 6:00pm UTC](https://discuss.elastic.co/t/elasticsearch-performance-tuning-doubts/182299 "2019-05-22T18:00:53Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![elk11](https://avatars.discourse-cdn.com/v4/letter/e/45deac/32.png) [@elk11](https://discuss.elastic.co/u/elk11)\
**Post date:** [May 22, 2019, 6:00pm UTC](https://discuss.elastic.co/t/elasticsearch-performance-tuning-doubts/182299/1 "2019-05-22T18:00:53Z")

</div>

Hi,

I have a 5 node cluster running elasticsearch 7.0.0.

Indices: 700  
Total size: 350GB (Primary storage excluding replicas)

Some of the biggest indices are around 2.5-3GB in size.

Now i have 1 primary shard and 1 replica for each indices. And on kibana it takes lot of time to load those big indices.

Does increasing the number of primary shards improve the search performance? (from kibana discover).

One big shard vs Many small shards, whose performance is better? (I'm talking about search performance here)

Thanks.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 22, 2019, 6:23pm UTC](https://discuss.elastic.co/t/elasticsearch-performance-tuning-doubts/182299/2 "2019-05-22T18:23:32Z")

</div>

> [@elk11](#):
>
> Does increasing the number of primary shards improve the search performance? (from kibana discover).

Yes.

> [@elk11](#):
>
> One big shard vs Many small shards, whose performance is better? (I'm talking about search performance here)

May I suggest you look at the following resources about sizing:

> **[Quantitative Cluster Sizing](https://www.elastic.co/elasticon/conf/2016/sf/quantitative-cluster-sizing)**

> **[How many shards should I have in my Elasticsearch cluster?](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster)**

> **[Managing your black Friday logs - CloudConf.IT](https://www.slideshare.net/dadoonet/managing-your-black-friday-logs-cloudconfit-93309899)**
>
> Managing your black Friday logs - CloudConf.IT - Download as a PDF or view online for free

[![](https://us1.discourse-cdn.com/elastic/original/3X/7/c/7c2edebd5bac194c58a5df87ede6872cb41d61a8.jpeg "Managing your Black Friday Logs, Pablo Musa, Elastic, TechSummit Amsterdam") ](https://www.youtube.com/watch?v=ilP7tG6tabI)

And [https://www.elastic.co/webinars/using-rally-to-get-your-elasticsearch-cluster-size-right](https://www.elastic.co/webinars/using-rally-to-get-your-elasticsearch-cluster-size-right)

> I have a 5 node cluster running elasticsearch 7.0.0.  
> Indices: 700  
> Total size: 350GB (Primary storage excluding replicas)

After watching the videos, you will probably understand that in general you can have around 50gb per shard (it depends as usual) and at most 20 shards per gb of RAM.

Here you have 1400 shards on 5 nodes. 280 shards per node. Which means that you need at least 14gb of HEAP  
But with 50gb per shard, most likely only 7 primaries are needed, so 14 shards including replicas. Which is around 3 shards per node.

That would require probably less memory. I'd say that 8 gb of HEAP could be enough.

Again, it depends. So you need to test that against your own scenarii. Look at the resources I linked to.

---

<div class="post-metadata">

**Author:** ![elk11](https://avatars.discourse-cdn.com/v4/letter/e/45deac/32.png) [@elk11](https://discuss.elastic.co/u/elk11)\
**Post date:** [May 22, 2019, 6:26pm UTC](https://discuss.elastic.co/t/elasticsearch-performance-tuning-doubts/182299/3 "2019-05-22T18:26:25Z")

</div>

Wonderful! Thank you so much for the detailed info @dadoonet

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 22, 2019, 9:16pm UTC](https://discuss.elastic.co/t/elasticsearch-performance-tuning-doubts/182299/4 "2019-05-22T21:16:23Z")

</div>

I would like to add a few clarifications. As described [in this webinar](https://www.elastic.co/webinars/optimizing-storage-efficiency-in-elasticsearch) the amount of heap used will depend on how you have indexed and mapped your data as well as how large and optimized your shards are. Larger shards often use less heap per document than smaller ones, which is why using large shards are generally recommended for efficiency.

The rule-of-thumb of 20 shards per GB of heap comes from users often having far too many small shards and once you reach this limit the system is generally still working well. This recommendation is however a maximum number of shards and not a level you necessarily should expect to be able to reach. If you are using large shards I would expect the number of shards per node to be lower than the prescribed limit.

I sometimes hear the recommendation interpreted as "I should be able to have 20 50GB shards per GB of heap" which is not correct. If this was the case a node could have 1TB of data per GB of heap, which is generally very hard to achieve, at least without using frozen indices or extensive optimizations.

Each query is executed single-threaded against each shard, although multiple shard queries are run in parallel. The optimal number of shards for query performance therefore depends on the number of concurrent queries that compete for resources as well as whether it is CPU, heap or disk I/O that is limiting performance. Having multiple shards can often be faster than having a single one, but if you have too many shards performance is likely top start deteriorating. Best way to find out is to benchmark with as realistic data and queries as possible.

---

<div class="post-metadata">

**Author:** ![elk11](https://avatars.discourse-cdn.com/v4/letter/e/45deac/32.png) [@elk11](https://discuss.elastic.co/u/elk11)\
**Post date:** [May 25, 2019, 6:15pm UTC](https://discuss.elastic.co/t/elasticsearch-performance-tuning-doubts/182299/5 "2019-05-25T18:15:15Z")

</div>

@dadoonet @Christian_Dahlqvist I tried to implement some of the suggestions for performance but I haven't noticed a considerable performance improvements. Can you please tell me if am doing something wrong ? The new configurations as the compared to the old ones described above are:

Indices: 200 (525 primary shards, 525 replica shards)

Primary store size = 140GB (decreased from 220 GB after reidexing from 5.x cluster)

I have 5 nodes each with 16GB of JVM heap configured.

So, around 1000 shards (including replicas) in total cluster. ---\> 200 shards/Node ----\> 12 shards/GB of heap.

So I was expecting a very fast query results. But some queries take around 40-50 sec to complete.

I know i haven't considered docs count. But my each indices have only around 50k-60k docs.

I even ran 'forcemerge' on my old indices. That didn't help too.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 25, 2019, 8:32pm UTC](https://discuss.elastic.co/t/elasticsearch-performance-tuning-doubts/182299/6 "2019-05-25T20:32:28Z")

</div>

That still sounds like far too many shards given the size of the data. If you had s total of 10 indices with a single primary shard each the average shard size would be 14GB which is quite reasonable.

---

<div class="post-metadata">

**Author:** ![elk11](https://avatars.discourse-cdn.com/v4/letter/e/45deac/32.png) [@elk11](https://discuss.elastic.co/u/elk11)\
**Post date:** [May 27, 2019, 12:47pm UTC](https://discuss.elastic.co/t/elasticsearch-performance-tuning-doubts/182299/7 "2019-05-27T12:47:34Z")

</div>

Thanks for the suggestion @Christian_Dahlqvist

I reduced the number of primary shards to 70 (I can still reduce, but i have time based indices and i would like some granularity). But I'm not seeing any performance improvements while querying. Instead the the query time increased a bit.

Then i increased replicas from 1 to 2. The query time improved a bit. But increasing replicas is not preferred choice for me as the store size increases.

I have 8 CPU cores for each node. Which i think is enough.

i ran 'docker stats' and:

```
CONTAINER ID NAME CPU % MEM USAGE / LIMIT MEM % NET I/O BLOCK I/O PIDS
345313a89d89 kibana 0.73% 132.8MiB / 43.07GiB 0.30% 1.63GB / 210MB 210MB / 8.19kB 11
123410886856 elasticsearch 16.03% 19.02GiB / 43.07GiB 44.16% 1.44TB / 1.16TB 218GB / 1.52TB 125
345bc40c4c31 bbbbbbbbb 0.04% 102.2MiB / 43.07GiB 0.23% 2.45MB / 656B 161MB / 0B 22
5678069c72f7 aaaaaaa 0.66% 1.663GiB / 43.07GiB 3.86% 8GB / 1.51GB 586MB / 938MB 84
364067897a56 some_other_container 0.63% 315.6MiB / 43.07GiB 0.72% 102GB / 8.61GB 2.84GB / 164GB 44
456778b25f30 some_container 0.61% 481.8MiB / 43.07GiB 1.09% 307GB / 4.5GB 1.79GB / 457GB 46

```

What am i missing?

---

<div class="post-metadata">

**Author:** ![elk11](https://avatars.discourse-cdn.com/v4/letter/e/45deac/32.png) [@elk11](https://discuss.elastic.co/u/elk11)\
**Post date:** [June 2, 2019, 12:16pm UTC](https://discuss.elastic.co/t/elasticsearch-performance-tuning-doubts/182299/8 "2019-06-02T12:16:27Z")

</div>

@Christian_Dahlqvist @dadoonet, I think the problem I was having with performance was because I has some large text fields (Large enough that they violated max field length! [https://www.elastic.co/guide/en/elasticsearch/reference/current/breaking-changes-7.0.html#\_limiting\_the\_length\_of\_an\_analyzed\_text\_during\_highlighting](https://www.elastic.co/guide/en/elasticsearch/reference/current/breaking-changes-7.0.html#_limiting_the_length_of_an_analyzed_text_during_highlighting))

I solved that length issue by indexing those large fields as term\_vectors.

Now the question is, how to tune the query performance from kibana with these term\_vector fields?

Is there a way I can exclude these large text fields when it is unnecessary? (But those large text fields should still remain searchable)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 30, 2019, 12:21pm UTC](https://discuss.elastic.co/t/elasticsearch-performance-tuning-doubts/182299/9 "2019-06-30T12:21:04Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
