# Elasticsearch Cluster Performance Tuning Help required

**URL:** <https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867>\
**Category:** Elasticsearch\
**Created:** [November 22, 2018, 11:32am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867 "2018-11-22T11:32:29Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![nuwancs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nuwancs/32/46419_2.png) [@nuwancs](https://discuss.elastic.co/u/nuwancs)\
**Post date:** [November 22, 2018, 11:32am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/1 "2018-11-22T11:32:29Z")

</div>

Hi All,  
I have 3 nodes elastic cluster each node assigned with 1Tb hard disk, 15GB Ram and 4 CPU  
Elastic version using is 6.3.0.  
At the moment I have 3,304,624,313 documents using 2.1 TB disk space  
This data is collected within a month.

Problems is  
Doing a search on cluster takes over 5 minutes.

1. In order to optimize search performance, what can I do?

2. what is the maximum data size 3 node cluster can handle?

3. Is it ok to split the indices vertically so small fields are grouped in one index? will it help improving performance

---

<div class="post-metadata">

**Author:** ![mjunaidmuzammil](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mjunaidmuzammil/32/56910_2.png) [@mjunaidmuzammil](https://discuss.elastic.co/u/mjunaidmuzammil)\
**Post date:** [November 22, 2018, 12:03pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/2 "2018-11-22T12:03:20Z")

</div>

For answering 1) & 2), need answer to how many indexes & shards your cluster contain? What is the total heap memory allocated for ES?

I am not sure whether I understand your question 3). Can you explain that with an example?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 22, 2018, 12:09pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/3 "2018-11-22T12:09:35Z")

</div>

What is the output of the [cluster health API](https://www.elastic.co/guide/en/elasticsearch/reference/6.5/cluster-health.html)?

---

<div class="post-metadata">

**Author:** ![nuwancs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nuwancs/32/46419_2.png) [@nuwancs](https://discuss.elastic.co/u/nuwancs)\
**Post date:** [November 22, 2018, 12:12pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/4 "2018-11-22T12:12:09Z")

</div>

Thank you Christian for quick response.  
Please find below cluster health status  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/a/3/a3062c3b649e3f94bc7a426e879f75d40b3d5ad7.png)

---

<div class="post-metadata">

**Author:** ![nuwancs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nuwancs/32/46419_2.png) [@nuwancs](https://discuss.elastic.co/u/nuwancs)\
**Post date:** [November 22, 2018, 12:15pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/5 "2018-11-22T12:15:17Z")

</div>

Thanks Junaid for quick response

please find below requested information

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/d/1/d1f0376719864f787e544d50f2f7563cb22ad607.png)

Regarding my 3 rd questiom, think my document has field1 and field2, where field2 is a long text. So is it ok to split the index into two indices index1 will contain filed1 only and index2 will contain field2 only. will it help improving my search performance as i will do the search on less number of fields?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 22, 2018, 12:19pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/6 "2018-11-22T12:19:11Z")

</div>

That gives an average shard size of just over 2GB, which is a bit on the small side.

Have you looked at monitoring to see what is limiting performance? Is it CPU or perhaps slow storage resulting in significant iowait?

---

<div class="post-metadata">

**Author:** ![nuwancs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nuwancs/32/46419_2.png) [@nuwancs](https://discuss.elastic.co/u/nuwancs)\
**Post date:** [November 22, 2018, 12:38pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/7 "2018-11-22T12:38:01Z")

</div>

Hi Christian,  
I'm using google cloud basic hard disk, hope it can do this job well 🙂

When a search is made all cpu reaches almost 400%  
I have allocated 8GB out of 15GB ram for the elastic but it doesn't pass 60%

Regarding shard size, is it worth increasing the shard size (~40GB) by doing a reindexing?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 22, 2018, 12:43pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/8 "2018-11-22T12:43:04Z")

</div>

Elasticsearch is generally very I/O intensive, so having fast storage is very important. Run `iostat -x` to see how the storage is performing. I would not be surprised to see a lot of iowait indicating that this is the bottleneck. If that is confirmed I would recommend to upgrading ton more performant storage.

---

<div class="post-metadata">

**Author:** ![nuwancs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nuwancs/32/46419_2.png) [@nuwancs](https://discuss.elastic.co/u/nuwancs)\
**Post date:** [November 22, 2018, 12:59pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/9 "2018-11-22T12:59:08Z")

</div>

Please find the iostat -x output

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/1/f14d8bff04770be7b5bf429dc54bed2df09daa5f.png)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 22, 2018, 1:14pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/10 "2018-11-22T13:14:21Z")

</div>

That doesn't look too bad assuming it was taken while a query was running. Then you may be limited by CPU, so may need to scale out or up the cluster.

---

<div class="post-metadata">

**Author:** ![nuwancs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nuwancs/32/46419_2.png) [@nuwancs](https://discuss.elastic.co/u/nuwancs)\
**Post date:** [November 22, 2018, 1:43pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/11 "2018-11-22T13:43:23Z")

</div>

@Christian_Dahlqvist Thanks for the support. I will do a test after scaling.

Finally, I am planning have 3 times more day in the future as 3 months retention is required (the day we are looking at is one month)  
if i am to stick to the same hardware spec will the following make any performance improvement?

1. splitting the index and putting filed1 2 in one index and filed 3 and 4 for i another index. out search queries are mostly based on a single filed which has a json payoad.
2. increasing the shard size to larger value and reducing the number of shards handling as for the moment i got 888 shards

---

<div class="post-metadata">

**Author:** ![nuwancs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nuwancs/32/46419_2.png) [@nuwancs](https://discuss.elastic.co/u/nuwancs)\
**Post date:** [November 22, 2018, 1:48pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/12 "2018-11-22T13:48:34Z")

</div>

Btw hope this is you 🙂

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/c/b/cb3744dd71bd39e51e301f863866de2c312f1be1.jpeg)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 22, 2018, 2:10pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/13 "2018-11-22T14:10:08Z")

</div>

It is indeed.

---

<div class="post-metadata">

**Author:** ![jeroen1](https://avatars.discourse-cdn.com/v4/letter/j/47e85d/32.png) [@jeroen1](https://discuss.elastic.co/u/jeroen1)\
**Post date:** [November 27, 2018, 7:39pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/14 "2018-11-27T19:39:04Z")

</div>

> In order to optimize search performance, what can I do?

You've got a ~2T dataset and ~50G RAM. This means: lots of I/O (dataset does not fit in RAM). Two options for increasing performance without other changes:

- More RAM (= more data in memory and / or file system caches). RAM is way way faster than anything else. So everything coming from RAM is a big plus.
- Faster disks (e.g. SSD, SSD in RAID). If the total amount of RAM is \< 2 TB, significant disk I/O is needed. Spinning disks = ~125MB/s, single SATA SSD = ~500MB/s, SSD RAID sets of PCIe SSD = way way faster. This way everything NOT coming from RAM can still load sort of fast.

Mapping (change requires re-indexing):

- According to other posts (I do not know the reason): use a max shard size of ~50GB.
- With rule above in mind: keep as close to 1 shard per CPU core as you can (1 shard = 1 process).
- Rough estimate in your case: ~2T / 0.05 = optimal is ~40 shards (if the dataset will not grow).
- Since you've got CPU 12 cores this is not the most efficient setup. So more cores will help as well.

---

<div class="post-metadata">

**Author:** ![nuwancs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nuwancs/32/46419_2.png) [@nuwancs](https://discuss.elastic.co/u/nuwancs)\
**Post date:** [December 2, 2018, 5:28am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/15 "2018-12-02T05:28:43Z")

</div>

Thanks, @jeroen1 for the detailed answer. I will update my setup as per your instructions.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 30, 2018, 5:28am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-performance-tuning-help-required/157867/16 "2018-12-30T05:28:43Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
