# ES 6.1.2 Cluster shows performance bottleneck

**URL:** <https://discuss.elastic.co/t/es-6-1-2-cluster-shows-performance-bottleneck/153208>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [October 19, 2018, 5:47pm UTC](https://discuss.elastic.co/t/es-6-1-2-cluster-shows-performance-bottleneck/153208 "2018-10-19T17:47:18Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![lmuthuraman](https://avatars.discourse-cdn.com/v4/letter/l/439d5e/32.png) [@lmuthuraman](https://discuss.elastic.co/u/lmuthuraman)\
**Post date:** [October 19, 2018, 5:47pm UTC](https://discuss.elastic.co/t/es-6-1-2-cluster-shows-performance-bottleneck/153208/1 "2018-10-19T17:47:19Z")

</div>

We were seeing some real slowness in the queries in our new ES 6.1.2 cluster. In order to understand more, we started running Rally Geoname test bed to get some performance benchmark of the cluster

The test results shows the querying is pretty slow.

Here is our test configuration that is run on AWS

> |Elastic Search Version|6.1.2|  
> Master Node = 3 (r4.large)  
> Data Node = 3 (i3.2xlarge)  
> |Front Node = 1 (r4.2xlarge) == rally host  
> Rally Test Car - Geonames

> java version "1.8.0\_161"

Summary/Our Analysis of the data

1. We ran the same rally tests on the 2.4 ES version legacy 3 data node cluster using the same hardware in AWS, Our 2.x rally tests performed way better than 6.x cluster
2. We also ran the same rally tests on 1 data node 6.1.2 cluster. We found the performance numbers in 1 data node 6.1.2 cluster is way better than 3 data node 6.1.2 cluster, but poorer than 2.4 cluster
3. For the country\_aggregrate uncached numbers, there is a huge difference between latency and service time. We checked the CPU utilization and all the system metrics. CPU utilization is hovering only around 40-50%.
4. We have run these tests multiple time over the last 1 week and found the results from rally are consistent.
5. We are running the basic configuration with nothing much changed in the Elasticsearch configuration for the 6.1.2 cluster.

I am attaching the part of the test results.

 ![RallyTestResults%20](https://us1.discourse-cdn.com/elastic/original/3X/9/f/9f87347945936dae09a56358ca82eaa8654f31a5.jpeg)

Any pointers or thoughts on what might be going on in our cluster. We are migrating our users from 2.4 to 6.1.2 and want to get a good handle before we roll everyone to new cluster and shut down the old cluster.

---

<div class="post-metadata">

**Author:** ![dliappis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dliappis/32/56174_2.png) [@dliappis](https://discuss.elastic.co/u/dliappis)\
**Post date:** [October 22, 2018, 9:09am UTC](https://discuss.elastic.co/t/es-6-1-2-cluster-shows-performance-bottleneck/153208/2 "2018-10-22T09:09:16Z")

</div>

Hello,

I don't know if you checked already for comparison, but we have an archive of release benchmarks in the usual page ([https://elasticsearch-benchmarks.elastic.co](https://elasticsearch-benchmarks.elastic.co)) on our bare metal environment; in particular looking at the 99th percentile service\_time for the `geonames` track between `2.4.6` and `6.4.0` on 1node our own benchmarks show:

`scroll` service time is less on `6.4.0`: `666.008ms` vs `751.116ms` on `2.4.6`.  
`country_agg_cached` is basically the same (`3.796`ms vs `3.783`ms)  
`country_agg_uncached` service time is a bit slower in 6.4 (and 5.6) giving `222.651ms` compared to `190.085ms` on `2.4.6` in service time, but nowhere near the 115% increase you are observing.

The first observation is that since your latency is \>\> service\_time in your 6.1.2 Elasticsearch (and this is not observed in the 2.4 setup), the cluster is bottlenecked somewhere (see also [here](https://discuss.elastic.co/t/why-the-percentile-latency-is-several-times-more-than-service-time/69630)).

My first thought would be to check if the environment setup is precisely the same (in terms of h/w) between your 2.4 and 6.1 cluster; e.g. are you using exactly the same instance types for ES nodes and in the same region and availability zone as well?

In addition to that, is the operating system the same (inc. version) for both environments? Apart from differences arising from different kernels and settings, the `i3.2xlarge` instance you are using for the data node benefits from NVMe instance store, however, this can not be efficiently utilized in older Linux kernels.

You mentioned you checked the system metrics, have you in particular looked at io metrics (`iostat -xz 1`)? I am linking here a useful [performance checklist](http://www.brendangregg.com/blog/2016-05-04/srecon2016-perf-checklists-for-sres.html) written by Brendan Gregg for checking resource utilization.

Dimitris

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 19, 2018, 9:22am UTC](https://discuss.elastic.co/t/es-6-1-2-cluster-shows-performance-bottleneck/153208/3 "2018-11-19T09:22:40Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
