# ES performance reliablity issue

**URL:** <https://discuss.elastic.co/t/es-performance-reliablity-issue/247974>\
**Category:** Elasticsearch\
**Created:** [September 9, 2020, 8:06am UTC](https://discuss.elastic.co/t/es-performance-reliablity-issue/247974 "2020-09-09T08:06:17Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![wangxr1985](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wangxr1985/32/117798_2.png) [@wangxr1985](https://discuss.elastic.co/u/wangxr1985)\
**Post date:** [September 9, 2020, 8:06am UTC](https://discuss.elastic.co/t/es-performance-reliablity-issue/247974/1 "2020-09-09T08:06:17Z")

</div>

My query is very simple, just like this:  
{ "query": {"bool": { "filter": {"term": {"user\_guid": "xxxx" }}}}}  
the user\_guid is equal to document id, and the query contains routing argument.

If I do this query several times(use different user\_guid in each query), the "took" value in return json is less than 10ms.

But, If I do this query 100,000 times(different user\_guid each query), there are about 500-1000 queries which take more than 50ms. I want to reduce the number to 0 or less than 10.

If there are 1000 high latency quries(more than 50ms):  
just 100 of them in slowlog, 40% in fetch and 60% in query.  
the other 900 not in slowlog, I can't get any information since there are no logs to analyze.

I do several tests:

1. use different ES versions: 5.6.3 and 7.5.2.
2. use different query types: get , query, filter.
3. use different number of nodes: one node cluster, 10 nodes cluster
4. use different qps: 10/s and 2000/s

All of the tests have the same issue, so I want to know if ES can not keep 100% low latency performance. The reason which causes the issue, and how to inprove it.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 9, 2020, 11:51pm UTC](https://discuss.elastic.co/t/es-performance-reliablity-issue/247974/2 "2020-09-09T23:51:08Z")

</div>

5.X is [EOL](https://www.elastic.co/support/eol), you should not be using it.

> [@wangxr1985](#):
>
> I want to reduce the number to 0

That's going to be impossible unfortunately.

> [@wangxr1985](#):
>
> But, If I do this query 100,000 times(different user\_guid each query), there are about 500-1000 queries which take more than 50ms. I want to reduce the number to 0 or less than 10.

How many records are these queries running against? What is the mapping of the field? How many indices and shards? What does `hot_nodes` look like when you run them? What are your node specs?

---

<div class="post-metadata">

**Author:** ![wangxr1985](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wangxr1985/32/117798_2.png) [@wangxr1985](https://discuss.elastic.co/u/wangxr1985)\
**Post date:** [September 10, 2020, 5:54am UTC](https://discuss.elastic.co/t/es-performance-reliablity-issue/247974/3 "2020-09-10T05:54:08Z")

</div>

400 million documents in the index, each document has 12 fields, the type of each field is keyword.  
The index has 40 shards. the size of index is 70GB.  
The node server is 16vcpu 64Gmem with one 1T ssd disk. The cpu usage is 20% under testing.  
The server performance is not the bottleneck, even if my script runs very slowly (qps 10/s), there are also 0.01%-0.1% of the queries which take more than 50ms.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 10, 2020, 5:59am UTC](https://discuss.elastic.co/t/es-performance-reliablity-issue/247974/4 "2020-09-10T05:59:22Z")

</div>

Why do you have 40 shards for an index of 70GB? I would expect 2-3 shards to be more appropriate. How large is your heap? Garbage collection can cause temporary slowdowns but is often faster the smaller the heap is so [it is important to set the heap size correctly](https://www.elastic.co/blog/a-heap-of-trouble). Larger is not always better.

You may also want to test the most recent version as it has switched to G1GC, which could also affect the latency profile.

---

<div class="post-metadata">

**Author:** ![wangxr1985](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wangxr1985/32/117798_2.png) [@wangxr1985](https://discuss.elastic.co/u/wangxr1985)\
**Post date:** [September 10, 2020, 7:36am UTC](https://discuss.elastic.co/t/es-performance-reliablity-issue/247974/5 "2020-09-10T07:36:20Z")

</div>

Now I do another test, I use a filter query like this  
{  
"profile": true,  
"bool": {  
"filter": {  
"query": {  
"term": {  
"user\_guid": "xxx"  
}  
}  
}  
}  
}

and then I do a query with the same user\_guid 1,600,000 times.  
The results of filter query should hit the query cache, but there are still 233 records in slowlog(more than 50ms, some in query and some in fetch), and the total queries which take more than 50ms is 6926.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 10, 2020, 7:50am UTC](https://discuss.elastic.co/t/es-performance-reliablity-issue/247974/6 "2020-09-10T07:50:37Z")

</div>

Since you have a single node there will at some point be GC which will affect latencies. Please try the things I suggested and see if it makes any difference.

---

<div class="post-metadata">

**Author:** ![wangxr1985](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wangxr1985/32/117798_2.png) [@wangxr1985](https://discuss.elastic.co/u/wangxr1985)\
**Post date:** [September 11, 2020, 7:27am UTC](https://discuss.elastic.co/t/es-performance-reliablity-issue/247974/7 "2020-09-11T07:27:31Z")

</div>

GC is most likely the cause of the issue.  
The frequency of young gc is 1-2/s when I use jstat to print the gc info, and then I switch it to G1GC ,or tune the cms jvm arguments (e.g. -Xmn10g -XX:SurvivorRatio=10 -XX:+UseParNewGC，-XX:MaxTenuringThreshold=15). The frequency of young gc reduces to 0.2/s and the number of queries which take more than 50ms reduce to 1/10-1/5.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 9, 2020, 7:27am UTC](https://discuss.elastic.co/t/es-performance-reliablity-issue/247974/8 "2020-10-09T07:27:53Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
