# Term query by \_id very slow (30s+) occasionally

**URL:** <https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305>\
**Category:** Elasticsearch\
**Created:** [September 18, 2023, 11:33pm UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305 "2023-09-18T23:33:02Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![May\_Zeng](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/may_zeng/32/123964_2.png) [@May\_Zeng](https://discuss.elastic.co/u/May_Zeng)\
**Post date:** [September 18, 2023, 11:33pm UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/1 "2023-09-18T23:33:02Z")

</div>

Mapping:

```auto
{
  "dynamic": "strict",
  "_source": {
    "enabled": false
  },
  "properties": {
    "data": {
      "type": "binary",
      "doc_values": false,
      "store": true
    }
  }
}

```

Query:

```auto
{"query":{
  "bool" : {
    "filter" : [
      {
        "term" : {
          "_id" : {
            "value" : "a7884205-3d10-11ee-8fe7-0b299f77c536",
            "boost" : 1.0
          }
        }
      }
    ],
    "adjust_pure_negative" : true,
    "boost" : 1.0
  }
}
}

```

---

<div class="post-metadata">

**Author:** ![May\_Zeng](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/may_zeng/32/123964_2.png) [@May\_Zeng](https://discuss.elastic.co/u/May_Zeng)\
**Post date:** [September 19, 2023, 5:11am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/2 "2023-09-19T05:11:01Z")

</div>

3 ES nodes (1 replica, 3 shards)  
Each node: 4Core, 32G RAM and 14G heap  
1T data (100 \* 10G)  
JVM statics as following:

```auto
"jvm": {
                  "timestamp": 1695089656498,
                  "uptime_in_millis": 8293249814,
                  "mem": {
                      "heap_used_in_bytes": 5938777088,
                      "heap_used_percent": 39,
                      "heap_committed_in_bytes": 15032385536,
                      "heap_max_in_bytes": 15032385536,
                      "non_heap_used_in_bytes": 270824264,
                      "non_heap_committed_in_bytes": 277282816,
                      "pools": {
                          "young": {
                              "used_in_bytes": 3430940672,
                              "max_in_bytes": 0,
                              "peak_used_in_bytes": 9009364992,
                              "peak_max_in_bytes": 0
                          },
                          "old": {
                              "used_in_bytes": 2340064256,
                              "max_in_bytes": 15032385536,
                              "peak_used_in_bytes": 10826308096,
                              "peak_max_in_bytes": 15032385536
                          },
                          "survivor": {
                              "used_in_bytes": 167772160,
                              "max_in_bytes": 0,
                              "peak_used_in_bytes": 700494512,
                              "peak_max_in_bytes": 0
                          }
                      }
                  },
                  "threads": {
                      "count": 194,
                      "peak_count": 197
                  },
                  "gc": {
                      "collectors": {
                          "young": {
                              "collection_count": 27822,
                              "collection_time_in_millis": 1431077
                          },
                          "old": {
                              "collection_count": 0,
                              "collection_time_in_millis": 0
                          }
                      }
                  },
                  "buffer_pools": {
                      "mapped": {
                          "count": 9938,
                          "used_in_bytes": 511147464733,
                          "total_capacity_in_bytes": 511147464733
                      },
                      "direct": {
                          "count": 173,
                          "used_in_bytes": 11555891,
                          "total_capacity_in_bytes": 11555889
                      },
                      "mapped - 'non-volatile memory'": {
                          "count": 0,
                          "used_in_bytes": 0,
                          "total_capacity_in_bytes": 0
                      }
                  },
                  "classes": {
                      "current_loaded_count": 29654,
                      "total_loaded_count": 31313,
                      "total_unloaded_count": 1659
                  }
              }

```

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 19, 2023, 5:32am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/3 "2023-09-19T05:32:16Z")

</div>

Which version of Elasticsearch are you using?

Do you have a single index with 3 primary shards and 1 replica in the cluster? If so, what is the size of this index in trem of shard size in GB as well as document count? What is the average size of the data in this binary field?

What type of storage are you using? Local SSD?

What load is the cluster under (read and write) when you see these long latencies?

This usage pattern is quite unusual. It looks to me like you are trying to use Elasticsearch as a key-value store for some binary data. Why are you using Elasticsearch this way?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood1/32/101255_2.png) [@Mark\_Harwood1](https://discuss.elastic.co/u/Mark_Harwood1)\
**Post date:** [September 19, 2023, 6:50am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/4 "2023-09-19T06:50:45Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> This usage pattern is quite unusual. It looks to me like you are trying to use Elasticsearch as a key-value store for some binary data

Yeah I’d be interested to know how big the blobs are in this doc.

---

<div class="post-metadata">

**Author:** ![May\_Zeng](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/may_zeng/32/123964_2.png) [@May\_Zeng](https://discuss.elastic.co/u/May_Zeng)\
**Post date:** [September 19, 2023, 8:42pm UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/5 "2023-09-19T20:42:20Z")

</div>

Hi Christian,  
Thanks for you reply.  
Here are the details:  
ES version: 8.4.2  
Each index 3 primary shards, 1 replica  
ILM:

```auto
 {
      "max_size": "5gb",
      "max_primary_shard_size": "5gb",
      "max_age": "1d",
      "max_docs": 10000000
  }

```

Average size of binary field data: 8kb. The followings are part of the indices.

```auto
health status index uuid pri rep docs.count docs.deleted store.size pri.store.size
green open messages-000370 DneyyWoHQp-T17_kjn9kCg 3 1 1471584 0 12gb 6gb
green open messages-000371 -h50Q3JiSrmsu0mmt7ZCMw 3 1 1394995 0 10.9gb 5.4gb
green open messages-000372 XQFIvqAvSeCRMe_GCNYZ5A 3 1 1380346 0 11gb 5.5gb
green open messages-000373 wTQC9L3dS_CiRMbC8rSJpQ 3 1 1655118 0 10.3gb 5.1gb
green open messages-000374 Wuft_rovRDqs1i8mX_Ygyw 3 1 1182124 0 10gb 5gb

```

When doing indexing and search - long latencies, CPU usage is less than 20%, OS memory is roughly 18G. The hardware resource is reasonable but the performance is not good. The worst case is 30s+, normal situation is seconds (2-10s).

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 19, 2023, 9:51pm UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/6 "2023-09-19T21:51:21Z")

</div>

How many indices do you have in the cluster? What type of storage do you have?

---

<div class="post-metadata">

**Author:** ![May\_Zeng](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/may_zeng/32/123964_2.png) [@May\_Zeng](https://discuss.elastic.co/u/May_Zeng)\
**Post date:** [September 19, 2023, 9:59pm UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/7 "2023-09-19T21:59:10Z")

</div>

Currently the total indices is 258 and message itself is 185.  
The entry point: 20-40 messages per second.

---

<div class="post-metadata">

**Author:** ![May\_Zeng](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/may_zeng/32/123964_2.png) [@May\_Zeng](https://discuss.elastic.co/u/May_Zeng)\
**Post date:** [September 20, 2023, 12:48am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/8 "2023-09-20T00:48:22Z")

</div>

The storage is SSD + SATA ( HCI: hyper-converged infrastructure)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 20, 2023, 4:34am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/9 "2023-09-20T04:34:58Z")

</div>

You did not answer the following questions:

> [@Christian\_Dahlqvist](#):
>
> What load is the cluster under (read and write) when you see these long latencies?

How many documents are indexed per second? What bulk size are you using? How many concurrent queries are you running when you see the large latencies?

> [@Christian\_Dahlqvist](#):
>
> This usage pattern is quite unusual. It looks to me like you are trying to use Elasticsearch as a key-value store for some binary data. Why are you using Elasticsearch this way?

Can you please describe the use case and why you are using Elasticsearch this way?

> [@May\_Zeng](#):
>
> ```auto
> {
> "max_size": "5gb",
> "max_primary_shard_size": "5gb",
> "max_age": "1d",
> "max_docs": 10000000
> }
> 
> ```

Why have you limited your shard size to 5GB? That seems quite small and will result in a large number of shards that need to be searched in every query.

---

<div class="post-metadata">

**Author:** ![May\_Zeng](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/may_zeng/32/123964_2.png) [@May\_Zeng](https://discuss.elastic.co/u/May_Zeng)\
**Post date:** [September 20, 2023, 4:53am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/10 "2023-09-20T04:53:33Z")

</div>

How many documents are indexed per second? What bulk size are you using? How many concurrent queries are you running when you see the large latencies?  
At least 40 message/s at the entry points. The total must be more than that. Bulk action: 5000, flushInterval: 5s; Only one query when you see the large latencies.

Can you please describe the use case and why you are using Elasticsearch this way?  
Sure. The system is kind of trace e.g. a message come to the system and the system processes the message via A-\>B-\>C-\>D-\>E. And at each point, we trace the event and keep the message content in order to reprocess the message. It may not be the perfect design but it is our current design. We may improve it to extract message out to other DB but not at this stage.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 20, 2023, 5:09am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/11 "2023-09-20T05:09:29Z")

</div>

What is the retention period for this data in the cluster?

Is the amount of data indexed static or expected to grow in the future?

---

<div class="post-metadata">

**Author:** ![May\_Zeng](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/may_zeng/32/123964_2.png) [@May\_Zeng](https://discuss.elastic.co/u/May_Zeng)\
**Post date:** [September 20, 2023, 5:37am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/12 "2023-09-20T05:37:28Z")

</div>

Retention days is 14. The system has run 7 days which means the data would be kind of double i.e. 2T.

---

<div class="post-metadata">

**Author:** ![May\_Zeng](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/may_zeng/32/123964_2.png) [@May\_Zeng](https://discuss.elastic.co/u/May_Zeng)\
**Post date:** [September 20, 2023, 5:39am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/13 "2023-09-20T05:39:39Z")

</div>

Why have you limited your shard size to 5GB?  
Is there any formula we can follow to set the right value of the shard size?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 20, 2023, 5:46am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/14 "2023-09-20T05:46:24Z")

</div>

Given the number of indices in the cluster I will assume your retention period is reasonably long, potentially up to 6 months.

Although I am not sure Elasticsearch is the ideal choice for this use case, I think there are a few things you could try in order to improve performance.

**1. Limit the number of shards queried**  
When Elasticsearch indexes a document it routs the document to one of the shards based on the document ID by default. Your query searches for one spcific document ID but queries all shards. You can try adding the document ID as a [routing value](https://www.elastic.co/guide/en/elasticsearch/reference/8.9/search-shard-routing.html#search-routing) with each search request. This should allow Elasticsearch to only search one shard (the one that could hold the document based on the ID) for each index.

This should work even if you change the number of primary shards.

**2. Increase shard size**  
You currently have a lot of very small shards, which can be slow to query. I would recommend increasing the shard size so you get fewer shards to query. At the same time I would also recommend you increase the number of primary shards as that will make the routing change described earlier more efficient. If you go to e.g. 6 primary shards per index only 1/6 of all shards would need to be searched instead of 1/3. This does not mean that you need to reindex - you can simply let the smaller indices with 3 primary shards age out over time.

This likely mean that you will need to make each index cover a larger time period, but that should not be a problem if your retention period is quite long. Maybe change the rollover criteria to something like this:

```auto
{
    "max_primary_shard_size": "50gb",
    "max_age": "10d"
}

```

**3. Forcemerge indices no longer written to**  
It may also be beneficial to add a step to ILM that forcemerges indices that have rolled over down to a single segment if you do not already have this in place.

---

<div class="post-metadata">

**Author:** ![May\_Zeng](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/may_zeng/32/123964_2.png) [@May\_Zeng](https://discuss.elastic.co/u/May_Zeng)\
**Post date:** [September 20, 2023, 6:20am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/15 "2023-09-20T06:20:12Z")

</div>

Really appreciate your professional suggestions.

1. We do include routing in request.

2. About the shard size and shard count, is there any consideration? Like nodes count, hard drive or memory? Currently most of the case, we keep data for 14 days, maybe shorter.

3. We do a little bit upsert to data. From hot to warm stage, I can add merge. Will this affect upsert?

4. The performance issues only happens when messages come through i.e. indexing. I am thinking whether it helps if increasing refresh interval and increasing query cache size or not. Any suggestions?  
Thanks.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 20, 2023, 6:36am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/16 "2023-09-20T06:36:00Z")

</div>

> [@May\_Zeng](#):
>
> We do include routing in request.

If you are already doing this that is good.

> [@May\_Zeng](#):
>
> About the shard size and shard count, is there any consideration? Like nodes count, hard drive or memory? Currently most of the case, we keep data for 14 days, maybe shorter.

5GB is a quite small shard size. Given that you search by ID I would expect you to benefit from larger shards. 50GB shard size is not uncommon and something I would try to achieve.

A larger number of primary shards would make your routing more efficient, so would probably be a worthwhile change. I would look to increase this gradually, e.g. initially go from 3 to 6. If you with this achieve the larger shard size you may later increase this to 9.

If you have a short retention period, e.g. 14 days, you need to make sure that you get a suitable number of indices to cover this period. In this case I would probably keep the rollover period at 1 day so you get at least 14 indices covering the retention period.

> [@May\_Zeng](#):
>
> We do a little bit upsert to data. From hot to warm stage, I can add merge. Will this affect upsert?

If you are updating existing data it does not make any sense to forcemerge down to a single segment. Forcemerging down to a single segment is I/O intensive and completely undone if you update or write to the index.

> [@May\_Zeng](#):
>
> The performance issues only happens when messages come through i.e. indexing. I am thinking whether it helps if increasing refresh interval and increasing query cache size or not. Any suggestions?

If that is the case I would look at storage I/O performance, especially at times you are performing indexing and see slow queries. What does `iostat -x` show?

---

<div class="post-metadata">

**Author:** ![May\_Zeng](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/may_zeng/32/123964_2.png) [@May\_Zeng](https://discuss.elastic.co/u/May_Zeng)\
**Post date:** [September 21, 2023, 3:25am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/17 "2023-09-21T03:25:37Z")

</div>

Thanks Christian. I will try 6 shards and shard size = 30G first.

> [@Christian\_Dahlqvist](#):
>
> If that is the case I would look at storage I/O performance, especially at times you are performing indexing and see slow queries. What does `iostat -x` show?

Here is one of the nodes iostat:

```auto
avg-cpu: %user %nice %system %iowait %steal %idle
          28.18 0.00 8.49 43.96 0.00 19.37

Device r/s rMB/s rrqm/s %rrqm r_await rareq-sz w/s wMB/s wrqm/s %wrqm w_await wareq-sz d/s dMB/s drqm/s %drqm d_await dareq-sz f/s f_await aqu-sz %util
dm-0 0.00 0.00 0.00 0.00 0.00 0.00 1.50 0.01 0.00 0.00 39.00 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.06 2.00
dm-1 0.00 0.00 0.00 0.00 0.00 0.00 46.50 0.18 0.00 0.00 66.11 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 3.07 19.25
dm-2 402.50 177.94 0.00 0.00 103.34 452.70 136.00 15.49 0.00 0.00 77.03 116.66 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 52.07 95.50 
sda 570.50 181.68 0.00 0.00 100.72 326.11 44.00 13.04 150.00 77.32 71.44 303.55 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 60.60 96.50
sr0 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00

```

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 21, 2023, 5:15am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/18 "2023-09-21T05:15:47Z")

</div>

That is very high (bad) values for `r_await` and `w_await`, which indicates that the performance of the storage you are using likely is the bottleneck.

I would looking into concentrating on improving the storage you are using. If the storage is this slow I do not think any of the other optimisations will make much difference.

---

<div class="post-metadata">

**Author:** ![May\_Zeng](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/may_zeng/32/123964_2.png) [@May\_Zeng](https://discuss.elastic.co/u/May_Zeng)\
**Post date:** [September 24, 2023, 10:46pm UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/19 "2023-09-24T22:46:23Z")

</div>

Hi Christian,  
Thanks for your help. We setup an environment with 30G + 6 shards with SSD had drive. The performance of query while writing is not ideal still (20s +). I am thinking that might help if we can separate the write and read. Does Es support write and read separately? like one node is setup for writing, another node is setup for reading whose data is syncing from writing node.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 25, 2023, 3:14am UTC](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305/20 "2023-09-25T03:14:07Z")

</div>

No, that kind of separation is not possible.

[Next page](https://discuss.elastic.co/t/term-query-by-id-very-slow-30s-occasionally/343305.md?page=2)
