# Why would the docs.deleted count in some indices exceed 10 times the docs.count?

**URL:** <https://discuss.elastic.co/t/why-would-the-docs-deleted-count-in-some-indices-exceed-10-times-the-docs-count/358223>\
**Category:** Elasticsearch\
**Created:** [April 25, 2024, 1:49pm UTC](https://discuss.elastic.co/t/why-would-the-docs-deleted-count-in-some-indices-exceed-10-times-the-docs-count/358223 "2024-04-25T13:49:17Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![wangxr1985](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wangxr1985/32/117798_2.png) [@wangxr1985](https://discuss.elastic.co/u/wangxr1985)\
**Post date:** [April 25, 2024, 1:49pm UTC](https://discuss.elastic.co/t/why-would-the-docs-deleted-count-in-some-indices-exceed-10-times-the-docs-count/358223/1 "2024-04-25T13:49:17Z")

</div>

OS: CentOS 7.8  
ES version: 7.5.2 with bundled JDK(openjdk version "13.0.1" 2019-10-15)

The default segment merge parameter index.merge.policy.deletes\_pct\_allowed is set to 33, so theoretically the docs.deleted count should not exceed half of the docs.count.

In most of my indices, the docs.deleted count follows this rule, but there are some indices where the docs.deleted count significantly exceeds the threshold for segment merging. For example:  
{  
"docs": {  
"count": 104309,  
"deleted": 1337471  
}

I tried changing the deletes\_pct\_allowed value to 20, but from the monitoring data, it seems that the docs.deleted count is still being cleared only after reaching the original value, indicating that this parameter may not be taking effect.

There are no nested types in my index mapping.

What could be the reason for this discrepancy?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 25, 2024, 3:27pm UTC](https://discuss.elastic.co/t/why-would-the-docs-deleted-count-in-some-indices-exceed-10-times-the-docs-count/358223/2 "2024-04-25T15:27:54Z")

</div>

If you are doing updates, that could explain it.

And update is 1 delete + 1 insert basically...  
So the number of deletion might be high if you are doing a lot of updates.

---

<div class="post-metadata">

**Author:** ![wangxr1985](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wangxr1985/32/117798_2.png) [@wangxr1985](https://discuss.elastic.co/u/wangxr1985)\
**Post date:** [April 26, 2024, 6:14am UTC](https://discuss.elastic.co/t/why-would-the-docs-deleted-count-in-some-indices-exceed-10-times-the-docs-count/358223/3 "2024-04-26T06:14:38Z")

</div>

The index indeed has frequent update operations, but I observed that the docs.deleted cleanup frequency is very low, being cleaned up only every few hours, as shown in the following figure:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/a/3/a3d7a0a419b2beeab60d27753ee3e4074ec6842d.png)

What could be the reason for the docs.deleted count significantly exceeding the docs.count and not being expunged for a long time? Are there any parameter settings that can make the expunge of deleted documents more frequent?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 26, 2024, 7:00am UTC](https://discuss.elastic.co/t/why-would-the-docs-deleted-count-in-some-indices-exceed-10-times-the-docs-count/358223/4 "2024-04-26T07:00:56Z")

</div>

The [Force Merge API](https://www.elastic.co/guide/en/elasticsearch/reference/8.13/indices-forcemerge.html) could help. See the `only_expunge_deletes` option.

May be you can control `index.merge.policy.expunge_deletes_allowed` index setting? It defaults to 10. It's not mentioned in the [merge settings documentation](https://www.elastic.co/guide/en/elasticsearch/reference/8.13/index-modules-merge.html) though. So may be it's not recommended to change the value...

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 26, 2024, 7:07am UTC](https://discuss.elastic.co/t/why-would-the-docs-deleted-count-in-some-indices-exceed-10-times-the-docs-count/358223/5 "2024-04-26T07:07:20Z")

</div>

Merging can be an expensive and I/O intensive operation so there is a good reason it is not necessarily performed very aggressively by default. Why are you looking to change this and make it more aggressive? Is it using up too much disk space? Is it affecting the page cache hit rate? Are you seeing a performance impact as the number of deleted documents grow?

---

<div class="post-metadata">

**Author:** ![wangxr1985](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wangxr1985/32/117798_2.png) [@wangxr1985](https://discuss.elastic.co/u/wangxr1985)\
**Post date:** [April 26, 2024, 7:33am UTC](https://discuss.elastic.co/t/why-would-the-docs-deleted-count-in-some-indices-exceed-10-times-the-docs-count/358223/6 "2024-04-26T07:33:36Z")

</div>

The main purpose is to improve query performance, as a high number of `docs.deleted` can result in increased CPU usage during queries. For example, in the chart below, each time `docs.deleted` is expunged, there is a noticeable decrease in CPU usage.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/1/2/127e16e036897c6c5271e6c6c4d0271e62f21288.png)  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/c/9/c92873c5f3466b5b26c974b048155aa7cbc5d6d2.png)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 26, 2024, 7:43am UTC](https://discuss.elastic.co/t/why-would-the-docs-deleted-count-in-some-indices-exceed-10-times-the-docs-count/358223/7 "2024-04-26T07:43:19Z")

</div>

The CPU usage does by itself not seem excessive or problematic. Is it resulting in longer query latencies? Are you tracking query latency as function of deleted document percentage? Is there a correlation?

What are the issues around query performance you are seeing?

---

<div class="post-metadata">

**Author:** ![wangxr1985](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wangxr1985/32/117798_2.png) [@wangxr1985](https://discuss.elastic.co/u/wangxr1985)\
**Post date:** [April 26, 2024, 8:09am UTC](https://discuss.elastic.co/t/why-would-the-docs-deleted-count-in-some-indices-exceed-10-times-the-docs-count/358223/8 "2024-04-26T08:09:19Z")

</div>

The query QPS of this ES cluster is very high, with many data nodes in the cluster. If we can keep the CPU usage as shown in the chart above at 20% and prevent it from rising to 30%, we could potentially reduce the number of servers by close to one-third. Each server is equipped with NVMe SSD drives, so the disk I/O load from segment merging is not a bottleneck.

I aim to configure the cluster to prioritize the cleanup of deleted documents during regular segment merges, maintaining a low count of doc.deleted to keep the CPU usage at a lower level.
