# Tuning merge policy for large number of updates

**URL:** <https://discuss.elastic.co/t/tuning-merge-policy-for-large-number-of-updates/53715>\
**Category:** Elasticsearch\
**Created:** [June 22, 2016, 10:13pm UTC](https://discuss.elastic.co/t/tuning-merge-policy-for-large-number-of-updates/53715 "2016-06-22T22:13:33Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![mkelkar](https://avatars.discourse-cdn.com/v4/letter/m/ea666f/32.png) [@mkelkar](https://discuss.elastic.co/u/mkelkar)\
**Post date:** [June 22, 2016, 10:13pm UTC](https://discuss.elastic.co/t/tuning-merge-policy-for-large-number-of-updates/53715/1 "2016-06-22T22:13:34Z")

</div>

Hi All,  
I am using ES 1.7.3. In my cluster I have a large number of updates daily. Ever since we switched from ES 1.3.9 to ES 1.7.3, we are starting to notice about 150% more disk usage as these updates happen. I looked at the indices and segments, and it turned out that additional disk usage is due to number of deleted ( ie updated) documents. If I use `_optimize?only_expunge_deletes=true`, I am able to reclaim disk space. but ES does not seem to reclaim disk space when merges occur...

Here are our merge policy settings

```
        "index.merge.policy.reclaim_deletes_weight": "6.0"
        "index.merge.policy.max_merged_segment": "1gb"
        "index.store.throttle.type": "merge"     
    "index.store.throttle.max_bytes_per_sec": "200mb"

```

With ES 1.3.9, we used to have

```
    "index.merge.policy.reclaim_deletes_weight": "2.0"
    "index.merge.policy.max_merged_segment": "5gb"
    "index.store.throttle.type": "none"     
"index.store.throttle.max_bytes_per_sec": "20mb"

```

This worked perfectly fine with no disk growth issues.

With ES 1.3.9 disk usage was not a problem, despite large number of updates happening, but with 1.7.3 I see disk usage creep up alarmingly.

I do have multi-gigabyte shards due to large data volume. This cluster also has heavy query rate ( about 4-5 thousand queries per second ). I use SSD backed machines, and haven't seen disk I/O or memory bottlenecks. Its only the disk usage that is the problem. How do I tune my merge policy to keep up with updates? I am ok if indexing is a little slow due to additional merges, but I cannot keep up with this kind of disk usage. Please advice!

Thanks,  
Madhav.

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [June 23, 2016, 1:16pm UTC](https://discuss.elastic.co/t/tuning-merge-policy-for-large-number-of-updates/53715/2 "2016-06-23T13:16:15Z")

</div>

> [@mkelkar](#):
>
> "index.merge.policy.max\_merged\_segment": "1gb"

Could you have a look at [\_cat segments](https://www.elastic.co/guide/en/elasticsearch/reference/current/cat-segments.html) and see if the biggest segments are the ones with all the deleted documents?

---

<div class="post-metadata">

**Author:** ![mkelkar](https://avatars.discourse-cdn.com/v4/letter/m/ea666f/32.png) [@mkelkar](https://discuss.elastic.co/u/mkelkar)\
**Post date:** [June 23, 2016, 3:02pm UTC](https://discuss.elastic.co/t/tuning-merge-policy-for-large-number-of-updates/53715/3 "2016-06-23T15:02:41Z")

</div>

Hi Nik,  
Yes, this is what it looks like -

```
&part-largecatalog0 5 r 10.101.54.105 _1iqg 70936 2738676 346368 3.7gb 10107002 true true 4.10.4 true
&part-largecatalog0 5 p 10.101.56.193 _1n2o 76560 2737806 345876 3.7gb 10106554 true true 4.10.4 true
&part-largecatalog0 2 r 10.101.56.28 _1cao 62592 2737559 345841 3.7gb 10109746 true true 4.10.4 true
&part-largecatalog0 2 p 10.101.61.227 _1g8x 67713 2734559 345039 3.7gb 10098938 true true 4.10.4 true
&part-largecatalog0 5 r 10.101.60.45 _15za 54406 2733689 344958 3.7gb 10090434 true true 4.10.4 true
&part-largecatalog0 2 r 10.101.55.203 _1bgr 61515 2733323 344583 3.7gb 10097050 true true 4.10.4 true
&part-largecatalog0 1 p 10.101.53.160 _1e64 65020 2676213 333414 3.6gb 9854106 true true 4.10.4 true
&part-largecatalog0 0 r 10.101.59.128 _104z 46835 2689569 333226 3.6gb 9937554 true true 4.10.4 true
&part-largecatalog0 4 p 10.101.63.231 _1562 53354 2678859 332846 3.6gb 9896026 true true 4.10.4 true
&part-largecatalog0 4 r 10.101.59.42 _1af8 60164 2679411 332785 3.6gb 9900530 true true 4.10.4 true

```

Earlier I had all the default settings for ES 1.7, except

```
"index.store.throttle.type": "none" 

```

When this index was created, the max\_merged\_segment was default ( 5gb). When I saw deletes getting accumulated, I set it to 1gb and tried to tune settings in order to recover space. But this did not work...The percentage of deleted documents is spread evenly across segments and is about 12%. Considering this, I have a couple of questions -

1. How much percentage of deleted docs is needed for a segment to be considered for merging?
2. If I had 1gb max\_merged\_segment right before loading data, would it help increasing percentage of deleted docs? Would reindexing help?

Thanks,  
Madhav.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:41pm UTC](https://discuss.elastic.co/t/tuning-merge-policy-for-large-number-of-updates/53715/4 "2017-07-05T22:41:03Z")

</div>


