# How to defrag an index?

**URL:** <https://discuss.elastic.co/t/how-to-defrag-an-index/36732>\
**Category:** Elasticsearch\
**Created:** [December 9, 2015, 11:17am UTC](https://discuss.elastic.co/t/how-to-defrag-an-index/36732 "2015-12-09T11:17:22Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![igor\_k](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_k/32/1157_2.png) [@igor\_k](https://discuss.elastic.co/u/igor_k)\
**Post date:** [December 9, 2015, 11:17am UTC](https://discuss.elastic.co/t/how-to-defrag-an-index/36732/1 "2015-12-09T11:17:23Z")

</div>

Hi,

Our use case:

- pretty big cluster - billions of docs
- we update documents in place
- data are not time-sliced as we often do retrieve and modify old documents
- **Issue** : over time we accumulated a lot of deleted documents in the indices; it is close to 20%
- we are on 1.6.x

It turns out that we have a few segments close to 5GB and using default settings elasticsearch doesn't want to merge them.

We'd like to be able to defragment the cluster to avoid wasting space, especially that the number of deleted docs grows over time.

I see two solutions here:

1. Change the merge policy to something like this

This should help us right now, but it will really push the issue in time, as when we accumulate 20% of deleted docs in these 20gb segments we'll have the same as right now.

1. Manually optimize the indices using optimize API `_optimize?max_num_segments=1`

We can make it a weekly or monthly job, but I'm afraid that the segments will grow unbounded this way and eventually we will kill the cluster performance.

**Q1**. I guess what we really want is some kind of an `_optimize` which will turn e.g.

- 5 \* 5gb shards with 20% of deleted
- into 4 \* 5gb shards with 0% deleted

**Q2**. Is there any other way this usecase should be handled without reindexing?

**Q3**. Do big shards have any negative impact on the cluster?

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 9, 2015, 4:06pm UTC](https://discuss.elastic.co/t/how-to-defrag-an-index/36732/2 "2015-12-09T16:06:07Z")

</div>

Q1: you may want to send an `_optimize` with `only_expunge_deletes=true`

Q2: leave deleted documents in the index and filter them out by a criteria at search time, or rearrange your index organization so old/unneeded indices can be dropped

Q3: yes

---

<div class="post-metadata">

**Author:** ![igor\_k](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_k/32/1157_2.png) [@igor\_k](https://discuss.elastic.co/u/igor_k)\
**Post date:** [December 9, 2015, 5:16pm UTC](https://discuss.elastic.co/t/how-to-defrag-an-index/36732/3 "2015-12-09T17:16:59Z")

</div>

Thanks Jörg for the answers. Looks like, there is no way to merge 5 shards into 4 shards of similar size, right? You can only merge 5 segments into one big segment?

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [December 9, 2015, 5:39pm UTC](https://discuss.elastic.co/t/how-to-defrag-an-index/36732/4 "2015-12-09T17:39:54Z")

</div>

Each shard is an individual Lucene index. Each index is made up of small  
segments, which are immutable and are merged from time to time. The number  
of shards cannot be changed once an index has been created.

I cannot see your original question since this "mailing list" does not  
always deliver emails.

Ivan

---

<div class="post-metadata">

**Author:** ![igor\_k](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_k/32/1157_2.png) [@igor\_k](https://discuss.elastic.co/u/igor_k)\
**Post date:** [December 9, 2015, 7:56pm UTC](https://discuss.elastic.co/t/how-to-defrag-an-index/36732/5 "2015-12-09T19:56:03Z")

</div>

Instead of _shards_ I meant _segments_. Edited the post.

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [December 9, 2015, 8:29pm UTC](https://discuss.elastic.co/t/how-to-defrag-an-index/36732/6 "2015-12-09T20:29:24Z")

</div>

> [@igor\_k](#):
>
> Instead of shards I meant segments.

Segment merging always goes down to a single segment. I've run into this issue before, btw. Ultimately we lived with the delete overhead.

Calling `_optimize` can actually make the trouble worse because it makes even bigger segments which the merge scheduler wants to merge _even less_ than it wants to merge the ones around 5gb.

Its something I've talked about with @mikemccand a few times but never came up with a good solution for.

---

<div class="post-metadata">

**Author:** ![igor\_k](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_k/32/1157_2.png) [@igor\_k](https://discuss.elastic.co/u/igor_k)\
**Post date:** [December 10, 2015, 10:05am UTC](https://discuss.elastic.co/t/how-to-defrag-an-index/36732/7 "2015-12-10T10:05:09Z")

</div>

Thanks Guys, so probably we'll also need to leave with some overhead. We'll try at least to understand the merge policy in more details and maybe tweak it a but to much our use case.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:32pm UTC](https://discuss.elastic.co/t/how-to-defrag-an-index/36732/8 "2017-07-05T23:32:04Z")

</div>


