# \_source 50% bigger after reindex

**URL:** <https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805>\
**Category:** Elasticsearch\
**Created:** [May 3, 2022, 8:22am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805 "2022-05-03T08:22:19Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [May 3, 2022, 8:22am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/1 "2022-05-03T08:22:19Z")

</div>

Hello.

Context:  
I reindexed a 7.15 index into 7.16 one and suddenly it is taking about +50% space.  
By [Analyze index disk usage API | Elasticsearch Guide [7.16] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/7.16/indices-disk-usage.html) I found that \_source field alone is now taking +64% more disk space than before.

Problem:  
What happened? I don't see any difference of index settings that could explain the increase.  
I suspected a change in index.codec, but that is not showing in diff of settings either.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 3, 2022, 8:29am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/2 "2022-05-03T08:29:04Z")

</div>

What was the reindex you did?  
What is the mapping(s)?

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [May 3, 2022, 9:26am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/4 "2022-05-03T09:26:59Z")

</div>

I purpused the index.codec idea further.  
Per [Provide index level indication when index.codec is set at the node level · Issue #26130 · elastic/elasticsearch · GitHub](https://github.com/elastic/elasticsearch/issues/26130) and [Add missing fields to \_segments API · Issue #3160 · elastic/elasticsearch-net · GitHub](https://github.com/elastic/elasticsearch-net/issues/3160) index.codec should be available by Index segments api.

But no luck: both `GET /srcindex/_segments?pretty` and `GET /dstindex/_segments?pretty` return `"attributes" : { "Lucene87StoredFieldsFormat.mode" : "BEST_SPEED" }`

* * *

I already deleted the orignal index. But I am sure it will also happen for index I am reindexing right now - will send you (how?) mappings/settings diff after reindex completes.  
I am looking into another possibility - so far I always reindexed on debian-based node. This time I reindexed on redhat-based ones.

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [May 4, 2022, 10:13am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/5 "2022-05-04T10:13:28Z")

</div>

I reindexed a similar index, details below:

> **Index settings+mappings diff**
>
> ```auto
> @@ -1,4 +1,4 @@
> -GET /srcindex?pretty
> +GET /dstindex?pretty
> {
> - "srcindex" : {
> + "dstindex" : {
> "aliases" : { },
> @@ -109,7 +109,5 @@
> },
> "message" : {
> - "type" : "text",
> - "index" : false,
> - "norms" : false
> + "type" : "match_only_text"
> },
> "msg" : {
> @@ -152,9 +150,6 @@
> "type" : "long"
> },
> - "session_id" : {
> - "type" : "keyword"
> - },
> "sessionid" : {
> - "type" : "long"
> + "type" : "keyword"
> },
> "set" : {
> @@ -243,5 +238,5 @@
> "lifecycle" : {
> "name" : "ILM-LOG",
> - "rollover_alias" : ""
> + "origination_date" : "1640390409363"
> },
> "routing" : {
> @@ -258,6 +253,6 @@
> "read" : "false"
> },
> - "provided_name" : "srcindex",
> - "creation_date" : "1640390409363",
> + "provided_name" : "dstindex",
> + "creation_date" : "1651492970910",
> "unassigned" : {
> "node_left" : {
> @@ -275,7 +270,7 @@
> "priority" : "50",
> "number_of_replicas" : "1",
> - "uuid" : "xxxxxxxxxxxxxxxxxxxxxx",
> + "uuid" : "yyyyyyyyyyyyyyyyyyyyyy",
> "version" : {
> - "created" : "7150299"
> + "created" : "7160199"
> }
> }
> 
> ```

> **Disk usage**
>
> Comparison by field:
> 
> ```auto
> srcindex[MB] dstindex[MB] increase [MB]
> _seq_no 214971712 158950924 -560
> <snip> -2 to +2
> timestamp 289979371 291058317 12
> _field_names - 184978352 19
> _id 55045137 584582451 341
> message - 767938163 7679
> _source 337282685 424764811 8748
> total 670767153 83313122 16236
> 
> ```
> 
> Comparison by type:
> 
> ```auto
> srcindex[MB] dstindex[MB] increase [MB]
> points 540927139 497632488 -433
> doc_values 148172165 146996708 -118
> norms 0 0 0
> inverted_index 118459334 195780484 7733
> stored_fields 35005294 440580779 9053
> total 670767153 83313122 16236
> 
> ```

> **\_cat segments**
>
> Original:
> 
> ```auto
> /_cat/segments/srcindex?h=index,shard,prirep,segment,generation,docs.count,docs.deleted,size,size.memory,committed,searchable,version,compound&v
> index shard prirep segment generation docs.count docs.deleted size size.memory committed searchable version compound
> srcindex 0 p _xxaa 341509 134835092 0 21.8gb 21404 true true 8.9.0 false
> srcindex 0 r _xxaa 341509 134835092 0 21.8gb 21404 true true 8.9.0 false
> srcindex 1 p _xxbb 342018 134812277 0 21.8gb 21404 true true 8.9.0 false
> srcindex 1 r _xxbb 342018 134812277 0 21.8gb 21404 true true 8.9.0 false
> srcindex 2 p _xxcc 341409 134841646 0 21.8gb 21404 true true 8.9.0 false
> srcindex 2 r _xxcc 341409 134841646 0 21.8gb 21404 true true 8.9.0 false
> 
> ```
> 
> Reindexed:
> 
> ```auto
> /_cat/segments/dstindex?h=index,shard,prirep,segment,generation,docs.count,docs.deleted,size,size.memory,committed,searchable,version,compound&v
> index shard prirep segment generation docs.count docs.deleted size size.memory committed searchable version compound
> dstindex 0 p _xx 688 134835092 0 27.1gb 143212 true true 8.11.1 false
> dstindex 0 r _xx 688 134835092 0 27.1gb 143212 true true 8.11.1 false
> dstindex 1 p _xx 688 134812277 0 27.1gb 143212 true true 8.11.1 false
> dstindex 1 r _xx 688 134812277 0 27.1gb 143212 true true 8.11.1 false
> dstindex 2 p _yy 1063 134841646 0 27.1gb 143212 true true 8.11.1 false
> dstindex 2 r _yy 1063 134841646 0 27.1gb 143212 true true 8.11.1 false
> 
> ```

This time \_source increased only by 25%, taking about as many extra bytes as new inverted index of message.

* * *

I'll try reindexing without adding inverted index for message and see what happens.

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [May 4, 2022, 1:58pm UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/6 "2022-05-04T13:58:12Z")

</div>

I would suggest I would suggest If you want to compare the sizes, do a force merge to one segment on the source index and then a force merge to one segment on the destination index. After it finishes, then you'll have an apples to apples comparison.

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [May 5, 2022, 9:52am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/8 "2022-05-05T09:52:55Z")

</div>

All indices I compare are already forcemerged, sorry about the confusion.  
You can verify that that in `_cat segments` output.

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [May 6, 2022, 9:19am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/9 "2022-05-06T09:19:09Z")

</div>

This is the actual (and in my opinion valid) comparision, using different index with similar data.  
Plain reindex, no mapping changes, both indices forcemerged, still +25% size increase on \_source.

> **Reindexing command**
>
> ```auto
> POST /_reindex?slices=auto&pretty
> {
> "source": {
> "index": "srcindex"
> },
> "dest": {
> "index": "identmap"
> }
> }
> POST /identmap/_forcemerge?max_num_segments=1&pretty
> {
> "_shards" : {
> "total" : 6,
> "successful" : 6,
> "failed" : 0
> }
> }
> 
> ```

> **Index settings+mappings diff**
>
> ```auto
> @@ -1,5 +1,5 @@
> -GET /srcindex?pretty
> +GET /identmap?pretty
> {
> - "srcindex" : {
> + "identmap" : {
> "aliases" : { },
> "mappings" : {
> @@ -243,5 +243,5 @@
> "lifecycle" : {
> "name" : "SIEM-LOG",
> - "rollover_alias" : ""
> + "origination_date" : "1640390409363"
> },
> "routing" : {
> @@ -258,6 +258,6 @@
> "read" : "false"
> },
> - "provided_name" : "srcindex",
> - "creation_date" : "1640390409363",
> + "provided_name" : "identmap",
> + "creation_date" : "1651742088307",
> "unassigned" : {
> "node_left" : {
> @@ -275,7 +275,7 @@
> "priority" : "50",
> "number_of_replicas" : "1",
> - "uuid" : "xxxxxxxxxxxxxxxxxxxxxx",
> + "uuid" : "aaaaaaaaaaaaaaaaaaaaaa",
> "version" : {
> - "created" : "7150299"
> + "created" : "7160199"
> }
> }
> 
> ```

> **Disk usage**
>
> Comparison by field:
> 
> ```auto
> srcindex[MB] identmap[MB] increase [MB]
> _seq_no 2149 1587 -562
> <snip> -2 to +2
> timestamp 2899 2910 11
> _id 5505 5819 314
> _source 33729 42095 8366
> total 67077 75204 8127
> 
> ```
> 
> Comparison by type:
> 
> ```auto
> srcindex[MB] identmap[MB] increase [MB]
> points 5409 4974 -435
> doc_values 14817 14699 -118
> inverted_index 11845 11842 -3
> norms 0 0 0
> term_vectors 0 0 0
> stored_fields 35005 43688 8683
> total 67077 75204 8127
> 
> ```

> **\_cat segments**
>
> srcindex(original):
> 
> ```auto
> /_cat/segments/srcindex?h=index,shard,prirep,segment,generation,docs.count,docs.deleted,size,size.memory,committed,searchable,version,compound&v
> index shard prirep segment generation docs.count docs.deleted size size.memory committed searchable version compound
> srcindex 0 p _xxaa 341509 134835092 0 21.8gb 21404 true true 8.9.0 false
> srcindex 0 r _xxaa 341509 134835092 0 21.8gb 21404 true true 8.9.0 false
> srcindex 1 p _xxbb 342018 134812277 0 21.8gb 21404 true true 8.9.0 false
> srcindex 1 r _xxbb 342018 134812277 0 21.8gb 21404 true true 8.9.0 false
> srcindex 2 p _xxcc 341409 134841646 0 21.8gb 21404 true true 8.9.0 false
> srcindex 2 r _xxcc 341409 134841646 0 21.8gb 21404 true true 8.9.0 false
> 
> ```
> 
> identmap(reindexed)
> 
> ```auto
> /_cat/segments/identmap?h=index,shard,prirep,segment,generation,docs.count,docs.deleted,size,size.memory,committed,searchable,version,compound&v
> index shard prirep segment generation docs.count docs.deleted size size.memory committed searchable version compound
> identmap 0 p _aa 697 134835092 0 24.4gb 140556 true true 8.11.1 false
> identmap 0 r _aa 697 134835092 0 24.4gb 140556 true true 8.11.1 false
> identmap 1 p _bb 727 134812277 0 24.4gb 140556 true true 8.11.1 false
> identmap 1 r _bb 727 134812277 0 24.4gb 140556 true true 8.11.1 false
> identmap 2 p _cc 699 134841646 0 24.4gb 140556 true true 8.11.1 false
> identmap 2 r _cc 699 134841646 0 24.4gb 140556 true true 8.11.1 false
> 
> ```

I am sorry for all the previous unnecessary messages. I should have done all this triage before posting.

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [May 9, 2022, 7:49am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/10 "2022-05-09T07:49:07Z")

</div>

@stephenb Sorry to bother you, but how do I forcemerge a specific segment of an index?  
Both [Force merge API | Elasticsearch Guide [7.17] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/7.17/indices-forcemerge.html) and [Force merge API | Elasticsearch Guide [8.2] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-forcemerge.html) merge all segments of an index. I don't see any option to target a specific segment.

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [May 9, 2022, 1:22pm UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/11 "2022-05-09T13:22:07Z")

</div>

As Far as I know there is no way to target a specific segment.

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [May 9, 2022, 1:50pm UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/12 "2022-05-09T13:50:26Z")

</div>

Ah, okay. I though it would be weird to merge just one segment, but wanted to be sure.  
So, all shards I am comparing are merged down to one segment. Is that enough for comparison?

> [@stephenb](#):
>
> I would suggest I would suggest If you want to compare the sizes, do a force merge to one segment on the source index and then a force merge to one segment on the destination index. After it finishes, then you'll have an apples to apples comparison.

The source index was forcemerged back in 7.15. Does it make sense to re-forcemerge it (1-\>1)?  
So far I was trying to avoid that, because then I'd loose my test case (and some gigabytes of space)

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [May 9, 2022, 4:07pm UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/13 "2022-05-09T16:07:34Z")

</div>

Hi @nisow95612 I think it is up to you... I have kind of lost track of exactly what you are trying to accomplish .. apologies...

Just for a test I picked a random index, reindex it and then force merged it to 1 segment the difference in size is ~.2% which is a reasonable margin of error to me. This has been my experience over time.. I am not sure what you are seeing.

My cluster is 7.17 even though those indices are 7.15.2

```auto
GET /_cat/indices/file*156*?v&s=pri.store.size:desc&bytes=b
health status index uuid pri rep docs.count docs.deleted store.size pri.store.size
green open filebeat-7.15.2-2022.05.04-000156 Asgg7scZT16ikU08ei075g 1 1 9943814 0 9146589863 4573198998
green open filebeat-7.15.2-2022.05.04-000156-reindex L2Nx5t3hTNqnCvWRFzVsUQ 1 1 9943814 0 9138208269 4569024766

```

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [May 10, 2022, 9:40am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/14 "2022-05-10T09:40:38Z")

</div>

Yes, I have made a mess in it, sorry.

Let me try to write it down properly.  
1)  
I was reindexing my data to change a mapping, observed +50% increase instead of expected +15%.  
By Disk usage API I found out that this increase is caused by \_source being +50% bigger.  
I assumed this must be due to index.codec, but it was BEST\_SPEED for both.  
Hence the title of this topic.  
2)  
Meanwhile original index was deleted, forcing me to debug different one.  
That one increased only +25% in size, but that is still way more than expected.  
To eliminate effect of mapping I reindexed again, copying mappings of "srcindex".  
Disk usage API still reported that the main culprit is +25% increase of \_source.

I documented this reindexing operation here:

> [@nisow95612](#):
>
> This is the actual (and in my opinion valid) comparision, using different index with similar data.  
> Plain reindex, no mapping changes, both indices forcemerged, still +25% size increase on \_source.
> 
> Reindexing command
> 
> ```auto
> POST /_reindex?slices=auto&pretty
> {
> "source": {
> "index": "srcindex"
> },
> "dest": {
> "index": "identmap"
> }
> }
> POST /identmap/_forcemerge?max_num_segments=1&pretty
> {
> "_shards" : {
> "total" : 6,
> "successful" : 6,
> "failed" : 0
> }
> }
> 
> ```
> 
> Index settings+mappings diff
> 
> ```auto
> @@ -1,5 +1,5 @@
> -GET /srcindex?pretty
> +GET /identmap?pretty
> {
> - "srcindex" : {
> + "identmap" : {
> "aliases" : { },
> "mappings" : {
> @@ -243,5 +243,5 @@
> "lifecycle" : {
> "name" : "SIEM-LOG",
> - "rollover_alias" : ""
> + "origination_date" : "1640390409363"
> },
> "routing" : {
> @@ -258,6 +258,6 @@
> "read" : "false"
> },
> - "provided_name" : "srcindex",
> - "creation_date" : "1640390409363",
> + "provided_name" : "identmap",
> + "creation_date" : "1651742088307",
> "unassigned" : {
> "node_left" : {
> @@ -275,7 +275,7 @@
> "priority" : "50",
> "number_of_replicas" : "1",
> - "uuid" : "xxxxxxxxxxxxxxxxxxxxxx",
> + "uuid" : "aaaaaaaaaaaaaaaaaaaaaa",
> "version" : {
> - "created" : "7150299"
> + "created" : "7160199"
> }
> }
> 
> ```
> 
> Disk usage
> 
> Comparison by field:
> 
> ```auto
> srcindex[MB] identmap[MB] increase [MB]
> _seq_no 2149 1587 -562
> <snip> -2 to +2
> timestamp 2899 2910 11
> _id 5505 5819 314
> _source 33729 42095 8366
> total 67077 75204 8127
> 
> ```
> 
> Comparison by type:
> 
> ```auto
> srcindex[MB] identmap[MB] increase [MB]
> points 5409 4974 -435
> doc_values 14817 14699 -118
> inverted_index 11845 11842 -3
> norms 0 0 0
> term_vectors 0 0 0
> stored_fields 35005 43688 8683
> total 67077 75204 8127
> 
> ```
> 
> \_cat segments
> 
> srcindex(original):
> 
> ```auto
> /_cat/segments/srcindex?h=index,shard,prirep,segment,generation,docs.count,docs.deleted,size,size.memory,committed,searchable,version,compound&v
> index shard prirep segment generation docs.count docs.deleted size size.memory committed searchable version compound
> srcindex 0 p _xxaa 341509 134835092 0 21.8gb 21404 true true 8.9.0 false
> srcindex 0 r _xxaa 341509 134835092 0 21.8gb 21404 true true 8.9.0 false
> srcindex 1 p _xxbb 342018 134812277 0 21.8gb 21404 true true 8.9.0 false
> srcindex 1 r _xxbb 342018 134812277 0 21.8gb 21404 true true 8.9.0 false
> srcindex 2 p _xxcc 341409 134841646 0 21.8gb 21404 true true 8.9.0 false
> srcindex 2 r _xxcc 341409 134841646 0 21.8gb 21404 true true 8.9.0 false
> 
> ```
> 
> identmap(reindexed)
> 
> ```auto
> /_cat/segments/identmap?h=index,shard,prirep,segment,generation,docs.count,docs.deleted,size,size.memory,committed,searchable,version,compound&v
> index shard prirep segment generation docs.count docs.deleted size size.memory committed searchable version compound
> identmap 0 p _aa 697 134835092 0 24.4gb 140556 true true 8.11.1 false
> identmap 0 r _aa 697 134835092 0 24.4gb 140556 true true 8.11.1 false
> identmap 1 p _bb 727 134812277 0 24.4gb 140556 true true 8.11.1 false
> identmap 1 r _bb 727 134812277 0 24.4gb 140556 true true 8.11.1 false
> identmap 2 p _cc 699 134841646 0 24.4gb 140556 true true 8.11.1 false
> identmap 2 r _cc 699 134841646 0 24.4gb 140556 true true 8.11.1 false
> 
> ```
> 
> I am sorry for all the previous unnecessary messages. I should have done all this triage before posting.

I am trying to figure out why all attempts of reindexing these indices consume +25% more space,  
fix the cause if possible, and then reindex to change mappings without significant comsumption of space.

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [May 10, 2022, 5:22pm UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/15 "2022-05-10T17:22:44Z")

</div>

@nisow95612 Not quite sure what to tell you.

I do see the lucene version is different, but I guess I wouldn't expect that to make the difference you're seeing.

Perhaps and apologies @DavidTurner to ask, perhaps David would have an idea why your indexes are bigger after reindex.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [May 10, 2022, 6:38pm UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/16 "2022-05-10T18:38:40Z")

</div>

It's hard to say without seeing the actual data, but one possible explanation is that the docs are being reordered during the reindex. `BEST_SPEED` means the docs are compressed in 60kiB blocks using LZ4 so the compression ratio is going to be very sensitive to how similar the docs in each block are.

The segments also come from different Lucene versions. I'm not familiar enough with Lucene development to know whether there might have been any regressions between these versions.

If you're in a position to share the actual data then I think we could dig deeper. If not then unfortunately we can only guess.

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [May 11, 2022, 5:30pm UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/17 "2022-05-11T17:30:56Z")

</div>

Thank you for the offer, I'll sift through the data and see if it could be shared.

* * *

Meanwhile I ran an experiment that might help:

1. Start with a smaller index that has the same problem - let's call it `ix`
2. Clone it - let's call the copy `ixclone`
3. Delete one document from `ixclone`
4. Forcemerge

```auto
GET /_cat/segments/ix,ixclone?v&h=index,shard,prirep,segment,generation,docs.count,docs.deleted,size,size.memory,committed,searchable,version,compound'
index shard prirep segment generation docs.count docs.deleted size size.memory committed searchable version compound
ix 0 p _21b6 95010 43741735 0 6.8gb 8708 true true 8.9.0 false
ix 1 p _27zu 103674 43743538 0 6.8gb 8716 true true 8.9.0 false
ix 2 p _207c 93576 43749154 0 6.8gb 8708 true true 8.9.0 false
ixclone 0 p _0 0 43741735 0 6.8gb 8708 true true 8.9.0 false
ixclone 1 p _0 0 43743538 0 6.8gb 8716 true true 8.9.0 false
ixclone 2 p _2 2 43749153 0 7.7gb 45684 true true 8.11.1 false
                                                                ^^^^^ INCREASE

```

1. Voilá - segment that held the deleted document is now \>10% bigger, while the rest stayed the same.

Does this, by a chance, prove reordering is not responsible for size increase, or "nope"?

> **Details**
>
> Both `ix` and `ixclone` actually have replicas. I removed them from the table to keep it short.
> 
> > **List of segments including replicas**
> >
> > ```auto
> > GET /_cat/segments/ix,ixclone?v&h=index,shard,prirep,segment,generation,docs.count,docs.deleted,size,size.memory,committed,searchable,version,compound'
> > index shard prirep segment generation docs.count docs.deleted size size.memory committed searchable version compound
> > ix 0 p _21b6 95010 43741735 0 6.8gb 8708 true true 8.9.0 false
> > ix 0 r _21b6 95010 43741735 0 6.8gb 8708 true true 8.9.0 false
> > ix 1 p _27zu 103674 43743538 0 6.8gb 8716 true true 8.9.0 false
> > ix 1 r _27zu 103674 43743538 0 6.8gb 8716 true true 8.9.0 false
> > ix 2 p _207c 93576 43749154 0 6.8gb 8708 true true 8.9.0 false
> > ix 2 r _207c 93576 43749154 0 6.8gb 8708 true true 8.9.0 false
> > ixclone 0 p _0 0 43741735 0 6.8gb 8708 true true 8.9.0 false
> > ixclone 0 r _0 0 43741735 0 6.8gb 8708 true true 8.9.0 false
> > ixclone 1 p _0 0 43743538 0 6.8gb 8716 true true 8.9.0 false
> > ixclone 1 r _0 0 43743538 0 6.8gb 8716 true true 8.9.0 false
> > ixclone 2 r _2 2 43749153 0 7.7gb 45684 true true 8.11.1 false
> > ixclone 2 p _2 2 43749153 0 7.7gb 45684 true true 8.11.1 false
> > 
> > ```
> 
> > **Shards before forcemerge**
> >
> > ```auto
> > GET /_cat/shards/ixclone?v=true&h=index,sh,pr,state,sc,docs,store
> > index sh pr state sc docs store
> > ixclone 1 p STARTED 1 43743538 6.8gb
> > ixclone 1 r STARTED 1 43743538 6.8gb
> > ixclone 2 p STARTED 2 43749153 6.8gb
> > ixclone 2 r STARTED 2 43749153 6.8gb
> > ixclone 0 p STARTED 1 43741735 6.8gb
> > ixclone 0 r STARTED 1 43741735 6.8gb
> > 
> > ```
> 
> > **Shards after forcemerge**
> >
> > ```auto
> > GET /_cat/shards/ixclone?v=true&h=index,sh,pr,state,sc,docs,store
> > index sh pr state sc docs store
> > ixclone 1 p STARTED 1 43743538 6.8gb
> > ixclone 1 r STARTED 1 43743538 6.8gb
> > ixclone 2 p STARTED 1 43749153 7.7gb
> > ixclone 2 r STARTED 1 43749153 7.7gb
> > ixclone 0 p STARTED 1 43741735 6.8gb
> > ixclone 0 r STARTED 1 43741735 6.8gb
> > 
> > ```

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [May 12, 2022, 6:06am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/18 "2022-05-12T06:06:19Z")

</div>

It does rather cast doubt on that idea, but I still don't think there's much we can do to help move this forward without seeing the data.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [May 12, 2022, 8:00am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/19 "2022-05-12T08:00:33Z")

</div>

A colleague has come up with a plausible explanation:

Lucene 8.7 changed some compression parameters to give smaller files at the expense of apparently-small performance drops (see [this blog post](https://www.elastic.co/blog/save-space-and-money-with-improved-storage-efficiency-in-elasticsearch-7-10)) but we later discovered some situations where the performance impact was not so small and partially reverted some of these changes in [LUCENE\_9917](https://issues.apache.org/jira/browse/LUCENE-9917) which landed in 8.10. Your original indices are written with the slower-but-more-highly-compressing Lucene 8.9 and the new ones are using the faster-but-less-compression Lucene 8.11.

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [May 12, 2022, 9:36am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/20 "2022-05-12T09:36:30Z")

</div>

That looks very plausible, thank you for figuring it out!

I guess there is nothing for me to do but to be happy that it is not +100% like with those nginx logs being compared in linked comment: [[LUCENE-9917] Reduce block size for BEST\_SPEED - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-9917?focusedCommentId=17403818#comment-17403818)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 9, 2022, 9:37am UTC](https://discuss.elastic.co/t/source-50-bigger-after-reindex/303805/21 "2022-06-09T09:37:17Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
