# Impact of deletion on rewriting data

**URL:** <https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015>\
**Category:** Elasticsearch\
**Created:** [August 4, 2018, 11:28am UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015 "2018-08-04T11:28:25Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![sambodhi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sambodhi/32/33873_2.png) [@sambodhi](https://discuss.elastic.co/u/sambodhi)\
**Post date:** [August 4, 2018, 11:28am UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/1 "2018-08-04T11:28:25Z")

</div>

I am ingesting 40 GB data at a time after deleting the previous data for the same period. I have to delete before since new version might be missing some rows. Since update wouldn't remove he unwanted rows, am deleting before re-ingesting. Can deleting right before ingestion can performance issues in writing data?

Cluster details:  
7 node x r4.2xlarge (8 vCPU, 61 GB - 32 assigned to ES) on AWS  
ES 6.0 30GB per node (not 32GB)  
210GB memory / 5TB disk space  
Linux Red Hat 4.8.3-9 4.4.15-25.57.amzn1.x86\_64 Java 1.8

Index size: 1.7b documents / 1.8 TB / 28 Shards

Thanks

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 4, 2018, 1:59pm UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/2 "2018-08-04T13:59:17Z")

</div>

If you are removing data based on a given date, I'd suggest to use time based indices and just drop the unneeded indices.

---

<div class="post-metadata">

**Author:** ![sambodhi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sambodhi/32/33873_2.png) [@sambodhi](https://discuss.elastic.co/u/sambodhi)\
**Post date:** [August 4, 2018, 3:32pm UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/3 "2018-08-04T15:32:16Z")

</div>

Yes we plan to make that change. But we have something in production already, where we are facing slow writes and am guessing it is because of the deletion we do juts before. So I am looking for how to get this sorted for now and we will fix it permanently by having time based indices

---

<div class="post-metadata">

**Author:** ![sambodhi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sambodhi/32/33873_2.png) [@sambodhi](https://discuss.elastic.co/u/sambodhi)\
**Post date:** [August 4, 2018, 6:00pm UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/4 "2018-08-04T18:00:34Z")

</div>

@dadoonet thanks for your reply. Can having multiple indices impact on performance? For example currently we have around 2 TB data in 28 primary shards. If we divide indices by week (lets day) and we have 52 indices. so querying 6 months of data, it will query 26\*5 = 130 shards or may be it is better to have lesser shards per index.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 4, 2018, 9:44pm UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/5 "2018-08-04T21:44:10Z")

</div>

Why would you keep so many shards?

---

<div class="post-metadata">

**Author:** ![sambodhi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sambodhi/32/33873_2.png) [@sambodhi](https://discuss.elastic.co/u/sambodhi)\
**Post date:** [August 5, 2018, 8:51am UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/6 "2018-08-05T08:51:23Z")

</div>

We started with 7 shards but our data was growing fast and we reached like 120GB per shard. We started facing problems with it with ingestion and slow read performance when cluster is relocating etc. To keep 50GB/shard, we chose 28 shards.

Ok, I realised we can't easily split by date because we do parent/child to do absolute distinct (no approximations but accurate) which normally ES would not allow since it uses hyperloglog with cardinality 40,000. Probably, ES was a wrong choice for this. a) there was no way to do absolute distinct for higher cardinality, we worked around with parent/child but its not great b) Because of this we cannot even split the index.

So with current problem of deleting and re-ingesting, I see the difference in indexing speed when 1) Just ingest data 2) I delete and ingest. In the graph below index rate is consistent in case 1 and slows down in 2

 ![45](https://us1.discourse-cdn.com/elastic/original/3X/0/f/0fd4d3a3527972721a44c7fa039b57eff97a5562.png)

Can you please help me understand what causes this and if there can anything done to sort this for now (for example, may be deleting the data in advance)? Thank you

---

<div class="post-metadata">

**Author:** ![sambodhi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sambodhi/32/33873_2.png) [@sambodhi](https://discuss.elastic.co/u/sambodhi)\
**Post date:** [August 5, 2018, 11:46am UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/7 "2018-08-05T11:46:40Z")

</div>

![47](https://us1.discourse-cdn.com/elastic/original/3X/4/c/4cdf07ee9228b46b86281463f7ea0292bae463d1.png)

peaks before red line is 1 and after is 2  
even after deletion has finished, ingestion is slow

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 6, 2018, 9:07am UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/8 "2018-08-06T09:07:42Z")

</div>

Did you run a force merge after delete operation?

---

<div class="post-metadata">

**Author:** ![sambodhi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sambodhi/32/33873_2.png) [@sambodhi](https://discuss.elastic.co/u/sambodhi)\
**Post date:** [August 6, 2018, 10:05am UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/9 "2018-08-06T10:05:15Z")

</div>

ah no we didn't force merge. Thanks for pointing! Probably that leads to this overlap between deletion and ingestion. probably we should call this api [https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-forcemerge.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-forcemerge.html) after after deletion and before re-ingesting new data? I am guessing could be a heavy operation to do over 2-3 TB data.

Also we run deletion and ingestion with refresh\_interval = -1. Does that needs to be reset as well or force merge would be sufficient?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 6, 2018, 12:02pm UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/10 "2018-08-06T12:02:08Z")

</div>

Use [Force merge API | Elasticsearch Guide [8.11] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-forcemerge.html) with `only_expunge_deletes` so you will "just" remove documents that needs to be erased.

I'm not sure if this will help or not, but I'd give it a try.

> Also we run deletion and ingestion with refresh\_interval = -1. Does that needs to be reset as well or force merge would be sufficient?

When do you refresh the index then? Are calling refresh manually?

---

<div class="post-metadata">

**Author:** ![sambodhi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sambodhi/32/33873_2.png) [@sambodhi](https://discuss.elastic.co/u/sambodhi)\
**Post date:** [August 7, 2018, 2:23pm UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/11 "2018-08-07T14:23:03Z")

</div>

We have 4 ingestion jobs running weekly across 2 indices. We set refresh\_interval=-1 before these set of jobs runs and reset it to refresh\_interval=1s at the end.

Actually we have tried force merge before (with only\_expunge\_deletes) when our cluster went into yellow state (that time we had very big shards, holding around 220GB data each so that could be a reason), we ended up reindexing to a new index with took around 2 days. From that experience we realised this is an heavy operation and little concerned to do it every week.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 7, 2018, 5:07pm UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/12 "2018-08-07T17:07:27Z")

</div>

Only "good" solution IMO is still:

> If you are removing data based on a given date, I'd suggest to use time based indices and just drop the unneeded indices.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 4, 2018, 5:07pm UTC](https://discuss.elastic.co/t/impact-of-deletion-on-rewriting-data/143015/13 "2018-09-04T17:07:29Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
