# Missing data from replica shards after delete by query and index

**URL:** <https://discuss.elastic.co/t/missing-data-from-replica-shards-after-delete-by-query-and-index/149402>\
**Category:** Elasticsearch\
**Created:** [September 21, 2018, 6:58am UTC](https://discuss.elastic.co/t/missing-data-from-replica-shards-after-delete-by-query-and-index/149402 "2018-09-21T06:58:57Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![ashokm](https://avatars.discourse-cdn.com/v4/letter/a/ecae2f/32.png) [@ashokm](https://discuss.elastic.co/u/ashokm)\
**Post date:** [September 21, 2018, 6:58am UTC](https://discuss.elastic.co/t/missing-data-from-replica-shards-after-delete-by-query-and-index/149402/1 "2018-09-21T06:58:57Z")

</div>

Hi All,

We are on ES 2.4.4. We are using delete by query to delete all docs of given type in an index using delete by query. Immediately, we index the data again. Sometimes, doc count on replica is less than primary. I am using "\_primary/\_replica" preference to find out the counts. If we delete the entire index and index the data gain, things are fine.

In Pre-Prod, we have 2 Node cluster and in Prod we have 6 Node cluster. Issue happens on both environments, Each index has 2 shards and 1 replica . Can you please suggest what could the root cause and how to either troubleshoot or fix the issue?

Thank you  
Ashok

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 21, 2018, 7:02am UTC](https://discuss.elastic.co/t/missing-data-from-replica-shards-after-delete-by-query-and-index/149402/2 "2018-09-21T07:02:30Z")

</div>

Have you waited for the operation to complete and run a refresh before getting the count?

---

<div class="post-metadata">

**Author:** ![ashokm](https://avatars.discourse-cdn.com/v4/letter/a/ecae2f/32.png) [@ashokm](https://discuss.elastic.co/u/ashokm)\
**Post date:** [September 21, 2018, 1:00pm UTC](https://discuss.elastic.co/t/missing-data-from-replica-shards-after-delete-by-query-and-index/149402/3 "2018-09-21T13:00:54Z")

</div>

Yes, we do wait and refresh. Here is the exact code we are using

```
					DeleteByQueryResponse rsp = new DeleteByQueryRequestBuilder(client, DeleteByQueryAction.INSTANCE)
													.setIndices(INDEX)
													.setTypes(TYPE)
													.setSource(new SearchSourceBuilder().query(QueryBuilders.matchAllQuery()).size(5000).toString())
													.execute()
													.actionGet();
				RefreshResponse refreshResponse = client.admin().indices().refresh(new RefreshRequest(INDEX)).actionGet();
```

---

<div class="post-metadata">

**Author:** ![ashokm](https://avatars.discourse-cdn.com/v4/letter/a/ecae2f/32.png) [@ashokm](https://discuss.elastic.co/u/ashokm)\
**Post date:** [September 21, 2018, 8:33pm UTC](https://discuss.elastic.co/t/missing-data-from-replica-shards-after-delete-by-query-and-index/149402/4 "2018-09-21T20:33:32Z")

</div>

Couple of observations

1. Instead of delete by query, I scroll through all docs and delete using bulk request and still same issue is seen
2. We are not relying on auto generated doc id
3. If we completely delete the index and re-index the whole data, there are no issues

Please suggest what could be going wrong for us? Thank you so much

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 22, 2018, 6:20am UTC](https://discuss.elastic.co/t/missing-data-from-replica-shards-after-delete-by-query-and-index/149402/5 "2018-09-22T06:20:54Z")

</div>

Elasticsearch 2.4 is quite old and a lot of effort has gone into improving resiliency and durability in later versions. If I recall correctly, replication in Elasticsearch 2.x was asynchronous, so could be more susceptible to network issues. Is your cluster deployed within a single DC with fast and reliable connections between the nodes?

To be sure that you are indeed waiting for the job to complete and replication of the changes to finish, can you run the steps manually (verifying that they all have completed before continuing) and verify you see the same problem then?

---

<div class="post-metadata">

**Author:** ![ashokm](https://avatars.discourse-cdn.com/v4/letter/a/ecae2f/32.png) [@ashokm](https://discuss.elastic.co/u/ashokm)\
**Post date:** [September 22, 2018, 6:43am UTC](https://discuss.elastic.co/t/missing-data-from-replica-shards-after-delete-by-query-and-index/149402/6 "2018-09-22T06:43:03Z")

</div>

Thank you Christian for you response.

All our machines are AWS EC2 instances in a single region but on different availability zones. As I mentioned above, we don't have this issue when index is deleted completely and indexed again. Issue only happens when we delete data for few types and added them back. So, I am assuming network connectivity is not an issue

I will explore the manual option.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 20, 2018, 6:43am UTC](https://discuss.elastic.co/t/missing-data-from-replica-shards-after-delete-by-query-and-index/149402/7 "2018-10-20T06:43:05Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
