# Issues with scan and scroll as well as count API

**URL:** <https://discuss.elastic.co/t/issues-with-scan-and-scroll-as-well-as-count-api/29034>\
**Category:** Elasticsearch\
**Created:** [September 10, 2015, 1:18pm UTC](https://discuss.elastic.co/t/issues-with-scan-and-scroll-as-well-as-count-api/29034 "2015-09-10T13:18:09Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![d\_bharvi](https://avatars.discourse-cdn.com/v4/letter/d/7feea3/32.png) [@d\_bharvi](https://discuss.elastic.co/u/d_bharvi)\
**Post date:** [September 10, 2015, 1:18pm UTC](https://discuss.elastic.co/t/issues-with-scan-and-scroll-as-well-as-count-api/29034/1 "2015-09-10T13:18:09Z")

</div>

Hi,  
I am using scan\_scroll API for data re-indexing using python client. The total data is of 90 GB which contains 40 Million documents. Since it is query based re-indexing, i usually get less than 10000 documents per query. Below are the index and machine configurations.  
Elasticsearch version : 1.4.2  
No. of primary shards: 8  
No. of replica shards: 8  
No of total segments: 16  
There re two data nodes with 26 GB of RAM and 8 core CPU each. 3 master and 1 client nodes also exist in the cluster.

My problem is scan\_scroll API is not consistent at all. on 20% of the time it does not give me the complete data for the same query. The same thing happens with the \_count API too. Hitting the same query to get the count of data returns different results many a time.

Have anyone faced this issue?

Please let me know if someone can help.

Regards,  
Bharvi

---

<div class="post-metadata">

**Author:** ![javanna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javanna/32/4698_2.png) [@javanna](https://discuss.elastic.co/u/javanna)\
**Post date:** [September 10, 2015, 3:58pm UTC](https://discuss.elastic.co/t/issues-with-scan-and-scroll-as-well-as-count-api/29034/2 "2015-09-10T15:58:14Z")

</div>

Hi Bharvi,  
that makes me think that maybe some replicas got out of sync with their primaries at some point. What if you use the search api, with search\_type count and you specify the preference, so that you always hit the same shards? Does the number of result change every time again?

Also I'm assuming that you haven't been indexing while querying at the moment (although not a problem with the scan/scroll).

---

<div class="post-metadata">

**Author:** ![d\_bharvi](https://avatars.discourse-cdn.com/v4/letter/d/7feea3/32.png) [@d\_bharvi](https://discuss.elastic.co/u/d_bharvi)\
**Post date:** [September 10, 2015, 6:21pm UTC](https://discuss.elastic.co/t/issues-with-scan-and-scroll-as-well-as-count-api/29034/3 "2015-09-10T18:21:00Z")

</div>

Thanks Luca for responding. I haven't tried searching with preference yet.  
I will try it soon.  
And there is no indexing going on while search. Its complete static data. I  
had indexed it using only primary shard and replicated it on other node  
after complete indexing is done.  
Earlier there were 334 segments in that index. But after optimizing there  
are one segment per shard. Still no luck.

But using preference parameter is very valid point. I will let you know the  
results. Thanks again for pointinh out.

---

<div class="post-metadata">

**Author:** ![javanna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javanna/32/4698_2.png) [@javanna](https://discuss.elastic.co/u/javanna)\
**Post date:** [September 10, 2015, 6:27pm UTC](https://discuss.elastic.co/t/issues-with-scan-and-scroll-as-well-as-count-api/29034/4 "2015-09-10T18:27:07Z")

</div>

You can also check how many docs you have on each shard using indices stats api e.g. to figure out if some replicas are out of sync.

---

<div class="post-metadata">

**Author:** ![d\_bharvi](https://avatars.discourse-cdn.com/v4/letter/d/7feea3/32.png) [@d\_bharvi](https://discuss.elastic.co/u/d_bharvi)\
**Post date:** [September 11, 2015, 5:44am UTC](https://discuss.elastic.co/t/issues-with-scan-and-scroll-as-well-as-count-api/29034/5 "2015-09-11T05:44:55Z")

</div>

Hi Luca,

Tried Everything. Got to know that my documents have not been distributed  
equally but there is no replication issue at all.  
Here is the document distribution for the index:  
Shard 0: 23596919  
Shard 1: 23597019  
Shard 2: 23593214  
Shard 3: 23598522  
Shard 4: 15684207  
Shard 5: 7294415  
Shard 6: 7293062  
Shard 7: 7294274

I tried all the parameters of preference : nodes, shards, primary .. But no  
luck. There are enough resources in the cluster.  
But, the results are still inconsistent.

## Regards

_Bharvi Dixit_  
Software Engineer  
596, Udyog Vihar Phase V, Sector 19, Gurgaon 122016, India  
_Tel:_ +91 (124) 438 4534 _Web:_ [www.grownout.com](http://www.grownout.com)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:50pm UTC](https://discuss.elastic.co/t/issues-with-scan-and-scroll-as-well-as-count-api/29034/6 "2017-07-05T23:50:57Z")

</div>


