# Query Millions of records in Elasticsearch

**URL:** <https://discuss.elastic.co/t/query-millions-of-records-in-elasticsearch/21145>\
**Category:** Elasticsearch\
**Created:** [December 8, 2014, 2:11pm UTC](https://discuss.elastic.co/t/query-millions-of-records-in-elasticsearch/21145 "2014-12-08T14:11:19Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![sushmitha](https://avatars.discourse-cdn.com/v4/letter/s/a9adbd/32.png) [@sushmitha](https://discuss.elastic.co/u/sushmitha)\
**Post date:** [December 8, 2014, 2:11pm UTC](https://discuss.elastic.co/t/query-millions-of-records-in-elasticsearch/21145/1 "2014-12-08T14:11:19Z")

</div>

Hi,

I have an index with 6 Crores of records. My usecase is to read the  
entire index, check each record, whether it is present in new index or  
not.If not I have to index into new index. I used scan and scroll operation  
to read the index using JAVA Api. But this process is taking lot of time  
i.e., to process 50,000 rcds it is taking 8 min. Can anyone suggest me how  
I can configure or change my queries.

Thanks in advance.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/4014e3f0-2ce6-48d9-afdd-e438857e85f0%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/4014e3f0-2ce6-48d9-afdd-e438857e85f0%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [December 8, 2014, 2:26pm UTC](https://discuss.elastic.co/t/query-millions-of-records-in-elasticsearch/21145/2 "2014-12-08T14:26:36Z")

</div>

On Mon, Dec 8, 2014 at 9:11 AM, Sushmitha Chakka \<  
[sushmitha@sigmoidanalytics.com](mailto:sushmitha@sigmoidanalytics.com)\>

> Hi,
> 
> I have an index with 6 Crores of records. My usecase is to read the  
> entire index, check each record, whether it is present in new index or  
> not.If not I have to index into new index. I used scan and scroll operation  
> to read the index using JAVA Api. But this process is taking lot of time  
> i.e., to process 50,000 rcds it is taking 8 min. Can anyone suggest me how  
> I can configure or change my queries.

I know from experience that scan/scroll can handle batch sizes in the low  
thousands without trouble so you should give that a shot. Each scroll call  
should be quite quick. It might be a good idea to post a JSON recreation  
of your problem so we can see what is happening. Usually the slow part of  
the scan/scroll into new index is the batch calls to add the documents into  
the new index. And whether or not that is "slow" is really dependant on  
the size of the documents, the complexity of any scripts you use on import,  
your disk speed, the complexity of your analysis, your cpu speed, the merge  
settings you use. That list is roughly in order of how likely I've seen  
things effect import speed.

Nik

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAPmjWd0twuaDy\_qL-Gjv1Bv2qpee%2BREY14s8T2%2BUbgDgsp8kSw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAPmjWd0twuaDy_qL-Gjv1Bv2qpee%2BREY14s8T2%2BUbgDgsp8kSw%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:44am UTC](https://discuss.elastic.co/t/query-millions-of-records-in-elasticsearch/21145/3 "2017-07-06T00:44:59Z")

</div>


