# Elasticsearch scroll offset

**URL:** <https://discuss.elastic.co/t/elasticsearch-scroll-offset/60255>\
**Category:** Elasticsearch\
**Created:** [September 12, 2016, 6:41am UTC](https://discuss.elastic.co/t/elasticsearch-scroll-offset/60255 "2016-09-12T06:41:31Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![dangnhdev](https://avatars.discourse-cdn.com/v4/letter/d/90ced4/32.png) [@dangnhdev](https://discuss.elastic.co/u/dangnhdev)\
**Post date:** [September 12, 2016, 6:41am UTC](https://discuss.elastic.co/t/elasticsearch-scroll-offset/60255/1 "2016-09-12T06:41:31Z")

</div>

Hello, can I scroll with a predefined offset in Elasticsearch 2.x? For example, I want to scroll from offset 101 limit 100. We want to scroll over an enormous dataset for analytic purpose. Our JVM's heap size can't store over 100 million document at a time.

Since we can't estimate the time for a scroll, use a old scroll id after a network failure event will make our services become a mess. Is there any gracefully way to handle this?

I found this closed issue : [Parallel/concurrent reads](https://github.com/elastic/elasticsearch/issues/13494#issuecomment-217368934)  
But **slice** is not available in Elasticsearch 2.x.

Additional info: We'are using Java driver.

---

<div class="post-metadata">

**Author:** ![jimczi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jimczi/32/47985_2.png) [@jimczi](https://discuss.elastic.co/u/jimczi)\
**Post date:** [September 13, 2016, 1:35pm UTC](https://discuss.elastic.co/t/elasticsearch-scroll-offset/60255/2 "2016-09-13T13:35:11Z")

</div>

If you are on ES 2.x one way to achieve this is to pick a field where you can perform multiple ranges that return the same amount of documents. A date field for instance with whom you can slice your data by doing a range query per month/year/day depending on the repartition of you data within the date field. This way you can split your scroll query in smaller queries that you can parallelize.

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [September 13, 2016, 2:32pm UTC](https://discuss.elastic.co/t/elasticsearch-scroll-offset/60255/3 "2016-09-13T14:32:14Z")

</div>

> [@dangnhdev](#):
>
> We want to scroll over an enormous dataset for analytic purpose.

This means you should use aggregations. Pulling the results back locally is almost always going to result in unacceptable performance even if you can get it working. You can totally write custom aggregations as plugins if you _really_ need to.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:20pm UTC](https://discuss.elastic.co/t/elasticsearch-scroll-offset/60255/4 "2017-07-05T22:20:39Z")

</div>


