# How to scroll through an Elasticsearch index using elasticsearch-spark?

**URL:** <https://discuss.elastic.co/t/how-to-scroll-through-an-elasticsearch-index-using-elasticsearch-spark/144618>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [August 16, 2018, 4:55am UTC](https://discuss.elastic.co/t/how-to-scroll-through-an-elasticsearch-index-using-elasticsearch-spark/144618 "2018-08-16T04:55:40Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Swapnil\_Nawale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/swapnil_nawale/32/23260_2.png) [@Swapnil\_Nawale](https://discuss.elastic.co/u/Swapnil_Nawale)\
**Post date:** [August 16, 2018, 4:55am UTC](https://discuss.elastic.co/t/how-to-scroll-through-an-elasticsearch-index-using-elasticsearch-spark/144618/1 "2018-08-16T04:55:40Z")

</div>

With the Java `Client.prepareSearch()` and `Client.prepareSearchScroll()` APIs, we can query an Elasticsearch index using the scrolls as mentioned in the [documentation](https://www.elastic.co/guide/en/elasticsearch/client/java-api/5.6/java-search-scrolling.html). With these APIs, we can select only a specific number of hits per request by setting `SearchRequestBuilder.setSize()` . The `SearchResponse` provides the scroll Id, which is then used in the subsequent request.

How can one use elasticsearch-spark to implement a similar functionality ? All `JavaEsSpark.esRDD()` methods return `JavaPairRDD` , which would contain all hits. Is there a way to request only a specific number of hits per request and then continue scrolling with further request?

I found the configuration `es.scroll.size` , which seems equivalent to `SearchRequestBuilder.setSize()` but I am not sure how to use it and how the scroll ids would be used in the context of elasticsearch-spark?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [August 17, 2018, 8:05pm UTC](https://discuss.elastic.co/t/how-to-scroll-through-an-elasticsearch-index-using-elasticsearch-spark/144618/2 "2018-08-17T20:05:01Z")

</div>

ES-Hadoop uses the scroll endpoint to collect all the data for processing within Spark. ES-Hadoop performs the multiple scroll requests under the hood on its own, requesting the next scroll entry after the data in the current scroll response is fully consumed. I'm not sure I understand what you're looking for in terms of advancing the scroll request on your own. Could you elaborate on your use case?

---

<div class="post-metadata">

**Author:** ![Swapnil\_Nawale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/swapnil_nawale/32/23260_2.png) [@Swapnil\_Nawale](https://discuss.elastic.co/u/Swapnil_Nawale)\
**Post date:** [August 21, 2018, 5:48pm UTC](https://discuss.elastic.co/t/how-to-scroll-through-an-elasticsearch-index-using-elasticsearch-spark/144618/3 "2018-08-21T17:48:17Z")

</div>

Thanks @james.baiera for your reply. The Elasticsearch documents in our indices have a field that stores a list of objects. I want to be able to fetch the documents from ES using a query, access the list of objects from above-mentioned field. And then apply certain filters on that list of objects. I also want to stop the execution once I get intended number of objects from the fetched documents. Because of certain restrictions related data storage formats, I can't apply the filters on the ES side.

For this, I could use `Client.prepareSearch()` and `Client.prepareSearchScroll()` APIs to fetch a specific number of documents in memory per scroll and apply filters on the list field. I can continue scrolling until the expected number of objects are extracted.

I am not sure how a similar functionality can be implemented with elasticsearch-spark ?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [August 21, 2018, 6:00pm UTC](https://discuss.elastic.co/t/how-to-scroll-through-an-elasticsearch-index-using-elasticsearch-spark/144618/4 "2018-08-21T18:00:36Z")

</div>

You could always access the documents from Elasticsearch, then extract the list items into their own records with a `flatmap` operation. Then once your filters are applied to the RDD/Dataset, tell Spark that you want to take X number of results from the RDD.

If you have a substantial number of records to go through, asking Spark to take a number of results may still compute the full result set. You can try to limit the total number of records returned from ES-hadoop to Spark by setting `es.scroll.limit`, which will tell the connector to preemptively stop reading results from a scroll request once it has read that many documents. You might need to play with the value you set in there so that you always get your minimum number of documents.

---

<div class="post-metadata">

**Author:** ![Swapnil\_Nawale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/swapnil_nawale/32/23260_2.png) [@Swapnil\_Nawale](https://discuss.elastic.co/u/Swapnil_Nawale)\
**Post date:** [August 21, 2018, 6:24pm UTC](https://discuss.elastic.co/t/how-to-scroll-through-an-elasticsearch-index-using-elasticsearch-spark/144618/5 "2018-08-21T18:24:32Z")

</div>

Thank you. I will experiment with the `es.scroll.limit` config.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 18, 2018, 6:24pm UTC](https://discuss.elastic.co/t/how-to-scroll-through-an-elasticsearch-index-using-elasticsearch-spark/144618/6 "2018-09-18T18:24:38Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
