# What is the best approach for streaming large number of docs from ES 7.1v scrolling vs slicing?

**URL:** <https://discuss.elastic.co/t/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scrolling-vs-slicing/210115>\
**Category:** Elasticsearch\
**Created:** [December 2, 2019, 6:18am UTC](https://discuss.elastic.co/t/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scrolling-vs-slicing/210115 "2019-12-02T06:18:02Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![yash.tandon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yash.tandon/32/53428_2.png) [@yash.tandon](https://discuss.elastic.co/u/yash.tandon)\
**Post date:** [December 2, 2019, 6:18am UTC](https://discuss.elastic.co/t/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scrolling-vs-slicing/210115/1 "2019-12-02T06:18:02Z")

</div>

> <https://stackoverflow.com/questions/59132907/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scroll>

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 2, 2019, 6:30am UTC](https://discuss.elastic.co/t/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scrolling-vs-slicing/210115/2 "2019-12-02T06:30:14Z")

</div>

Can you elaborate a bit more on the use case? How many documents are you retrieving? What do you do with this data? Are you indexing and/or updating concurrently? How frequently are you running this type of operation?

---

<div class="post-metadata">

**Author:** ![yash.tandon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yash.tandon/32/53428_2.png) [@yash.tandon](https://discuss.elastic.co/u/yash.tandon)\
**Post date:** [December 2, 2019, 6:38am UTC](https://discuss.elastic.co/t/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scrolling-vs-slicing/210115/3 "2019-12-02T06:38:16Z")

</div>

I am retrieving approx 1 million docs and currently i am doing nothing but just streaming the documents from ES, also I am getting variable number of docs each time i make a request and so i am not able to figure it out what i am doing wrong using the above code. Currently i am looking to get consistency in my results.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 2, 2019, 6:40am UTC](https://discuss.elastic.co/t/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scrolling-vs-slicing/210115/4 "2019-12-02T06:40:04Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> Are you indexing and/or updating concurrently? How frequently are you running this type of operation?

How are you intending to use this bulk retrieval?

The reason I am asking this is that Elasticsearch is a search engine optimized for fast retrieval of smaller result sets. Scroll is the right choice to return large data sets but as it locks segments in order to get a consistent view (impacts the ability to merge segments while a scroll is running) it is [as the docs describe](https://www.elastic.co/guide/en/elasticsearch/reference/7.1/search-request-scroll.html) not suitable for real-time user requests.

> Scrolling is not intended for real time user requests, but rather for processing large amounts of data, e.g. in order to reindex the contents of one index into a new index with a different configuration.

---

<div class="post-metadata">

**Author:** ![yash.tandon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yash.tandon/32/53428_2.png) [@yash.tandon](https://discuss.elastic.co/u/yash.tandon)\
**Post date:** [December 2, 2019, 6:48am UTC](https://discuss.elastic.co/t/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scrolling-vs-slicing/210115/5 "2019-12-02T06:48:55Z")

</div>

no i don't have issues with the real-time user requests according to my use case. Its just that i need to retrieve the existing docs in my cluster efficiently, without choking my cluster resources.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 2, 2019, 6:51am UTC](https://discuss.elastic.co/t/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scrolling-vs-slicing/210115/6 "2019-12-02T06:51:19Z")

</div>

How large are your documents? What is the specification of your cluster? How many shards are you actively fetching from?

---

<div class="post-metadata">

**Author:** ![yash.tandon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yash.tandon/32/53428_2.png) [@yash.tandon](https://discuss.elastic.co/u/yash.tandon)\
**Post date:** [December 2, 2019, 6:55am UTC](https://discuss.elastic.co/t/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scrolling-vs-slicing/210115/7 "2019-12-02T06:55:27Z")

</div>

actually i am using parent child relationship and on the basis of some conditions that i apply on my child docs, those document that get qualified, i am retrieving parent docs of those child docs by using has\_child query moreover my child docs are big(approx 40 fields) but parent docs are small(5-10).

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 2, 2019, 7:53am UTC](https://discuss.elastic.co/t/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scrolling-vs-slicing/210115/8 "2019-12-02T07:53:23Z")

</div>

How come you are using parent-child? Are the patent documents updated frequently?

Parent-child used more memory than flat documents but as I have never used them together with scroll queries I do not know how they interact.

---

<div class="post-metadata">

**Author:** ![yash.tandon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yash.tandon/32/53428_2.png) [@yash.tandon](https://discuss.elastic.co/u/yash.tandon)\
**Post date:** [December 2, 2019, 9:40am UTC](https://discuss.elastic.co/t/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scrolling-vs-slicing/210115/9 "2019-12-02T09:40:24Z")

</div>

no a document indexed once is never updated its just that the number of docs to be retrieved is big so i am issues related to heap and timeouts so i would request you to please help me with scrolling vs slicing.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 30, 2019, 9:40am UTC](https://discuss.elastic.co/t/what-is-the-best-approach-for-streaming-large-number-of-docs-from-es-7-1v-scrolling-vs-slicing/210115/10 "2019-12-30T09:40:24Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
