# Scan/scrolling: many small scrolls, or one large scroll?

**URL:** <https://discuss.elastic.co/t/scan-scrolling-many-small-scrolls-or-one-large-scroll/6729>\
**Category:** Elasticsearch\
**Created:** [February 16, 2012, 9:08am UTC](https://discuss.elastic.co/t/scan-scrolling-many-small-scrolls-or-one-large-scroll/6729 "2012-02-16T09:08:37Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Matt1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matt1/32/2986_2.png) [@Matt1](https://discuss.elastic.co/u/Matt1)\
**Post date:** [February 16, 2012, 9:08am UTC](https://discuss.elastic.co/t/scan-scrolling-many-small-scrolls-or-one-large-scroll/6729/1 "2012-02-16T09:08:37Z")

</div>

I was wondering, from a performance perspective (more specifically,  
cranking through the data as quickly as possible), which one is better if I  
wanted to scroll through a large-ish (hundred of gigabytes to a few  
terabytes) index with an ordered field (e.g. all docs have a date field):

1. Do many small scrolls, one each for each non-overlapping interval  
(again, for example, if I know beforehand that there's only 1 year of data,  
then 12 scrolls, one for each month)
2. Go ahead and just do a normal scroll with match\_all or something similar.

The reason I axk is because it was mentioned in previous posts that the  
deeper you go into a scoll, the slower it gets. Would this technique  
alleviate that? What are the tradeoffs? Also, would the answer change if it  
was a single node cluster versus a multi-node cluster?

Thanks in advance!

Matt

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 16, 2012, 7:51pm UTC](https://discuss.elastic.co/t/scan-scrolling-many-small-scrolls-or-one-large-scroll/6729/2 "2012-02-16T19:51:09Z")

</div>

First, the performance problem with "deeper" scrolling has been fixed (or greatly improved). In general, the benefit of what you suggest comes from the fact that you do things in parallel, so if you handle it on the client side in parallel as well (multiple processes / threads / machines), then do it.

On Thursday, February 16, 2012 at 11:08 AM, Matt wrote:

> I was wondering, from a performance perspective (more specifically, cranking through the data as quickly as possible), which one is better if I wanted to scroll through a large-ish (hundred of gigabytes to a few terabytes) index with an ordered field (e.g. all docs have a date field):
> 
> 1. Do many small scrolls, one each for each non-overlapping interval (again, for example, if I know beforehand that there's only 1 year of data, then 12 scrolls, one for each month)
> 2. Go ahead and just do a normal scroll with match\_all or something similar.
> 
> The reason I axk is because it was mentioned in previous posts that the deeper you go into a scoll, the slower it gets. Would this technique alleviate that? What are the tradeoffs? Also, would the answer change if it was a single node cluster versus a multi-node cluster?
> 
> Thanks in advance!
> 
> Matt

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [February 16, 2012, 8:01pm UTC](https://discuss.elastic.co/t/scan-scrolling-many-small-scrolls-or-one-large-scroll/6729/3 "2012-02-16T20:01:00Z")

</div>

On Thu, 2012-02-16 at 21:51 +0200, Shay Banon wrote:

> First, the performance problem with "deeper" scrolling has been fixed  
> (or greatly improved). In general, the benefit of what you suggest  
> comes from the fact that you do things in parallel, so if you handle  
> it on the client side in parallel as well (multiple processes /  
> threads / machines), then do it.

This performance improvement also applies to non scan requests?

clint

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 16, 2012, 8:21pm UTC](https://discuss.elastic.co/t/scan-scrolling-many-small-scrolls-or-one-large-scroll/6729/4 "2012-02-16T20:21:44Z")

</div>

No, just for the scan type.

On Thursday, February 16, 2012 at 10:01 PM, Clinton Gormley wrote:

> On Thu, 2012-02-16 at 21:51 +0200, Shay Banon wrote:
> 
> > First, the performance problem with "deeper" scrolling has been fixed  
> > (or greatly improved). In general, the benefit of what you suggest  
> > comes from the fact that you do things in parallel, so if you handle  
> > it on the client side in parallel as well (multiple processes /  
> > threads / machines), then do it.
> 
> This performance improvement also applies to non scan requests?
> 
> clint

---

<div class="post-metadata">

**Author:** ![Matt1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matt1/32/2986_2.png) [@Matt1](https://discuss.elastic.co/u/Matt1)\
**Post date:** [February 17, 2012, 3:53am UTC](https://discuss.elastic.co/t/scan-scrolling-many-small-scrolls-or-one-large-scroll/6729/5 "2012-02-17T03:53:41Z")

</div>

Got it, thanks for the response. I see the enhancement lined up for the  
proper 0.19 release, will give it a try when it's out. Cheers!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:38am UTC](https://discuss.elastic.co/t/scan-scrolling-many-small-scrolls-or-one-large-scroll/6729/6 "2017-07-06T03:38:58Z")

</div>


