# Scan over 1Mio records get's slower and slower

**URL:** <https://discuss.elastic.co/t/scan-over-1mio-records-gets-slower-and-slower/5926>\
**Category:** Elasticsearch\
**Created:** [November 21, 2011, 4:06pm UTC](https://discuss.elastic.co/t/scan-over-1mio-records-gets-slower-and-slower/5926 "2011-11-21T16:06:26Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Chris\_3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chris_3/32/3070_2.png) [@Chris\_3](https://discuss.elastic.co/u/Chris_3)\
**Post date:** [November 21, 2011, 4:06pm UTC](https://discuss.elastic.co/t/scan-over-1mio-records-gets-slower-and-slower/5926/1 "2011-11-21T16:06:26Z")

</div>

Hi list,

just wondering if the scan search type is supposed to get slower when  
reading like 1000 times 5000 records, as this is what i'm seeing.  
The time needed to get the next resultset roughly doubles after every  
second resultset (and of course reaches the timeout before i get all  
documents).

I'm running ES 0.18.4 on a single node, 2 shards, no replicas, with  
around 1000 types inside a single index (each with 40 to 60 fields),  
10mio rows totalling in 40gb, while scanning by type.

I suspect it's my own fault, but a short yes/no (or some pointer)  
would help, thanks.

Greets, Chris

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 21, 2011, 6:42pm UTC](https://discuss.elastic.co/t/scan-over-1mio-records-gets-slower-and-slower/5926/2 "2011-11-21T18:42:38Z")

</div>

Its not your fault, it will take longer as you scan further into the  
resultset.

On Mon, Nov 21, 2011 at 6:06 PM, Chris [pc@matt-schwarz.com](mailto:pc@matt-schwarz.com) wrote:

> Hi list,
> 
> just wondering if the scan search type is supposed to get slower when  
> reading like 1000 times 5000 records, as this is what i'm seeing.  
> The time needed to get the next resultset roughly doubles after every  
> second resultset (and of course reaches the timeout before i get all  
> documents).
> 
> I'm running ES 0.18.4 on a single node, 2 shards, no replicas, with  
> around 1000 types inside a single index (each with 40 to 60 fields),  
> 10mio rows totalling in 40gb, while scanning by type.
> 
> I suspect it's my own fault, but a short yes/no (or some pointer)  
> would help, thanks.
> 
> Greets, Chris

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [November 22, 2011, 10:47am UTC](https://discuss.elastic.co/t/scan-over-1mio-records-gets-slower-and-slower/5926/3 "2011-11-22T10:47:58Z")

</div>

On Mon, 2011-11-21 at 20:42 +0200, Shay Banon wrote:

> Its not your fault, it will take longer as you scan further into the  
> resultset.

For a scrolled 'scan' search? I thought the point of a scan (ie not  
being sorted) was that it was an efficient way to retrieve all docs?

clint

> On Mon, Nov 21, 2011 at 6:06 PM, Chris [pc@matt-schwarz.com](mailto:pc@matt-schwarz.com) wrote:  
> Hi list,
> 
> ```
> just wondering if the scan search type is supposed to get
> slower when
> reading like 1000 times 5000 records, as this is what i'm
> seeing.
> The time needed to get the next resultset roughly doubles
> after every
> second resultset (and of course reaches the timeout before i
> get all
> documents).
>     
> I'm running ES 0.18.4 on a single node, 2 shards, no replicas,
> with
> around 1000 types inside a single index (each with 40 to 60
> fields),
> 10mio rows totalling in 40gb, while scanning by type.
>     
> I suspect it's my own fault, but a short yes/no (or some
> pointer)
> would help, thanks.
>     
> Greets, Chris
> 
> ```

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 22, 2011, 12:12pm UTC](https://discuss.elastic.co/t/scan-over-1mio-records-gets-slower-and-slower/5926/4 "2011-11-22T12:12:32Z")

</div>

It is efficient, certainly compared to when you do sorting, but, there is  
still an overhead as you scroll "deeper".

On Tue, Nov 22, 2011 at 12:47 PM, Clinton Gormley [clint@traveljury.com](mailto:clint@traveljury.com)wrote:

> On Mon, 2011-11-21 at 20:42 +0200, Shay Banon wrote:
> 
> > Its not your fault, it will take longer as you scan further into the  
> > resultset.
> 
> For a scrolled 'scan' search? I thought the point of a scan (ie not  
> being sorted) was that it was an efficient way to retrieve all docs?
> 
> clint
> 
> > On Mon, Nov 21, 2011 at 6:06 PM, Chris [pc@matt-schwarz.com](mailto:pc@matt-schwarz.com) wrote:  
> > Hi list,
> > 
> > ```
> > just wondering if the scan search type is supposed to get
> > slower when
> > reading like 1000 times 5000 records, as this is what i'm
> > seeing.
> > The time needed to get the next resultset roughly doubles
> > after every
> > second resultset (and of course reaches the timeout before i
> > get all
> > documents).
> > 
> > I'm running ES 0.18.4 on a single node, 2 shards, no replicas,
> > with
> > around 1000 types inside a single index (each with 40 to 60
> > fields),
> > 10mio rows totalling in 40gb, while scanning by type.
> > 
> > I suspect it's my own fault, but a short yes/no (or some
> > pointer)
> > would help, thanks.
> > 
> > Greets, Chris
> > 
> > ```

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [November 23, 2011, 8:04am UTC](https://discuss.elastic.co/t/scan-over-1mio-records-gets-slower-and-slower/5926/5 "2011-11-23T08:04:49Z")

</div>

On 22 Nov., 13:12, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> It is efficient, certainly compared to when you do sorting, but, there is  
> still an overhead as you scroll "deeper".

Yes, although I didn't have that feeling in my case with some million  
documents.

Peter.

BTW: there is a new search option available in the upcoming lucene.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 23, 2011, 10:13am UTC](https://discuss.elastic.co/t/scan-over-1mio-records-gets-slower-and-slower/5926/6 "2011-11-23T10:13:00Z")

</div>

On Wed, Nov 23, 2011 at 10:04 AM, Karussell [tableyourtime@googlemail.com](mailto:tableyourtime@googlemail.com)wrote:

> On 22 Nov., 13:12, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > It is efficient, certainly compared to when you do sorting, but, there is  
> > still an overhead as you scroll "deeper".
> 
> Yes, although I didn't have that feeling in my case with some million  
> documents.

Its not that bad, it simply does an early exit during the collection part  
once enough docs have been "collected". A regular search always goes  
through all of them (to sort things properly).

> Peter.
> 
> BTW: there is a new search option available in the upcoming lucene.

You mean the searchAfter one? It does something similar.

---

<div class="post-metadata">

**Author:** ![rmuir](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rmuir/32/44949_2.png) [@rmuir](https://discuss.elastic.co/u/rmuir)\
**Post date:** [November 23, 2011, 9:22pm UTC](https://discuss.elastic.co/t/scan-over-1mio-records-gets-slower-and-slower/5926/7 "2011-11-23T21:22:18Z")

</div>

this is not correct. searchAfter is not an optimization, and it doesn't  
early exit. it uses a fixed size priority queue and because of this, the  
100 millionth page takes the same time as the first.

but you must pass the last result (bottom result from the previous page) so  
that it knows which entries are 'too competitive' to enter the pq.

On Nov 23, 2011 5:13 AM, "Shay Banon" [kimchy@gmail.com](mailto:kimchy@gmail.com)

> You mean the searchAfter one? It does something similar.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 24, 2011, 2:04pm UTC](https://discuss.elastic.co/t/scan-over-1mio-records-gets-slower-and-slower/5926/8 "2011-11-24T14:04:37Z")

</div>

I meant in terms of cost.

On Wed, Nov 23, 2011 at 11:22 PM, Robert Muir [rcmuir@gmail.com](mailto:rcmuir@gmail.com) wrote:

> this is not correct. searchAfter is not an optimization, and it doesn't  
> early exit. it uses a fixed size priority queue and because of this, the  
> 100 millionth page takes the same time as the first.
> 
> but you must pass the last result (bottom result from the previous page)  
> so that it knows which entries are 'too competitive' to enter the pq.
> 
> On Nov 23, 2011 5:13 AM, "Shay Banon" [kimchy@gmail.com](mailto:kimchy@gmail.com)
> 
> > You mean the searchAfter one? It does something similar.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:47am UTC](https://discuss.elastic.co/t/scan-over-1mio-records-gets-slower-and-slower/5926/9 "2017-07-06T03:47:37Z")

</div>


