# SCROLL internals

**URL:** <https://discuss.elastic.co/t/scroll-internals/72173>\
**Category:** Elasticsearch\
**Created:** [January 19, 2017, 1:50pm UTC](https://discuss.elastic.co/t/scroll-internals/72173 "2017-01-19T13:50:36Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![kasia](https://avatars.discourse-cdn.com/v4/letter/k/f07891/32.png) [@kasia](https://discuss.elastic.co/u/kasia)\
**Post date:** [January 19, 2017, 1:50pm UTC](https://discuss.elastic.co/t/scroll-internals/72173/1 "2017-01-19T13:50:36Z")

</div>

Hi,  
The question is about scrolling in ES v.5.1 cluster architecture.

In the docs, there's a nice explanation of the deep pagination problem with from/size:  
you know from that what the coordinating node and each shard computes whenever a new page is being requested.

Could someone explain how scrolling works at the same level of abstraction?  
For instance:

1. Which node does the search context reside on in the cluster?
2. What does the search context contain after the initial search? (all documents matching the query? some of them? any reference to the segments cotaining the documents matching the query? something else?)
3. What do particular shards do when the initial search is launched?
4. What's the responsability of the coordination node when the initial search is launched?
5. What do particular shards do when the scroll API to get the next bunch of data is lauched?
6. What's the responsability of the coordination node when the scroll API to get the next bunch of data is lauched?
7. Does scroll support sorting in 5.1 (only sorting by \_doc is mentioned as the most efficient option)?
8. How does the scroll relates to query\_then\_fetch | dfs\_query\_then\_fetch search types? In 2.x the search\_type had to be SCAN, right?
9. Can scrolling be combined with routing?
10. What is the sliced scroll for? Any use case where it could be useful would be appreciated.
11. Finally, is it for maintenance purposes exclusively?

Thanks in advance,  
Kasia

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [January 19, 2017, 2:35pm UTC](https://discuss.elastic.co/t/scroll-internals/72173/2 "2017-01-19T14:35:29Z")

</div>

> [@kasia](#):
>
> Which node does the search context reside on in the cluster?

For scrolling, the search context is not used. It's the scroll context. The scroll context resides on all shards involved in the index.

> [@kasia](#):
>
> What does the search context contain after the initial search? (all documents matching the query? some of them? any reference to the segments cotaining the documents matching the query? something else?)

The scroll context contains the total hits, the max score, and last emitted doc number (of the Lucene shard), and the keep-alive time for timing out the scroll context.

> [@kasia](#):
>
> What do particular shards do when the initial search is launched?

Nothing special. The scroll context is opened.

> [@kasia](#):
>
> What's the responsability of the coordination node when the initial search is launched?

The coordination node has no responsibility. It works like in a usual search operation.

> [@kasia](#):
>
> What do particular shards do when the scroll API to get the next bunch of data is lauched?

The shards deliver the next documents, like in a usual search operation, continuing from the last emitted doc.

> [@kasia](#):
>
> What's the responsability of the coordination node when the scroll API to get the next bunch of data is lauched?

See 4. There is nothing special regarding scroll.

> [@kasia](#):
>
> Does scroll support sorting in 5.1 (only sorting by \_doc is mentioned as the most efficient option)?

Yes.

> [@kasia](#):
>
> How does the scroll relates to query\_then\_fetch | dfs\_query\_then\_fetch search types? In 2.x the search\_type had to be SCAN, right?

SCAN is gone in 5.x.

`query_then_fetch` and `dfs_query_then_fetch` are modes for the search operation how to compute relevance scoring on the shards and how to retrieve the computed documents. This is not specific to scroll, it's a mode that can be set for all kinds of search operations.

> [@kasia](#):
>
> Can scrolling be combined with routing?

Document routing allows to index a document on a certain shard. This has no relationship with scroll search, so it can be "combined" (indexing and search can not be combined in a single operation).

> [@kasia](#):
>
> What is the sliced scroll for? Any use case where it could be useful would be appreciated.

It's for concurrent scrolling, which can be executed with higher performance.

> **[Request body search | Elasticsearch Guide \[8.11\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-body.html#request-body-search-scroll)**

Imagine a 10 million doc index and 5 shards on one server node. Then consider a thread pool on the server. To iterate over the index in the naive way, each shard uses one thread with a total of 5 threads. Each thread would have to traverse 2 million docs. Then imagine a sliced scroll with max = 2. Each thread on a shard would have to iterate over only 1 million documents, with a total of 10 threads. If you can run two scroll queries on client side, one for each slice, and join the result set on client side, you can finish the total scroll in approximately half of the time.

> [@kasia](#):
>
> Finally, is it for maintenance purposes exclusively?

Nope. It's for retrieving large result sets, for whatever purpose.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 16, 2017, 2:35pm UTC](https://discuss.elastic.co/t/scroll-internals/72173/3 "2017-02-16T14:35:59Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
