# Scroll time increment effect on Elastic Search

**URL:** <https://discuss.elastic.co/t/scroll-time-increment-effect-on-elastic-search/150178>\
**Category:** Elasticsearch\
**Created:** [September 27, 2018, 10:51am UTC](https://discuss.elastic.co/t/scroll-time-increment-effect-on-elastic-search/150178 "2018-09-27T10:51:55Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![AMIT\_KEWAL](https://avatars.discourse-cdn.com/v4/letter/a/ce73a5/32.png) [@AMIT\_KEWAL](https://discuss.elastic.co/u/AMIT_KEWAL)\
**Post date:** [September 27, 2018, 10:51am UTC](https://discuss.elastic.co/t/scroll-time-increment-effect-on-elastic-search/150178/1 "2018-09-27T10:51:56Z")

</div>

I am working on a project using ElasticSearch and querying it to fetch the member information. It has **~30 Lakhs records**.

Basically, I am running a campaign for 20L users and the user data is present on **elasticsearch6.2**. I query the ES and fetches the records in batches(50 records at a time) using the **scroll**. Also, I want to keep the **SEARCH context for 1 day** because if the campaign running process fails due to any reason, I can resume the campaign from where it was stopped. In this way, I will escape from starting the campaign again from starting. I am also saving the **scrollID** and will use it to resume campaign.

While testing I found CPU Utilization increased by 50% (ES config: 2 nodes with 4 shards running on aws, Instance Type: **i3.xlarge.elasticsearch** ) and its CPU Utilization remains consistent to 50%.

Is there any relation between CPU Utilization and keeping the search context for 1day. BTW campaigns take 6 hours to finish.

---

<div class="post-metadata">

**Author:** ![Bernt\_Rostad](https://avatars.discourse-cdn.com/v4/letter/b/3ab097/32.png) [@Bernt\_Rostad](https://discuss.elastic.co/u/Bernt_Rostad)\
**Post date:** [September 27, 2018, 11:13am UTC](https://discuss.elastic.co/t/scroll-time-increment-effect-on-elastic-search/150178/2 "2018-09-27T11:13:17Z")

</div>

I've never used a Scroll with more than 1 minute timeout and there is probably a price to pay for keeping the search context open for much longer. The official [Scroll](https://www.elastic.co/guide/en/elasticsearch/reference/6.4/search-request-scroll.html#scroll-search-context) documentations warns that:

`an open search context prevents the old segments from being deleted while they are still in use. [...] Keeping older segments alive means that more file handles are needed.`

But I'm not sure if this explains the increased CPU you're seeing.

An alternative to using Scroll is the light weight [Search After](https://www.elastic.co/guide/en/elasticsearch/reference/6.2/search-request-search-after.html) mechanism, which is very useful if you can order your records on a unique field - for instance a sequence number or a date timestamp. With this mechanism you're not keeping an expensive open state in Elasticsearch, because each search request knows where to start (after). Hence you can perform one search today and the next in a week, with no extra cost.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 25, 2018, 11:13am UTC](https://discuss.elastic.co/t/scroll-time-increment-effect-on-elastic-search/150178/3 "2018-10-25T11:13:24Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
