# Correct setting of "es.scroll.size" with for optimal Spark read performance

**URL:** <https://discuss.elastic.co/t/correct-setting-of-es-scroll-size-with-for-optimal-spark-read-performance/91077>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [June 28, 2017, 10:07am UTC](https://discuss.elastic.co/t/correct-setting-of-es-scroll-size-with-for-optimal-spark-read-performance/91077 "2017-06-28T10:07:42Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ntim](https://avatars.discourse-cdn.com/v4/letter/n/d9b06d/32.png) [@ntim](https://discuss.elastic.co/u/ntim)\
**Post date:** [June 28, 2017, 10:07am UTC](https://discuss.elastic.co/t/correct-setting-of-es-scroll-size-with-for-optimal-spark-read-performance/91077/1 "2017-06-28T10:07:43Z")

</div>

Hi,

I am using ElasticSearch 5.4.1, Spark 2.1.1 and Elasticsearch Hadoop 5.4.2. I have three nodes with ~250 indices with ~750 shards (no replication) containing ~1.7 billion documents.

If I try to query one day worth of data (~5m documents) in pyspark

```
docs = spark.read\
    .format('org.elasticsearch.spark.sql')\
    .option('pushdown', True)\
    .option('double.filtering', True)\
    .load('index-name-*/type')
docs\
    .filter(docs.timestamp > datetime.datetime(2017, 5, 1, 0, 0))\
    .filter(docs.timestamp <= datetime.datetime(2017, 5, 2, 0, 0))\
    .count() # solely for demonstration

```

it takes virtually forever (\>30m ?). The solution I found is to increase the "es.scroll.size" from its default value of 50 to the default setting of "index.max\_result\_window" of 10000. The operation then is completed in \< 200s!

Is there any particular reason for the default setting of 50? Maybe this could be updated in the documentation?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [June 29, 2017, 2:12pm UTC](https://discuss.elastic.co/t/correct-setting-of-es-scroll-size-with-for-optimal-spark-read-performance/91077/2 "2017-06-29T14:12:24Z")

</div>

Glad to hear you were able to fix your performance issue! 50 is a reasonable starting number for most users. In many cases you don't want the scroll size to default to 10000 due to the increased memory pressure on anything that is reading that much data all at once.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 27, 2017, 2:12pm UTC](https://discuss.elastic.co/t/correct-setting-of-es-scroll-size-with-for-optimal-spark-read-performance/91077/3 "2017-07-27T14:12:31Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
