# How to script export of \> 10,000 records - 5 mil?

**URL:** <https://discuss.elastic.co/t/how-to-script-export-of-10-000-records-5-mil/56029>\
**Category:** Elasticsearch\
**Created:** [July 20, 2016, 9:19pm UTC](https://discuss.elastic.co/t/how-to-script-export-of-10-000-records-5-mil/56029 "2016-07-20T21:19:27Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![zoplex](https://avatars.discourse-cdn.com/v4/letter/z/bc8723/32.png) [@zoplex](https://discuss.elastic.co/u/zoplex)\
**Post date:** [July 20, 2016, 9:19pm UTC](https://discuss.elastic.co/t/how-to-script-export-of-10-000-records-5-mil/56029/1 "2016-07-20T21:19:27Z")

</div>

I need to batch export of all record for one hour, one day, etc .. Often record counts are \> 10,000 (Much more than that); I cannot change the 10,000 limit on size parameter for curl - and since this is automated - any other option except scroll approach? Scroll seems complex for batching/scripting ...

Thanks,

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [July 20, 2016, 9:37pm UTC](https://discuss.elastic.co/t/how-to-script-export-of-10-000-records-5-mil/56029/2 "2016-07-20T21:37:43Z")

</div>

There isn't any option other than scroll. I bet some of the language clients like, python, perl, or ruby have scroll helpers that'd make it simpler.

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [July 20, 2016, 10:07pm UTC](https://discuss.elastic.co/t/how-to-script-export-of-10-000-records-5-mil/56029/3 "2016-07-20T22:07:22Z")

</div>

Why don't you use combination of `from`/`size` and page through the result set?

---

<div class="post-metadata">

**Author:** ![zoplex](https://avatars.discourse-cdn.com/v4/letter/z/bc8723/32.png) [@zoplex](https://discuss.elastic.co/u/zoplex)\
**Post date:** [July 20, 2016, 10:58pm UTC](https://discuss.elastic.co/t/how-to-script-export-of-10-000-records-5-mil/56029/4 "2016-07-20T22:58:13Z")

</div>

It needs to run in batch file ... assuming what you are talking about is interactive?

---

<div class="post-metadata">

**Author:** ![zoplex](https://avatars.discourse-cdn.com/v4/letter/z/bc8723/32.png) [@zoplex](https://discuss.elastic.co/u/zoplex)\
**Post date:** [July 20, 2016, 11:43pm UTC](https://discuss.elastic.co/t/how-to-script-export-of-10-000-records-5-mil/56029/5 "2016-07-20T23:43:05Z")

</div>

is it possible to script it ? Also I tried simple scroll call in ES 5 and got error:

curl -v -X GET 'localhost:9200/filebeat-2016.07.20/\_search?search\_type=scan&scroll=1m' -d '{ "query": { "match\_all": {} }, "size": 100000 }'  
...

{"error":{"root\_cause":[{"type":"illegal\_argument\_exception","reason":"No search type for [scan]"}],"type":"illegal\_argument\_exception","reason":"No search type for [scan]"},"status":400}root@node01:/app\_2/query

---

<div class="post-metadata">

**Author:** ![zoplex](https://avatars.discourse-cdn.com/v4/letter/z/bc8723/32.png) [@zoplex](https://discuss.elastic.co/u/zoplex)\
**Post date:** [July 21, 2016, 1:56am UTC](https://discuss.elastic.co/t/how-to-script-export-of-10-000-records-5-mil/56029/6 "2016-07-21T01:56:35Z")

</div>

.. another problem is that documentation indicates that scroll does not sort data - so if data is needed sorted by let's say timestamp, then sorting would have to be done outside once the data is extracted:

[https://www.elastic.co/guide/en/elasticsearch/guide/1.x/scan-scroll.html](https://www.elastic.co/guide/en/elasticsearch/guide/1.x/scan-scroll.html)

...  
"The costly part of deep pagination is the global sorting of results, but if we disable sorting, then we can return all documents quite cheaply. To do this, we use the scan search type. Scan instructs Elasticsearch to do no sorting, but to just return the next batch of results from every shard that still has results to return."

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 21, 2016, 6:13am UTC](https://discuss.elastic.co/t/how-to-script-export-of-10-000-records-5-mil/56029/7 "2016-07-21T06:13:06Z")

</div>

Read: [Scroll | Elasticsearch Guide [2.3] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/2.3/search-request-scroll.html)

> Scroll requests have optimizations that make them faster when the sort order is \_doc. If you want to iterate over all documents regardless of the order, this is the most efficient option:

> ```
> curl -XGET 'localhost:9200/_search?scroll=1m' -d '
> {
> "sort": [
> "_doc"
> ]
> }
> '
> 
> ```

So you can scroll with any other sort criteria.

Note that you are not forced to use scan when you scroll. Scan has been removed in 2.x.

---

<div class="post-metadata">

**Author:** ![zoplex](https://avatars.discourse-cdn.com/v4/letter/z/bc8723/32.png) [@zoplex](https://discuss.elastic.co/u/zoplex)\
**Post date:** [July 28, 2016, 5:13pm UTC](https://discuss.elastic.co/t/how-to-script-export-of-10-000-records-5-mil/56029/8 "2016-07-28T17:13:42Z")

</div>

Thank you all for the answers; I am also looking at changing the system parameter per suggestion from the co-worker:

curl -XPUT "[http://localhost:9200/myindex/\_settings](http://localhost:9200/myindex/_settings) "-d '{ "index" : { "max\_result\_window" : 500000 } }'

While this has some impact on ES - if large extracts are used only at night/say 1/day to get the data that is to be input into further analysis, it would be probably be manageable.

Thanks you

---

<div class="post-metadata">

**Author:** ![Dilip\_Kumar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dilip_kumar/32/15459_2.png) [@Dilip\_Kumar](https://discuss.elastic.co/u/Dilip_Kumar)\
**Post date:** [February 13, 2017, 10:17am UTC](https://discuss.elastic.co/t/how-to-script-export-of-10-000-records-5-mil/56029/9 "2017-02-13T10:17:48Z")

</div>

Hi zoplex

You can use this to increase fetch documents

curl -XPUT [http://localhost:9200/indexname/\_settings](http://localhost:9200/indexname/_settings) -d '{ "index" : { "max\_result\_window" : 1000000}}'

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:03pm UTC](https://discuss.elastic.co/t/how-to-script-export-of-10-000-records-5-mil/56029/10 "2017-07-05T22:03:19Z")

</div>


