# How to improve Scroll runtime for 5 billion record retrieval?

**URL:** <https://discuss.elastic.co/t/how-to-improve-scroll-runtime-for-5-billion-record-retrieval/227734>\
**Category:** Elasticsearch\
**Created:** [April 13, 2020, 6:18am UTC](https://discuss.elastic.co/t/how-to-improve-scroll-runtime-for-5-billion-record-retrieval/227734 "2020-04-13T06:18:36Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sheraz\_Tariq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sheraz_tariq/32/46276_2.png) [@Sheraz\_Tariq](https://discuss.elastic.co/u/Sheraz_Tariq)\
**Post date:** [April 13, 2020, 6:18am UTC](https://discuss.elastic.co/t/how-to-improve-scroll-runtime-for-5-billion-record-retrieval/227734/1 "2020-04-13T06:18:36Z")

</div>

I want to retrieve 5 billion `_id`s from an index in Elasticsearch - the index itself is around 3TB and I'm using the scroll feature to do this. However, I'm getting pretty poor runtime. I'm seeing it take around 11 seconds to retrieve 100k entries (also around 0.5 seconds to retrieve 5k entries). I've tried using a smaller size, 5000 instead of 100k and sorted by `_doc` to get better performance, but I'm still not seeing runtime good enough to pull out 5 billion entries in anything less than 7 days of non stop running. I want to return all the `_ids` for all indices in my cluster, but I'm not even able to get even one index without a very long runtime. I'm not sure what else I can do to improve performance? Should I add more data or router nodes? Will that even help? Or is there something I can do with my scroll query?

Here's what my search looks like:

```auto
       result = @client.search index: indexname,
                            scroll: '1m',
                            body: {size: 5000,
                            sort: [
                                "_doc"
                              ],
                            query: {match_all: {},
                                   },
                            }

```

```auto
result_data = @client.scroll body: {scroll_id: scroll_id, scroll: '1m'}

```

I'm using the scroll\_id from the previous scroll as well?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 13, 2020, 7:43am UTC](https://discuss.elastic.co/t/how-to-improve-scroll-runtime-for-5-billion-record-retrieval/227734/2 "2020-04-13T07:43:31Z")

</div>

What does the latency look like if you do not sort and use the natural sorting order?

Why do you need to do this in the first place?

---

<div class="post-metadata">

**Author:** ![Sheraz\_Tariq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sheraz_tariq/32/46276_2.png) [@Sheraz\_Tariq](https://discuss.elastic.co/u/Sheraz_Tariq)\
**Post date:** [April 13, 2020, 7:51am UTC](https://discuss.elastic.co/t/how-to-improve-scroll-runtime-for-5-billion-record-retrieval/227734/3 "2020-04-13T07:51:56Z")

</div>

This is for a project to get metrics on the number of ids that make it into ES from a database that we use. Without sort, I get 13 seconds to get 100k entries.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 11, 2020, 7:52am UTC](https://discuss.elastic.co/t/how-to-improve-scroll-runtime-for-5-billion-record-retrieval/227734/4 "2020-05-11T07:52:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
