# Elasticsearch scroll

**URL:** <https://discuss.elastic.co/t/elasticsearch-scroll/157366>\
**Category:** Elasticsearch\
**Created:** [November 19, 2018, 1:29pm UTC](https://discuss.elastic.co/t/elasticsearch-scroll/157366 "2018-11-19T13:29:27Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![lonzodc](https://avatars.discourse-cdn.com/v4/letter/l/f08c70/32.png) [@lonzodc](https://discuss.elastic.co/u/lonzodc)\
**Post date:** [November 19, 2018, 1:29pm UTC](https://discuss.elastic.co/t/elasticsearch-scroll/157366/1 "2018-11-19T13:29:27Z")

</div>

Hi All,

I would like to ask the best approach to extract thousands of records in Elasticsearch? Also is it possible to perform 1000 concurrent request using scroll api to extract the data from the elasticsearch index?

Thank you,

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 19, 2018, 1:31pm UTC](https://discuss.elastic.co/t/elasticsearch-scroll/157366/2 "2018-11-19T13:31:30Z")

</div>

How large is your cluster? How much data are you looking to extract? How many indices and shards is this spread across?

---

<div class="post-metadata">

**Author:** ![lonzodc](https://avatars.discourse-cdn.com/v4/letter/l/f08c70/32.png) [@lonzodc](https://discuss.elastic.co/u/lonzodc)\
**Post date:** [November 19, 2018, 1:43pm UTC](https://discuss.elastic.co/t/elasticsearch-scroll/157366/3 "2018-11-19T13:43:51Z")

</div>

Hi Christian,

We are using AWS ES m4.large having 8GB memory and 300GB EBS. We are planning to extract 200 thousands of records. We also have 1000 plus indices having 5 shards each indices.

Thank you,

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 19, 2018, 1:44pm UTC](https://discuss.elastic.co/t/elasticsearch-scroll/157366/4 "2018-11-19T13:44:50Z")

</div>

How many nodes in the cluster? Just 1?

---

<div class="post-metadata">

**Author:** ![lonzodc](https://avatars.discourse-cdn.com/v4/letter/l/f08c70/32.png) [@lonzodc](https://discuss.elastic.co/u/lonzodc)\
**Post date:** [November 19, 2018, 1:46pm UTC](https://discuss.elastic.co/t/elasticsearch-scroll/157366/5 "2018-11-19T13:46:57Z")

</div>

We have 7 nodes in total. The details are as follows 3 master nodes and 4 datanodes

Thank you,

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 19, 2018, 1:50pm UTC](https://discuss.elastic.co/t/elasticsearch-scroll/157366/6 "2018-11-19T13:50:16Z")

</div>

The first thing I would like to point out is that you have far too many indices and shards for a cluster that size. This can be very inefficient. I recommend you [read this blog post about shards](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster) as it provides some practical guidelines.

Given the relatively low amount of heap available I would recommend running a few school queries at a time so you can determine how much the cluster can handle. I would not be surprised if you are suffering from heap pressure given the number of shards in the cluster. Performing 1000 requests in parallel would likely make it fall over.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 17, 2018, 2:00pm UTC](https://discuss.elastic.co/t/elasticsearch-scroll/157366/7 "2018-12-17T14:00:41Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
