# Is there a way to do scan with limit

**URL:** <https://discuss.elastic.co/t/is-there-a-way-to-do-scan-with-limit/120103>\
**Category:** Elasticsearch\
**Created:** [February 16, 2018, 6:49am UTC](https://discuss.elastic.co/t/is-there-a-way-to-do-scan-with-limit/120103 "2018-02-16T06:49:52Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![ryankauk](https://avatars.discourse-cdn.com/v4/letter/r/5daacb/32.png) [@ryankauk](https://discuss.elastic.co/u/ryankauk)\
**Post date:** [February 16, 2018, 6:49am UTC](https://discuss.elastic.co/t/is-there-a-way-to-do-scan-with-limit/120103/1 "2018-02-16T06:49:53Z")

</div>

Hey so basically I'm using elastic search to retrieve a lot of data fast in order to do mapped calculations on it on the server. I'm planning to prepare for millions to tens of millions having to be loaded at once.

I came across the scan function in python, so I do the scan on each shard as well as split them into separate processes.

However I still would like to put a threshold on this in case it ever reaches to hundreds of millions and would just like to get a data sample size.

Please let me know if this feature exists as I can't seem to find any documentation. So far elastic search is exactly what I need and am just missing this small feature.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 16, 2018, 7:54am UTC](https://discuss.elastic.co/t/is-there-a-way-to-do-scan-with-limit/120103/2 "2018-02-16T07:54:26Z")

</div>

When using the scroll API, you can just call [clear scroll](https://www.elastic.co/guide/en/elasticsearch/reference/6.2/search-request-scroll.html#_clear_scroll_api) anytime you think you are done (ie you have enough data).

---

<div class="post-metadata">

**Author:** ![ryankauk](https://avatars.discourse-cdn.com/v4/letter/r/5daacb/32.png) [@ryankauk](https://discuss.elastic.co/u/ryankauk)\
**Post date:** [March 7, 2018, 7:52pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-do-scan-with-limit/120103/4 "2018-03-07T19:52:07Z")

</div>

Using the python scan helper I'm not sure how you can get the scroll Id from it though. As well I split the scan into each shard the index has for maximum parallelism.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 4, 2018, 7:56pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-do-scan-with-limit/120103/5 "2018-04-04T19:56:40Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
