# Need help with scan/scroll using elasticsearch-py client

**URL:** <https://discuss.elastic.co/t/need-help-with-scan-scroll-using-elasticsearch-py-client/78548>\
**Category:** Elasticsearch\
**Created:** [March 14, 2017, 3:17pm UTC](https://discuss.elastic.co/t/need-help-with-scan-scroll-using-elasticsearch-py-client/78548 "2017-03-14T15:17:19Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ksnerd](https://avatars.discourse-cdn.com/v4/letter/k/ad7895/32.png) [@ksnerd](https://discuss.elastic.co/u/ksnerd)\
**Post date:** [March 14, 2017, 3:17pm UTC](https://discuss.elastic.co/t/need-help-with-scan-scroll-using-elasticsearch-py-client/78548/1 "2017-03-14T15:17:19Z")

</div>

I'm using elasticsearch 5.2.1 and elasticsearch-py 5.2.0. My index.max\_result\_window is set to the default (10,000).

I'd like to have the option in my script to return all matching documents for a query. When I execute my script, I consistently get a result like this: `<generator object scan at 0x00B5CE40>` instead of the dict I would expect. I know my query is fine, as I can return the first 10,000 results using es.search without an error.

All that said, I'm a relative noob, so any and all help is appreciated.

My code looks like this (simplified for brevity):

```
from elasticsearch import Elasticsearch, helpers

es = Elasticsearch('hostname', port=9200)

res = helpers.scan(
                client = es,
				scroll = '2m',
                query = {"query":{"bool":{"must": [{"query_string": {"query": escaped_query }},
                        {"range":{"@timestamp":{"gte": from_date, "format": "basic_date"}}}]}}}, 
                index = "custom_data*")

print(res)
```

---

<div class="post-metadata">

**Author:** ![ksnerd](https://avatars.discourse-cdn.com/v4/letter/k/ad7895/32.png) [@ksnerd](https://discuss.elastic.co/u/ksnerd)\
**Post date:** [March 14, 2017, 7:53pm UTC](https://discuss.elastic.co/t/need-help-with-scan-scroll-using-elasticsearch-py-client/78548/2 "2017-03-14T19:53:21Z")

</div>

Just updating my own post for anyone else looking for help on this. So the "scan" helper returns a Python generator object (didn't realize that was a thing). You can use the generator object that's returned to iterate over all matches. The code should look something like this:

```
from elasticsearch import Elasticsearch, helpers

es = Elasticsearch('hostname', port=9200)

res = helpers.scan(
                client = es,
				scroll = '2m',
                query = {"query":{"bool":{"must": [{"query_string": {"query": escaped_query }},
                        {"range":{"@timestamp":{"gte": from_date, "format": "basic_date"}}}]}}}, 
                index = "custom_data*")

for i in res:
    print(i)
```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 11, 2017, 7:53pm UTC](https://discuss.elastic.co/t/need-help-with-scan-scroll-using-elasticsearch-py-client/78548/3 "2017-04-11T19:53:23Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
