# Get all documents from an index

**URL:** <https://discuss.elastic.co/t/get-all-documents-from-an-index/86977>\
**Category:** Elasticsearch\
**Created:** [May 24, 2017, 12:48pm UTC](https://discuss.elastic.co/t/get-all-documents-from-an-index/86977 "2017-05-24T12:48:22Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![McElroy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mcelroy/32/24639_2.png) [@McElroy](https://discuss.elastic.co/u/McElroy)\
**Post date:** [May 24, 2017, 12:48pm UTC](https://discuss.elastic.co/t/get-all-documents-from-an-index/86977/1 "2017-05-24T12:48:23Z")

</div>

Is it possible to get all the documents from an index?  
I tried it with python and requests but always get  
`query_phase_execution_exception","reason":"Result window is too large, from + size must be less than or equal to: [10000] but was [11000]. See the scroll api for a more efficient way to request large data sets. This limit can be set by changing the [index.max_result_window] index level setting.`

I have no idea how the scroll api works and the documentation isn't helpful for me either.  
Could someone please help me.

---

<div class="post-metadata">

**Author:** ![Heinmci](https://avatars.discourse-cdn.com/v4/letter/h/6a8cbe/32.png) [@Heinmci](https://discuss.elastic.co/u/Heinmci)\
**Post date:** [May 24, 2017, 1:01pm UTC](https://discuss.elastic.co/t/get-all-documents-from-an-index/86977/2 "2017-05-24T13:01:08Z")

</div>

Hi,  
You said you tried with python, so I'll just show you what worked for me :

```
es = Elasticsearch(['http://yourElasticIP:9200/'])
doc = {
        'size' : 10000,
        'query': {
            'match_all' : {}
       }
   }
res = es.search(index='indexname', doc_type='typename', body=doc,scroll='1m')

```

Then you get a reponse with your matching documents and also an attribute named '\_scroll\_id'

So you can do

```
scrollId = res['_scroll_id']
es.scroll(scroll_id = scrollId, scroll = '1m')

```

Where res is the result of your previous es search.  
You can do the es.scroll as many times as you need, just remember to update the scrollId value each time you do a new request  
Sorry if I wasn't very clear

---

<div class="post-metadata">

**Author:** ![McElroy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mcelroy/32/24639_2.png) [@McElroy](https://discuss.elastic.co/u/McElroy)\
**Post date:** [May 24, 2017, 1:03pm UTC](https://discuss.elastic.co/t/get-all-documents-from-an-index/86977/3 "2017-05-24T13:03:00Z")

</div>

Thank you.  
Size indicates how many hits i get?

---

<div class="post-metadata">

**Author:** ![Heinmci](https://avatars.discourse-cdn.com/v4/letter/h/6a8cbe/32.png) [@Heinmci](https://discuss.elastic.co/u/Heinmci)\
**Post date:** [May 24, 2017, 1:03pm UTC](https://discuss.elastic.co/t/get-all-documents-from-an-index/86977/4 "2017-05-24T13:03:35Z")

</div>

Yes, but as you saw, it can't be over 10 000, so you have to use the scroll API, don't think you have another choice

---

<div class="post-metadata">

**Author:** ![McElroy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mcelroy/32/24639_2.png) [@McElroy](https://discuss.elastic.co/u/McElroy)\
**Post date:** [May 24, 2017, 1:06pm UTC](https://discuss.elastic.co/t/get-all-documents-from-an-index/86977/5 "2017-05-24T13:06:14Z")

</div>

Ok, so I will get the first 10 000 results. How do I get the rest?  
Sorry for my stupid asking, but I am missing the forest through the trees right now.

---

<div class="post-metadata">

**Author:** ![Heinmci](https://avatars.discourse-cdn.com/v4/letter/h/6a8cbe/32.png) [@Heinmci](https://discuss.elastic.co/u/Heinmci)\
**Post date:** [May 24, 2017, 1:16pm UTC](https://discuss.elastic.co/t/get-all-documents-from-an-index/86977/6 "2017-05-24T13:16:47Z")

</div>

If you look at the code above, the es.scroll function allows you to get results past 10 000.

```
es = Elasticsearch(['http://x.x.x.x:9200/'])
doc = {
    'size' : 10000,
    'query': {
        'match_all' : {}
    }
}

res = es.search(index="myIndex", doc_type='myType', body=doc,scroll='1m')
scroll = res['_scroll_id']
res2 = es.scroll(scroll_id = scroll, scroll = '1m')

```

In this example, you have your first 10 000 hits in res, and the next 10 000 in res2. If you want results from 20 000 to 30 000, you just get the new scroll id value from res 2!

---

<div class="post-metadata">

**Author:** ![McElroy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mcelroy/32/24639_2.png) [@McElroy](https://discuss.elastic.co/u/McElroy)\
**Post date:** [May 24, 2017, 1:18pm UTC](https://discuss.elastic.co/t/get-all-documents-from-an-index/86977/7 "2017-05-24T13:18:39Z")

</div>

AHHHH... forest, there it is.  
Thank you for your help. It finally made click.

---

<div class="post-metadata">

**Author:** ![Heinmci](https://avatars.discourse-cdn.com/v4/letter/h/6a8cbe/32.png) [@Heinmci](https://discuss.elastic.co/u/Heinmci)\
**Post date:** [May 24, 2017, 1:19pm UTC](https://discuss.elastic.co/t/get-all-documents-from-an-index/86977/8 "2017-05-24T13:19:32Z")

</div>

No problem, have a nice day!

---

<div class="post-metadata">

**Author:** ![McElroy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mcelroy/32/24639_2.png) [@McElroy](https://discuss.elastic.co/u/McElroy)\
**Post date:** [May 24, 2017, 1:43pm UTC](https://discuss.elastic.co/t/get-all-documents-from-an-index/86977/9 "2017-05-24T13:43:09Z")

</div>

on one of my indexes I get no data just  
`{'timed_out': False, 'hits': {'total': 1843, 'max_score': 1.0, 'hits': []}, '_shards': {'successful': 5, 'total': 5, 'failed': 0}, 'terminated_early': False, '_scroll_id': 'DnF1ZXJ5VGhlbkZldGNoBQAAAAAAARfQFm5rMUVCeUxTVDJHUm5qZ2dBQkpJMncAAAAAAAExGBZyNFIxMV93QVRqT0wtTTNoZ1dUenN3AAAAAAABF88WbmsxRUJ5TFNUMkdSbmpnZ0FCSkkydwAAAAAAAPrrFnpFTW9aaHRPUzd1X0Y0UHRORTFpSFEAAAAAAAExFxZyNFIxMV93QVRqT0wtTTNoZ1dUenN3', 'took': 2}`  
Any idea on that?

---

<div class="post-metadata">

**Author:** ![Heinmci](https://avatars.discourse-cdn.com/v4/letter/h/6a8cbe/32.png) [@Heinmci](https://discuss.elastic.co/u/Heinmci)\
**Post date:** [May 24, 2017, 1:49pm UTC](https://discuss.elastic.co/t/get-all-documents-from-an-index/86977/10 "2017-05-24T13:49:47Z")

</div>

Sorry, not sure why that is.  
Only thing that comes to mind is that size was set to 0, other than that I don't know

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 21, 2017, 1:50pm UTC](https://discuss.elastic.co/t/get-all-documents-from-an-index/86977/11 "2017-06-21T13:50:05Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
