# Fastest way to retrieve all ids in an index?

**URL:** <https://discuss.elastic.co/t/fastest-way-to-retrieve-all-ids-in-an-index/180955>\
**Category:** Elasticsearch\
**Created:** [May 14, 2019, 8:46am UTC](https://discuss.elastic.co/t/fastest-way-to-retrieve-all-ids-in-an-index/180955 "2019-05-14T08:46:21Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![akshaymaniyar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/akshaymaniyar/32/45836_2.png) [@akshaymaniyar](https://discuss.elastic.co/u/akshaymaniyar)\
**Post date:** [May 14, 2019, 8:46am UTC](https://discuss.elastic.co/t/fastest-way-to-retrieve-all-ids-in-an-index/180955/1 "2019-05-14T08:46:21Z")

</div>

I have an index which has around 300 million documents. What is the fastest recommended way to retrieve all the documentIDs from the index?

Currently I am using the below python script for doing a scan and scroll to retrieve all the IDs. However this takes around 20-24 hours to fetch all the IDs

```
import csv
import json
import sys
import requests

indexName = sys.argv[1]
url = "http://<<es_host>>/" + str(indexName) + "/_search"

querystring = {"scroll":"1m"}

payload = {
            "query": {},
            "size": 1000,
            "stored_fields": []
}

headers = {
    "Content-Type": "application/json"
}

response = requests.request("POST", url, data=json.dumps(payload), headers=headers, params=querystring)
response = json.loads(response.text)

#print response
f = open('results.txt','w')
while True:
    # print(len(response['hits']['hits']))
    if(len(response['hits']['hits']) == 0):
        break
    for hit in response['hits']['hits']:
        fsn = hit['_id']
        f.write(fsn)
        f.write("\n")

    scroll_id = response['_scroll_id']
    #print scroll_id

    payload = {
        "scroll_id": scroll_id,
        "scroll" : "1m"
    }

    #print (payload)
    url = "http://<<es>>/_search/scroll"
    response = requests.request("POST", url, data=json.dumps(payload), headers=headers)
```

---

<div class="post-metadata">

**Author:** ![thn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thn/32/8061_2.png) [@thn](https://discuss.elastic.co/u/thn)\
**Post date:** [May 14, 2019, 10:54am UTC](https://discuss.elastic.co/t/fastest-way-to-retrieve-all-ids-in-an-index/180955/2 "2019-05-14T10:54:46Z")

</div>

Don't know what's in your data but I think there are a few options that you can think about

- do a scan and roll on the document's timestamp so you can you have multiple processes running on different date/time related range

- do a scan and roll on the document's "type/category/topic/etc"

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 14, 2019, 11:06am UTC](https://discuss.elastic.co/t/fastest-way-to-retrieve-all-ids-in-an-index/180955/3 "2019-05-14T11:06:21Z")

</div>

I think you could be a little faster if you sort on the internal id `_doc` as the documentation says.  
Also have a look at [https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html#sliced-scroll](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html#sliced-scroll)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 11, 2019, 11:18am UTC](https://discuss.elastic.co/t/fastest-way-to-retrieve-all-ids-in-an-index/180955/4 "2019-06-11T11:18:03Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
