# All ids

**URL:** <https://discuss.elastic.co/t/all-ids/4449>\
**Category:** Elasticsearch\
**Created:** [May 20, 2011, 4:55pm UTC](https://discuss.elastic.co/t/all-ids/4449 "2011-05-20T16:55:07Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Albin\_Stigo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/albin_stigo/32/3227_2.png) [@Albin\_Stigo](https://discuss.elastic.co/u/Albin_Stigo)\
**Post date:** [May 20, 2011, 4:55pm UTC](https://discuss.elastic.co/t/all-ids/4449/1 "2011-05-20T16:55:07Z")

</div>

Hi,

What is the easiest/most efficient way of getting all document ids in an index?

Backstory:  
I have csv file where I want to index each (unique) line. I do a sha1  
on each line and use that as the id when I create a document.  
Periodically the csv file is updated and then I plan to rehash all  
lines and compare the set of ids from the file with the set of ids in  
the index, and in that way know which documents to delete and which to  
add.

--Albin

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [May 20, 2011, 5:18pm UTC](https://discuss.elastic.co/t/all-ids/4449/2 "2011-05-20T17:18:47Z")

</div>

Hi Albin

> What is the easiest/most efficient way of getting all document ids in an index?
> 
> Backstory:  
> I have csv file where I want to index each (unique) line. I do a sha1  
> on each line and use that as the id when I create a document.  
> Periodically the csv file is updated and then I plan to rehash all  
> lines and compare the set of ids from the file with the set of ids in  
> the index, and in that way know which documents to delete and which to  
> add.

How many IDs are we talking about? 100, 1000, 10 million?

If small numbers, then you could use search\_type=scan and scroll to  
retrieve the ids for all your docs, eg:

curl -XGET '[http://127.0.0.1:9200/\_all/\_search?scroll=5m&search\_type=scan](http://127.0.0.1:9200/_all/_search?scroll=5m&search_type=scan)' -d '  
{  
"fields" : {},  
"query" : {  
"match\_all" : {}  
},  
"size" : 100  
}  
'

The above would return 100 records from each shard at a time. You would  
need to get the scroll id from the response to the above query, and then  
retrieve 'tranches' of records using:

curl -XGET '[http://127.0.0.1:9200/\_search/scroll?scroll=5m&scroll\_id=xxx](http://127.0.0.1:9200/_search/scroll?scroll=5m&scroll_id=xxx)'

On every request, you will get a new scroll\_id which you need to pass to  
the next scroll request.

If we're talking about lots of records, then you probably don't want to  
retrieve all IDs at the same time, in which case you should probably  
read (eg) 1000 IDs from the CSV and search for those IDs.

By default, since 0.16, the \_id is no longer indexed, which means that  
you can't retrieve it with eg { term =\> { \_id =\> 123 }}

You can change that when you create the mapping by setting the \_id field  
to { index: "not\_analyzed" } (instead of the default "no")

However, if you know the \_type of your documents, than you can use the  
"ids" query instead, eg:

{ query: { ids: { type: "mydoc", values: [123,124,125] }}}

That doesn't help you with deleting IDs for records that no longer exist  
in the CSV file, because you need some way of marking an existing doc as  
'seen', which requires reindexing that doc.

Which begs the question: might it not be easier to just reindex all your  
data to a new index and then use an alias to point 'myindex' to  
'myindex\_2011\_05\_11'. That way you can reindex every day, repoint the  
alias, and just delete the old index.

hth

clint

---

<div class="post-metadata">

**Author:** ![Albin\_Stigo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/albin_stigo/32/3227_2.png) [@Albin\_Stigo](https://discuss.elastic.co/u/Albin_Stigo)\
**Post date:** [May 20, 2011, 9:05pm UTC](https://discuss.elastic.co/t/all-ids/4449/3 "2011-05-20T21:05:07Z")

</div>

Thanks for a very good answer.

We're talking about 20.000 very small documents. It's basically a  
dictionary with codes and their definition.

Right now I'm leaning towards your conclusion of recreating and using  
an alias. All in line with the KISS principle 🙂

Another idea I had was to use couchdb and the elasticsearch river  
plugin to keep them in sync... Since its very easy to get a list of  
all keys in couchdb. Any thoughts on that?

--Albin

On Fri, May 20, 2011 at 7:18 PM, Clinton Gormley  
[clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk) wrote:

> Hi Albin
> 
> > What is the easiest/most efficient way of getting all document ids in an index?
> > 
> > Backstory:  
> > I have csv file where I want to index each (unique) line. I do a sha1  
> > on each line and use that as the id when I create a document.  
> > Periodically the csv file is updated and then I plan to rehash all  
> > lines and compare the set of ids from the file with the set of ids in  
> > the index, and in that way know which documents to delete and which to  
> > add.
> 
> How many IDs are we talking about? 100, 1000, 10 million?
> 
> If small numbers, then you could use search\_type=scan and scroll to  
> retrieve the ids for all your docs, eg:
> 
> curl -XGET '[http://127.0.0.1:9200/\_all/\_search?scroll=5m&search\_type=scan](http://127.0.0.1:9200/_all/_search?scroll=5m&search_type=scan)' -d '  
> {  
> "fields" : {},  
> "query" : {  
> "match\_all" : {}  
> },  
> "size" : 100  
> }  
> '
> 
> The above would return 100 records from each shard at a time. You would  
> need to get the scroll id from the response to the above query, and then  
> retrieve 'tranches' of records using:
> 
> curl -XGET '[http://127.0.0.1:9200/\_search/scroll?scroll=5m&scroll\_id=xxx](http://127.0.0.1:9200/_search/scroll?scroll=5m&scroll_id=xxx)'
> 
> On every request, you will get a new scroll\_id which you need to pass to  
> the next scroll request.
> 
> If we're talking about lots of records, then you probably don't want to  
> retrieve all IDs at the same time, in which case you should probably  
> read (eg) 1000 IDs from the CSV and search for those IDs.
> 
> By default, since 0.16, the \_id is no longer indexed, which means that  
> you can't retrieve it with eg { term =\> { \_id =\> 123 }}
> 
> You can change that when you create the mapping by setting the \_id field  
> to { index: "not\_analyzed" } (instead of the default "no")
> 
> However, if you know the \_type of your documents, than you can use the  
> "ids" query instead, eg:
> 
> { query: { ids: { type: "mydoc", values: [123,124,125] }}}
> 
> That doesn't help you with deleting IDs for records that no longer exist  
> in the CSV file, because you need some way of marking an existing doc as  
> 'seen', which requires reindexing that doc.
> 
> Which begs the question: might it not be easier to just reindex all your  
> data to a new index and then use an alias to point 'myindex' to  
> 'myindex\_2011\_05\_11'. That way you can reindex every day, repoint the  
> alias, and just delete the old index.
> 
> hth
> 
> clint

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:05am UTC](https://discuss.elastic.co/t/all-ids/4449/4 "2017-07-06T04:05:42Z")

</div>


