# How to retrieve the content of Elasticsearch reverse indexes?

**URL:** <https://discuss.elastic.co/t/how-to-retrieve-the-content-of-elasticsearch-reverse-indexes/142758>\
**Category:** Elasticsearch\
**Created:** [August 2, 2018, 12:04pm UTC](https://discuss.elastic.co/t/how-to-retrieve-the-content-of-elasticsearch-reverse-indexes/142758 "2018-08-02T12:04:21Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![fmind](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fmind/32/34006_2.png) [@fmind](https://discuss.elastic.co/u/fmind)\
**Post date:** [August 2, 2018, 12:04pm UTC](https://discuss.elastic.co/t/how-to-retrieve-the-content-of-elasticsearch-reverse-indexes/142758/1 "2018-08-02T12:04:21Z")

</div>

Hello,

I want to run an analysis on Elasticsearch to retrieve the content of its reverse indexes.

I need this information to know the list of document ID associated to combination of field/value stored in the reverse index:

For instance, given these 4 documents in my cluster:  
{'\_id': 1, 'os': 'linux', 'lang': 'python'}  
{'\_id': 2, 'os': 'linux', 'lang': 'perl'}  
{'\_id': 3, 'os': 'mac', 'lang': 'ruby'}  
{'\_id': 4, 'os': 'bsd 'lang': 'python'}

I want to return the following results, where '\_ids' contains the list of document id:  
{'os': 'linux', '\_ids': [1, 2]}  
{'os': 'mac', '\_ids': [3]}  
{'os': 'bsd', '\_ids': [4]}  
{'lang': 'python', '\_ids': [1, 4]}  
{'lang': 'ruby', '\_ids': [3]}  
{'lang': 'perl', '\_ids': [2]}

I tested the Composite Aggregation API, but I was only able to return the document count and not the full list of document id:  
{'os': 'linux', 'docs': 2}  
{'os': 'mac, 'docs': 1}  
{'os': 'bsd, 'docs': 1}  
{'lang': 'python', docs: 2}  
{'lang': 'ruby', 'docs': 1}  
{'lang': 'perl', 'docs': 1}

At this point, my options are either to use Elasticsearch Hadoop or to migrate my data to Hadoop directly (the index contains more than 6 million document, and takes 1.7 TB). I could also run a single query per field / value, but that would be really inefficient (in my previous example, that would be 6 queries for 4 documents).

Do you know an Elasticsearch API I can use to extract this information ?

Is there a more lightweight alternative to using Hadoop for this case ?

Thank you for your help

---

<div class="post-metadata">

**Author:** ![fmind](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fmind/32/34006_2.png) [@fmind](https://discuss.elastic.co/u/fmind)\
**Post date:** [August 6, 2018, 8:31am UTC](https://discuss.elastic.co/t/how-to-retrieve-the-content-of-elasticsearch-reverse-indexes/142758/2 "2018-08-06T08:31:53Z")

</div>

bump

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 6, 2018, 9:05am UTC](https://discuss.elastic.co/t/how-to-retrieve-the-content-of-elasticsearch-reverse-indexes/142758/3 "2018-08-06T09:05:45Z")

</div>

May be a terms aggregation then a top hits inner aggregation would work for your case?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 3, 2018, 9:17am UTC](https://discuss.elastic.co/t/how-to-retrieve-the-content-of-elasticsearch-reverse-indexes/142758/4 "2018-09-03T09:17:34Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
