# How to extract the content of elasticsearch indexes?

**URL:** <https://discuss.elastic.co/t/how-to-extract-the-content-of-elasticsearch-indexes/123300>\
**Category:** Elasticsearch\
**Created:** [March 9, 2018, 4:08pm UTC](https://discuss.elastic.co/t/how-to-extract-the-content-of-elasticsearch-indexes/123300 "2018-03-09T16:08:09Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![fmind](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fmind/32/34006_2.png) [@fmind](https://discuss.elastic.co/u/fmind)\
**Post date:** [March 9, 2018, 4:08pm UTC](https://discuss.elastic.co/t/how-to-extract-the-content-of-elasticsearch-indexes/123300/1 "2018-03-09T16:08:10Z")

</div>

Hello,

I want to perform an analysis on elasticsearch indexes.

For instance, I have an index that stores two fields: 'name' and 'age'.

The result I want are the documents associated to each value of each field:

name =\> ['bob' =\> [Doc#1, Doc#2, Doc#3],  
'alice' =\> [Doc#4, Doc#5]]

age =\> ['20' =\> [Doc#2, Doc#4],  
'30' =\> [Doc#1, Doc#5]  
'40' =\> [Doc#3]]

Is there a way to perform this kind of query in elasticsearch ?

Thank you !

---

<div class="post-metadata">

**Author:** ![fmind](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fmind/32/34006_2.png) [@fmind](https://discuss.elastic.co/u/fmind)\
**Post date:** [March 13, 2018, 10:52am UTC](https://discuss.elastic.co/t/how-to-extract-the-content-of-elasticsearch-indexes/123300/2 "2018-03-13T10:52:33Z")

</div>

Bump

---

<div class="post-metadata">

**Author:** ![abdon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/abdon/32/9195_2.png) [@abdon](https://discuss.elastic.co/u/abdon)\
**Post date:** [March 13, 2018, 2:41pm UTC](https://discuss.elastic.co/t/how-to-extract-the-content-of-elasticsearch-indexes/123300/3 "2018-03-13T14:41:51Z")

</div>

You would an aggregation instead of a query for this. You could for example use a Terms aggregation.

Given these docs:

```auto
PUT my_index/doc/_bulk
{ "index" : { "_id" : "1" } }
{"name": "bob", "age": 30}
{ "index" : { "_id" : "2" } }
{"name": "bob", "age": 20}
{ "index" : { "_id" : "3" } }
{"name": "bob", "age": 30}
{ "index" : { "_id" : "4" } }
{"name": "alice", "age": 20}
{ "index" : { "_id" : "5" } }
{"name": "alice", "age": 30}

```

You can get the IDs by name using this aggregation:

```auto
GET my_index/_search
{
  "size": 0,
  "aggs": {
    "top_names": {
      "terms": {
        "field": "name.keyword",
        "size": 100
      },
      "aggs": {
        "top_ids": {
          "terms": {
            "field": "_id",
            "size": 100
          }
        }
      }
    }
  }
}

```

Which will return you:

```auto
"buckets": [
        {
          "key": "bob",
          "doc_count": 3,
          "top_ids": {
            "doc_count_error_upper_bound": 0,
            "sum_other_doc_count": 0,
            "buckets": [
              {
                "key": "1",
                "doc_count": 1
              },
              {
                "key": "2",
                "doc_count": 1
              },
              {
                "key": "3",
                "doc_count": 1
              }
            ]
          }
        },
        {
          "key": "alice",
          "doc_count": 2,
          "top_ids": {
            "doc_count_error_upper_bound": 0,
            "sum_other_doc_count": 0,
            "buckets": [
              {
                "key": "4",
                "doc_count": 1
              },
              {
                "key": "5",
                "doc_count": 1
              }
            ]
          }
        }
      ]

```

Aggregating on `_id` is not possible on older versions of Elasticsearch. You may need to replace `_id` with `_uid` (which is a concatenation of the `_type` and `_id`) if you're using an older version.

To aggregate on age you would use replace `name.keyword` with `age` in the request above.

Note the `"size": 100` in the request above. This will limit the response to contain the 100 most common names and will return you up to 100 IDs. You could increase that number if you need to retrieve more values (but you may run into memory limitations). Or alternatively, if you're on version 6.1 or later, you could also take a look at the [composite aggregation](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-composite-aggregation.html) to retrieve all values and IDs.

---

<div class="post-metadata">

**Author:** ![fmind](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fmind/32/34006_2.png) [@fmind](https://discuss.elastic.co/u/fmind)\
**Post date:** [March 14, 2018, 9:33am UTC](https://discuss.elastic.co/t/how-to-extract-the-content-of-elasticsearch-indexes/123300/4 "2018-03-14T09:33:16Z")

</div>

Perfect, this is exactly what I need !

Thank you

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 11, 2018, 9:33am UTC](https://discuss.elastic.co/t/how-to-extract-the-content-of-elasticsearch-indexes/123300/5 "2018-04-11T09:33:26Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
