# Generate Aggregation List for Large Index

**URL:** <https://discuss.elastic.co/t/generate-aggregation-list-for-large-index/70149>\
**Category:** Elasticsearch\
**Created:** [December 28, 2016, 3:34pm UTC](https://discuss.elastic.co/t/generate-aggregation-list-for-large-index/70149 "2016-12-28T15:34:50Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![krmathieu](https://avatars.discourse-cdn.com/v4/letter/k/b5e925/32.png) [@krmathieu](https://discuss.elastic.co/u/krmathieu)\
**Post date:** [December 28, 2016, 3:34pm UTC](https://discuss.elastic.co/t/generate-aggregation-list-for-large-index/70149/1 "2016-12-28T15:34:50Z")

</div>

I am trying to get all unique values in a given field using the following terms aggregations and it is returning a "can't communicate with server error" but I no that is not the actual issue because if I drop the size value to 2000000 it works:

{  
"size": 0,  
"aggs":{  
"names":{  
"terms": { "field": "name",  
"size": 4000000  
}  
}  
}  
} \> myResults.txt

I am using Elastic Cloud with 16GB of RAM and 384GB of disk space and I believe the cluster just can't handle the larger number of results. Is there anyway to get all 4 million unique values out that I need for post processing? Any help that anyone could provide would be appreciated. Thanks.

Kevin

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [December 28, 2016, 4:13pm UTC](https://discuss.elastic.co/t/generate-aggregation-list-for-large-index/70149/2 "2016-12-28T16:13:13Z")

</div>

In the next release (5.2) we have support for partitioning terms into an arbitrary number of sets and working with one set at a time. See [https://www.elastic.co/guide/en/elasticsearch/reference/5.x/search-aggregations-bucket-terms-aggregation.html#\_filtering\_values\_with\_partitions](https://www.elastic.co/guide/en/elasticsearch/reference/5.x/search-aggregations-bucket-terms-aggregation.html#_filtering_values_with_partitions)

---

<div class="post-metadata">

**Author:** ![krmathieu](https://avatars.discourse-cdn.com/v4/letter/k/b5e925/32.png) [@krmathieu](https://discuss.elastic.co/u/krmathieu)\
**Post date:** [December 28, 2016, 6:19pm UTC](https://discuss.elastic.co/t/generate-aggregation-list-for-large-index/70149/3 "2016-12-28T18:19:49Z")

</div>

Thanks for the info. In the interim, is there a way I can do separate queries and just merge the results? Maybe add a filter query to just do names starting with A-L and then another starting with M-Z? Not sure how to use REGEX in an ES query though. Please let me know if this would be doable. Thanks.

Kevin

---

<div class="post-metadata">

**Author:** ![krmathieu](https://avatars.discourse-cdn.com/v4/letter/k/b5e925/32.png) [@krmathieu](https://discuss.elastic.co/u/krmathieu)\
**Post date:** [December 28, 2016, 9:17pm UTC](https://discuss.elastic.co/t/generate-aggregation-list-for-large-index/70149/4 "2016-12-28T21:17:07Z")

</div>

This problem is solved, I was in fact able to use a REGEX filter and run two separate queries to do what I needed and here is what the latter query looks like:

```
{
    "size": 0,
    "aggs" : {
        "names" : {
            "filter" : { "regexp": { "name": "[m-zM-Z].*" } },
            "aggs" : {
                "filteredNames" : { "terms": {"field" : "name", "size": 2000000} }
            }
        }
    }
}
```

Just wanted to close the loop on this issue. Thanks.

Kevin

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 25, 2017, 9:17pm UTC](https://discuss.elastic.co/t/generate-aggregation-list-for-large-index/70149/5 "2017-01-25T21:17:12Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
