# Running cardinality for more than 10000 buckets

**URL:** <https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717>\
**Category:** Elasticsearch\
**Created:** [July 29, 2019, 2:03pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717 "2019-07-29T14:03:55Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![divyang](https://avatars.discourse-cdn.com/v4/letter/d/7ea924/32.png) [@divyang](https://discuss.elastic.co/u/divyang)\
**Post date:** [July 29, 2019, 2:03pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/1 "2019-07-29T14:03:55Z")

</div>

Hi, I am making a project to get the unique number of userids per url in in our index .  
Index has more than 5000 urls and far more number of users ids. To the get the unique number if user ids i am using cardinality on the buckets received for the urls . Right now testing on a single node but elastic shuts down when I am querying the data for a whole month . If i use Pagination , i wont be able to get the unique records . Surely , in production I would increase the number of nodes . But , is there any option you would like to suggest ?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 29, 2019, 2:08pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/2 "2019-07-29T14:08:32Z")

</div>

Hi divyang,  
Cardinality aggs nested underneath a field like URL which has a lot of unique values uses a lot of RAM.  
Your options are for reducing RAM usage are:

1. Trade space for some accuracy using the `precision_threshold` setting of the cardinality agg
2. Break your single request into multiple calls (using either the `composite` agg instead of the `terms` agg or use the `partition` feature of the `terms` aggregation).

---

<div class="post-metadata">

**Author:** ![divyang](https://avatars.discourse-cdn.com/v4/letter/d/7ea924/32.png) [@divyang](https://discuss.elastic.co/u/divyang)\
**Post date:** [July 29, 2019, 2:35pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/3 "2019-07-29T14:35:01Z")

</div>

Thanks for the prompt reply . Partioning feature is very helpful , Just to clarify , a given url will be present in only 1 partition , not any other , right

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 29, 2019, 2:52pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/4 "2019-07-29T14:52:43Z")

</div>

> [@divyang](#):
>
> a given url will be present in only 1 partition , not any other , right

Correct. The composite agg offers that guarantee too but can't sort by child agg (eg URLs sorted by number of unique visitors). The `terms` agg can sort by child agg but the order guarantees are only for the URLs that fall into the same partition.

---

<div class="post-metadata">

**Author:** ![divyang](https://avatars.discourse-cdn.com/v4/letter/d/7ea924/32.png) [@divyang](https://discuss.elastic.co/u/divyang)\
**Post date:** [July 29, 2019, 3:00pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/5 "2019-07-29T15:00:57Z")

</div>

i havent used sorting in the query , is it necessary ? on what basis are the partitions formed ?

---

<div class="post-metadata">

**Author:** ![divyang](https://avatars.discourse-cdn.com/v4/letter/d/7ea924/32.png) [@divyang](https://discuss.elastic.co/u/divyang)\
**Post date:** [July 29, 2019, 3:02pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/6 "2019-07-29T15:02:39Z")

</div>

Also , I am not sure how composite aggregation would be used with cardinality , so havent tried it

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 29, 2019, 3:06pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/7 "2019-07-29T15:06:39Z")

</div>

> [@divyang](#):
>
> on what basis are the partitions formed ?

Hash modulo N. The same technique used to route documents by ID to a choice of shard.  
You just pick what "N" is at query time and it's a way of evenly dividing up a set of values based on hashing the values.

> [@divyang](#):
>
> , I am not sure how composite aggregation would be used with cardinality ,

Composite agg processes values in value order - if terms partitioning is taking a random subset of all terms in each "page" then composite agg is getting the next N terms _after_ the last page's last term. You just swap the `composite` agg for your `terms` agg and make URL the choice of value by which it sorts buckets.

---

<div class="post-metadata">

**Author:** ![divyang](https://avatars.discourse-cdn.com/v4/letter/d/7ea924/32.png) [@divyang](https://discuss.elastic.co/u/divyang)\
**Post date:** [July 29, 2019, 3:25pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/8 "2019-07-29T15:25:26Z")

</div>

i tried the composite aggregation as well , it isnt giving me a direct answer , i will need to process the results in program code . Are you suggesting composite aggregation over partioning , Im in favour of partioning because its giving a direct answer

Result after a composite Query:

{

```
"after_key": {
    "page_urlpath": "/",
    "domain_sessionid": "00000d9f-7628-429e-8f97-f82aba38b2d6"
},
"buckets": [
    {
        "key": {
            "page_urlpath": "/",
            "domain_sessionid": "000006ba-cead-4577-8374-b5ae43434a47"
        },
        "doc_count": 4
    }
    ,
    {
        "key": {
            "page_urlpath": "/",
            "domain_sessionid": "00000d9f-7628-429e-8f97-f82aba38b2d6"
        },
        "doc_count": 3
    }
```

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 29, 2019, 3:46pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/9 "2019-07-29T15:46:35Z")

</div>

I need to see your query JSON - it looks like you haven't embedded the cardinality agg underneath the composite agg.

---

<div class="post-metadata">

**Author:** ![divyang](https://avatars.discourse-cdn.com/v4/letter/d/7ea924/32.png) [@divyang](https://discuss.elastic.co/u/divyang)\
**Post date:** [July 29, 2019, 3:48pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/10 "2019-07-29T15:48:35Z")

</div>

Sorry I tried it again its working , finally im getting answer through both methods , which one do you suggest or are both good ?

This is my new query json;

{  
"from": 0,  
"size": 0,  
"sort" : [{"page\_urlpath":{"order":"asc"}}],  
"aggs": {  
"my\_buckets": {  
"composite": {  
"size": 2,  
"sources": [  
{  
"page\_urlpath": {  
"terms": {  
"field": "page\_urlpath.keyword"  
}  
}  
}  
]   
},  
"aggregations": {  
"visitors": {  
"cardinality": {  
"field": "domain\_sessionid.keyword",  
"precision\_threshold": 40000  
}  
}  
}  
}  
}  
}

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 29, 2019, 3:55pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/11 "2019-07-29T15:55:27Z")

</div>

Not much in it I expect but I'd probably go with composite if order is unimportant.

> [@divyang](#):
>
> "size": 2,

Is that not on the small side?

---

<div class="post-metadata">

**Author:** ![divyang](https://avatars.discourse-cdn.com/v4/letter/d/7ea924/32.png) [@divyang](https://discuss.elastic.co/u/divyang)\
**Post date:** [July 29, 2019, 3:57pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/12 "2019-07-29T15:57:47Z")

</div>

No that was just for seeing if the query is working , i had tried composite queries before but were giving errors , Later i tried with a size of 100 , it worked , then a size of 1000 , elasrtic went down 😀

---

<div class="post-metadata">

**Author:** ![divyang](https://avatars.discourse-cdn.com/v4/letter/d/7ea924/32.png) [@divyang](https://discuss.elastic.co/u/divyang)\
**Post date:** [July 31, 2019, 3:09pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/13 "2019-07-31T15:09:29Z")

</div>

As I am running partitioning for querying the pageurl buckets , I am thinking to use 100-150 partitions , is that ok ? Is there a cap on the number of partitions or a max number that is advisable ?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 31, 2019, 4:07pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/14 "2019-07-31T16:07:25Z")

</div>

[The docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html#_filtering_values_with_partitions) talk about how to pick a suitable size

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 28, 2019, 4:07pm UTC](https://discuss.elastic.co/t/running-cardinality-for-more-than-10000-buckets/192717/15 "2019-08-28T16:07:31Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
