# Cardinality Aggregation Hashes Only

**URL:** <https://discuss.elastic.co/t/cardinality-aggregation-hashes-only/17975>\
**Category:** Elasticsearch\
**Created:** [June 7, 2014, 9:54pm UTC](https://discuss.elastic.co/t/cardinality-aggregation-hashes-only/17975 "2014-06-07T21:54:51Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![msukmanowsky](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/msukmanowsky/32/44771_2.png) [@msukmanowsky](https://discuss.elastic.co/u/msukmanowsky)\
**Post date:** [June 7, 2014, 9:54pm UTC](https://discuss.elastic.co/t/cardinality-aggregation-hashes-only/17975/1 "2014-06-07T21:54:51Z")

</div>

Hi there,

We're using ES for web analytics purposes and so far, have loved the  
experience. We create hourly indexes that contain only one type of "url"  
document which has multiple metrics fields like "page\_views". We've  
recently begun looking into how to store more complex metrics that require  
set arithmetic such as "unique views" or "unique visitors".

While the cardinality aggregation  
[http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/search-aggregations-metrics-cardinality-aggregation.html](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/search-aggregations-metrics-cardinality-aggregation.html) is  
awesome, it seems like it'd be crazy for us to store all the user IDs that  
we saw even for an hour on certain URLs as the number could grow to be very  
large, very quickly. Just to clarify, this is the document schema I'm  
saying would probably be silly:

{  
"url": "[http://example.com/](http://example.com/)",  
"hour": "2014-05-31T03:00:00"  
"user\_ids": [  
"e4c88ac4-ccc7-49e0-9a2e-34ab24420d2b",  
"252d0f6e-2e9d-487d-95f4-ac3d53cce977",  
"90b5d83b-44d6-4462-9f4b-3ab41e75143e",  
"b6c9d0f8-5e4f-4308-92eb-be68d7b06d78",  
"7a097ac1-7410-4918-a780-0020197d0b14"  
],  
"metrics": {  
"page\_views": 100  
}  
}

Being fairly new to Lucene and ES, I don't really know what a massive (\>  
100K) user\_ids array per document would do to ES/Lucene at indexing or  
query time. In addition, although that structure would allow us to query  
for hourly URLs that contained a certain user\_id, it's probably beyond our  
current scope. Precomputing the unique number per hour doesn't help us  
when we want to perform aggregations at query time and know unique users  
across a series of hours.

Toying around with two approaches in my head, and I wanted to get some  
feedback:

1. Find a way to store only the HLL object in ES but without the actual  
array of distinct values. This way, we have the benefit of the cardinality  
aggregations, but without storing the full set of user\_ids. Is there a way  
to do this?
2. Store a binary blob which represents a custom HLL that we'll create  
and index. Create a new aggregation for a bitwise OR operation on that  
binary object which would allow us to union the HLLs in the aggregation and  
return that result

I lean a little bit more to solution #2 only because we'd prefer to have  
the HLL's accuracy tuneable instead of rely in ES defaults.

Would love to hear some thoughts on how to solve this kind of issue.

Mike

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/1af2370f-c402-44ac-b05d-fe0b1bee00a8%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/1af2370f-c402-44ac-b05d-fe0b1bee00a8%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![photonic\_world\_2](https://avatars.discourse-cdn.com/v4/letter/p/e5b9ba/32.png) [@photonic\_world\_2](https://discuss.elastic.co/u/photonic_world_2)\
**Post date:** [July 16, 2016, 11:14pm UTC](https://discuss.elastic.co/t/cardinality-aggregation-hashes-only/17975/2 "2016-07-16T23:14:13Z")

</div>

Hello,

I have a similar problem, would love to hear how you got about solving this.

Thanks!

---

<div class="post-metadata">

**Author:** ![msukmanowsky](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/msukmanowsky/32/44771_2.png) [@msukmanowsky](https://discuss.elastic.co/u/msukmanowsky)\
**Post date:** [July 18, 2016, 2:55am UTC](https://discuss.elastic.co/t/cardinality-aggregation-hashes-only/17975/3 "2016-07-18T02:55:46Z")

</div>

Hi @photonic_world_2 . We ended up storing user IDs as an array but with a few important caveats in the mapping (note, we're still on Elasticsearch 1.7.2):

```auto
{
   "my_index": {
      "mappings": {
         "my_doc_type": {
            "properties": {
               "visitors": {
                  "type": "murmur3",
                  "index": "no",
                  "doc_values": true,
                  "fielddata": {
                     "format": "doc_values"
                  }
               }
            }
         }
      }
   }

```

To walk you through these:

- `"type": "murmur3"` ensures that values in this field pre-hashed using murmur3 and the result of that hash is stored in the `visitors.hash` field. This saves us from having to perform hashing at query time which significantly slows down cardinality aggregations.
- `"index": "no"` specifies that we don't need this field searchable, so don't add it to the inverted index. If you need the ability to search for specific visitors in docs, you'll have to set this to `not_analyzed`, but be prepared to pay an indexing and disk penalty.
- `doc_values` is critical for cardinality aggregates and is now the default for all properties in ES \> 2.0 (see [https://www.elastic.co/guide/en/elasticsearch/reference/current/doc-values.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/doc-values.html) for more info)

One last thing, if you end up storing 1,000s of user IDs per doc, you'll also likely want to disable `_source` for this field as storing the JSON blob of a huge set of user IDs takes up a ton of disk. I think we were able to cut storage requirements in half or more by disabling source and the inverted index for this field.

Hope that helps.

---

<div class="post-metadata">

**Author:** ![photonic\_world\_2](https://avatars.discourse-cdn.com/v4/letter/p/e5b9ba/32.png) [@photonic\_world\_2](https://discuss.elastic.co/u/photonic_world_2)\
**Post date:** [July 18, 2016, 3:41pm UTC](https://discuss.elastic.co/t/cardinality-aggregation-hashes-only/17975/4 "2016-07-18T15:41:11Z")

</div>

Thanks @msukmanowsky for the details. We have a similar mapping setup, while our requirement needs 2 cardinality aggregations on the same field. Even though these aggregations are on the same level (i.e not sub aggregations, see `vistors` under category\_agg and `visitors`). I see the effects of [combinatorial explosion](https://www.elastic.co/guide/en/elasticsearch/guide/current/_preventing_combinatorial_explosions.html) . Trying to understand how and why 🙂

> "aggregations": {  
> "category\_agg": {  
> "terms": {  
> "field": "category"  
> },  
> "aggregations": {  
> "visitors": {  
> "cardinality": {  
> "field": "visitors.hash",  
> "precision\_threshold": 10000  
> }  
> },  
> "total\_recipients": {  
> "value\_count": {  
> "field": "visitors.hash"  
> }  
> }  
> }  
> },  
> "visitors": {  
> "cardinality": {  
> "field": "visitors.hash",  
> "precision\_threshold": 10000  
> }  
> }  
> }

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:34pm UTC](https://discuss.elastic.co/t/cardinality-aggregation-hashes-only/17975/5 "2017-07-05T22:34:38Z")

</div>


