# Terms aggregation by long faster than by string?

**URL:** <https://discuss.elastic.co/t/terms-aggregation-by-long-faster-than-by-string/38525>\
**Category:** Elasticsearch\
**Created:** [January 6, 2016, 5:27pm UTC](https://discuss.elastic.co/t/terms-aggregation-by-long-faster-than-by-string/38525 "2016-01-06T17:27:11Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![dorr](https://avatars.discourse-cdn.com/v4/letter/d/d2c977/32.png) [@dorr](https://discuss.elastic.co/u/dorr)\
**Post date:** [January 6, 2016, 5:27pm UTC](https://discuss.elastic.co/t/terms-aggregation-by-long-faster-than-by-string/38525/1 "2016-01-06T17:27:12Z")

</div>

Hello,  
I have an ID field with very high cardinality, currently implemented as a string, containing content similar to a GUID.  
I wish to perform terms aggregations on a large data, and want to optimize this.

I read [this article](https://www.elastic.co/guide/en/elasticsearch/guide/1.x/preload-fielddata.html) that discusses ordinals and was wondering:  
If I change the field implementation to a long, would that help in terms of query speed / memory usage / anything?

Thanks.

---

<div class="post-metadata">

**Author:** ![bleskes](https://avatars.discourse-cdn.com/v4/letter/b/71c47a/32.png) [@bleskes](https://discuss.elastic.co/u/bleskes)\
**Post date:** [January 6, 2016, 7:34pm UTC](https://discuss.elastic.co/t/terms-aggregation-by-long-faster-than-by-string/38525/2 "2016-01-06T19:34:41Z")

</div>

Internally strings and numbers are treated as bytes. When matters is how the bytes are distributed. Numbers also have the "down side" of being chopped to multiple terms to speed up range searches (see [https://www.elastic.co/guide/en/elasticsearch/reference/2.1/precision-step.html](https://www.elastic.co/guide/en/elasticsearch/reference/2.1/precision-step.html) ). In general GUIDs are fine, but check this blog for advice on how to optimize them: [http://blog.mikemccandless.com/2014/05/choosing-fast-unique-identifier-uuid.html](http://blog.mikemccandless.com/2014/05/choosing-fast-unique-identifier-uuid.html)

---

<div class="post-metadata">

**Author:** ![dorr](https://avatars.discourse-cdn.com/v4/letter/d/d2c977/32.png) [@dorr](https://discuss.elastic.co/u/dorr)\
**Post date:** [January 7, 2016, 12:05pm UTC](https://discuss.elastic.co/t/terms-aggregation-by-long-faster-than-by-string/38525/3 "2016-01-07T12:05:08Z")

</div>

Hi Boaz, thanks for the info.  
I will look into the formatting of whatever type I choose. I see precision\_step is only for Elasticsearch 2.0+. Are there any recommendations for v1.7?

Also, I'm still wondering about this (from the link I posted):

> [@](#):
>
> Ordinals are only built and used for strings. Numerical data (integers, geopoints, dates, etc) doesn’t need an ordinal mapping, since the value itself acts as an intrinsic ordinal mapping.

Can switching to a numeric type help the performance of my query as well?

P.S. It's important to note I'm doing terms aggregation on a contextual ID field that is shared between multiple records (i.e. "session\_id"), not on the unique document ID itself, if that matters.  
Thanks.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:26pm UTC](https://discuss.elastic.co/t/terms-aggregation-by-long-faster-than-by-string/38525/4 "2017-07-05T23:26:25Z")

</div>


