# What is the difference between using "keyword" tokenizer and "not\_analyzed"

**URL:** <https://discuss.elastic.co/t/what-is-the-difference-between-using-keyword-tokenizer-and-not-analyzed/50870>\
**Category:** Elasticsearch\
**Created:** [May 24, 2016, 5:31pm UTC](https://discuss.elastic.co/t/what-is-the-difference-between-using-keyword-tokenizer-and-not-analyzed/50870 "2016-05-24T17:31:41Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![photonic\_world\_2](https://avatars.discourse-cdn.com/v4/letter/p/e5b9ba/32.png) [@photonic\_world\_2](https://discuss.elastic.co/u/photonic_world_2)\
**Post date:** [May 24, 2016, 5:31pm UTC](https://discuss.elastic.co/t/what-is-the-difference-between-using-keyword-tokenizer-and-not-analyzed/50870/1 "2016-05-24T17:31:41Z")

</div>

Hi,

What is the difference between using "keyword" tokenizer and "not\_analyzed" on the fields?

Does either provider better performance in aggregations compared with the other if the field size is the same (for e.g: if I am aggregating on email addresses, with a company of 1 million employees)

Thanks!

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [May 24, 2016, 6:08pm UTC](https://discuss.elastic.co/t/what-is-the-difference-between-using-keyword-tokenizer-and-not-analyzed/50870/2 "2016-05-24T18:08:10Z")

</div>

`not_analyzed` is slightly faster at index time.

The `keyword` tokenizer allows you to use token filters like `lowercase`.

It shouldn't matter either way for aggregations once you've paid the (comparatively low) price to build the query for the any filtering you do before the aggregation.

---

<div class="post-metadata">

**Author:** ![photonic\_world\_2](https://avatars.discourse-cdn.com/v4/letter/p/e5b9ba/32.png) [@photonic\_world\_2](https://discuss.elastic.co/u/photonic_world_2)\
**Post date:** [May 24, 2016, 6:35pm UTC](https://discuss.elastic.co/t/what-is-the-difference-between-using-keyword-tokenizer-and-not-analyzed/50870/3 "2016-05-24T18:35:46Z")

</div>

Thanks @nik9000

We are using keyword tokenizer just to use the filter lowercase.

Followup question which I have been wondering about:

Since an aggregation gets all values of a field aggregated on into doc\_values or field\_data.  
1-\> Does query filtering improve memory footprint and time taken to fetch aggregations apart from having large number of buckets?  
2-\> Is there any difference between using [Aggregation filter](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-filter-aggregation.html) and query filters on aggregation query? Will either change the aggregation performance.

---

<div class="post-metadata">

**Author:** ![photonic\_world\_2](https://avatars.discourse-cdn.com/v4/letter/p/e5b9ba/32.png) [@photonic\_world\_2](https://discuss.elastic.co/u/photonic_world_2)\
**Post date:** [May 26, 2016, 4:10pm UTC](https://discuss.elastic.co/t/what-is-the-difference-between-using-keyword-tokenizer-and-not-analyzed/50870/4 "2016-05-26T16:10:46Z")

</div>

I have instance of two mappings one with fields 'analyzer: keyword' analyzed and other with fields 'not\_analyzed'. In some instances why does aggregation on index with fields 'index:not\_analyzed' take less time and less field data space.

I am confused what is the real difference between these at search time @nik9000 ?

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [May 26, 2016, 7:24pm UTC](https://discuss.elastic.co/t/what-is-the-difference-between-using-keyword-tokenizer-and-not-analyzed/50870/5 "2016-05-26T19:24:49Z")

</div>

> [@photonic\_world\_2](#):
>
> 1-\> Does query filtering improve memory footprint and time taken to fetch aggregations apart from having large number of buckets? 2-\> Is there any difference between using Aggregation filter and query filters on aggregation query? Will either change the aggregation performance.

Using a query to filter is going to be faster because you never have to load anything from doc values.

> [@photonic\_world\_2](#):
>
> I have instance of two mappings one with fields 'analyzer: keyword' analyzed and other with fields 'not\_analyzed'. In some instances why does aggregation on index with fields 'index:not\_analyzed' take less time and less field data space.

Ah ha! A thing I forgot! `not_analyzed` supports doc\_values which will use way, way less memory. As soon as you use an analyzer you get field data only. Elasticsearch 5.0 is coming with a thing that lets you use an analyzer, but only one that emits only a single token, and still use doc values.

So my suggestion is `not_analyzed` all the time.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:48pm UTC](https://discuss.elastic.co/t/what-is-the-difference-between-using-keyword-tokenizer-and-not-analyzed/50870/6 "2017-07-05T22:48:47Z")

</div>


