# Summary of index types

**URL:** <https://discuss.elastic.co/t/summary-of-index-types/5142>\
**Category:** Elasticsearch\
**Created:** [August 12, 2011, 5:58pm UTC](https://discuss.elastic.co/t/summary-of-index-types/5142 "2011-08-12T17:58:42Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 12, 2011, 5:58pm UTC](https://discuss.elastic.co/t/summary-of-index-types/5142/1 "2011-08-12T17:58:42Z")

</div>

I don't have a Lucene background, so this may be rudimentary.

The index types are defined here with an explanation:  
[http://www.elasticsearch.org/guide/reference/mapping/core-types.html](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)

_analyzed_ (default)

Indexed, Searchable, Tokenized

\*not\_analyzed \*

Indexed, Searchable, Not Tokenized

_no_

Not Indexed, Not Searchable, Not Tokenized

The first option (analyzed) and the last (no) seem to make a lot of sense,  
and I understand them.

However, I do struggle to come up with a use case where not\_analyzed should  
be used. Since it is not tokenized, I would expect a match only when the  
exact same search term is provided in a query. In fact, since it is not  
tokenized, it has to match right down to stop words, whitespace and case,  
correct? Maybe for matching on a hash code, a case-insensitive username, or  
a zip code?

If I have a not\_analyzed field containing "XY&Z Company", will I only get a  
match if I query for "XY&Z Company"?

---

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [August 12, 2011, 7:28pm UTC](https://discuss.elastic.co/t/summary-of-index-types/5142/2 "2011-08-12T19:28:47Z")

</div>

We use not\_analyzed for generating facet results that can be used for  
display purposes. Also, there are for fields that are exact match  
filters, this can be appropriate, similar to what you were guessing  
below.

Best Regards,  
Paul

On Aug 12, 11:58 am, James Cook [jc...@tracermedia.com](mailto:jc...@tracermedia.com) wrote:

> I don't have a Lucene background, so this may be rudimentary.
> 
> The index types are defined here with an explanation:[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)
> 
> _analyzed_ (default)
> 
> Indexed, Searchable, Tokenized
> 
> \*not\_analyzed \*
> 
> Indexed, Searchable, Not Tokenized
> 
> _no_
> 
> Not Indexed, Not Searchable, Not Tokenized
> 
> The first option (analyzed) and the last (no) seem to make a lot of sense,  
> and I understand them.
> 
> However, I do struggle to come up with a use case where not\_analyzed should  
> be used. Since it is not tokenized, I would expect a match only when the  
> exact same search term is provided in a query. In fact, since it is not  
> tokenized, it has to match right down to stop words, whitespace and case,  
> correct? Maybe for matching on a hash code, a case-insensitive username, or  
> a zip code?
> 
> If I have a not\_analyzed field containing "XY&Z Company", will I only get a  
> match if I query for "XY&Z Company"?

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 12, 2011, 10:08pm UTC](https://discuss.elastic.co/t/summary-of-index-types/5142/3 "2011-08-12T22:08:59Z")

</div>

Just another note on not\_analyzed fields, those are exactly the same as  
fields that are analyzed with a keyword tokenizer (keyword tokenizer simply  
treats the whole text as a single token). People many times want this  
behavior, but also do things like lowercasing, in which case, one can create  
a custom analyzer that has a keyword tokenizer and a lowercase filter, and  
use that as the analyzer to the field (and the field will still be  
"analyzed").

On Fri, Aug 12, 2011 at 10:28 PM, ppearcy [ppearcy@gmail.com](mailto:ppearcy@gmail.com) wrote:

> We use not\_analyzed for generating facet results that can be used for  
> display purposes. Also, there are for fields that are exact match  
> filters, this can be appropriate, similar to what you were guessing  
> below.
> 
> Best Regards,  
> Paul
> 
> On Aug 12, 11:58 am, James Cook [jc...@tracermedia.com](mailto:jc...@tracermedia.com) wrote:
> 
> > I don't have a Lucene background, so this may be rudimentary.
> > 
> > The index types are defined here with an explanation:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)
> > 
> > _analyzed_ (default)
> > 
> > Indexed, Searchable, Tokenized
> > 
> > \*not\_analyzed \*
> > 
> > Indexed, Searchable, Not Tokenized
> > 
> > _no_
> > 
> > Not Indexed, Not Searchable, Not Tokenized
> > 
> > The first option (analyzed) and the last (no) seem to make a lot of  
> > sense,  
> > and I understand them.
> > 
> > However, I do struggle to come up with a use case where not\_analyzed  
> > should  
> > be used. Since it is not tokenized, I would expect a match only when the  
> > exact same search term is provided in a query. In fact, since it is not  
> > tokenized, it has to match right down to stop words, whitespace and case,  
> > correct? Maybe for matching on a hash code, a case-insensitive username,  
> > or  
> > a zip code?
> > 
> > If I have a not\_analyzed field containing "XY&Z Company", will I only get  
> > a  
> > match if I query for "XY&Z Company"?

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 13, 2011, 11:01am UTC](https://discuss.elastic.co/t/summary-of-index-types/5142/4 "2011-08-13T11:01:52Z")

</div>

I had found this Lucene presentation[http://www.slideshare.net/otisg/lucene-introduction](http://www.slideshare.net/otisg/lucene-introduction)by Otis Gospodnetic. He presented the built-in analyzers in a very clear  
way.

_The quick brown fox jumped over the lazy dog._

_WhitespaceAnalyzer_:

```
[The] [quick] [brown] [fox] [jumped] [over] [the] [lazy] [dogs]

```

_SimpleAnalyzer_:

```
[the] [quick] [brown] [fox] [jumped] [over] [the] [lazy] [dogs]

```

_StopAnalyzer_:

```
[quick] [brown] [fox] [jumped] [over] [lazy] [dogs] 

```

_StandardAnalyzer_:

```
[quick] [brown] [fox] [jumped] [over] [lazy] [dogs] 

```

"\*XY&Z Corporation - \*_[xyz@example.com](mailto:xyz@example.com)_ [xyz@example.com](mailto:xyz@example.com)"

_WhitespaceAnalyzer_:

```
[XY&Z] [Corporation] [-] [xyz@example.com] 

```

_SimpleAnalyzer_:

```
[xy] [z] [corporation] [xyz] [example] [com] 

```

_StopAnalyzer_:

```
[xy] [z] [corporation] [xyz] [example] [com] 

```

_StandardAnalyzer_:

```
[xy&z] [corporation] [xyz@example.com]

```

As Shay stated, we use a keyword tokenizer with a lowercase filter for  
username lookups. I suppose based on the analyzer example above, we could  
use the SimpleAnalyzer to achieve a similar result?

index.analysis.analyzer.lowercase\_keyword.type=custom  
index.analysis.analyzer.lowercase\_keyword.tokenizer=keyword  
index.analysis.analyzer.lowercase\_keyword.filter.0=lowercase

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 13, 2011, 11:06am UTC](https://discuss.elastic.co/t/summary-of-index-types/5142/5 "2011-08-13T11:06:51Z")

</div>

simple analyzer breaks text into tokens at non letters...., so its different  
than the keyword tokenizer, which will treat the whole text as a single  
token.

On Sat, Aug 13, 2011 at 2:01 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:

> I had found this Lucene presentation[http://www.slideshare.net/otisg/lucene-introduction](http://www.slideshare.net/otisg/lucene-introduction)by Otis Gospodnetic. He presented the built-in analyzers in a very clear  
> way.
> 
> _The quick brown fox jumped over the lazy dog._
> 
> _WhitespaceAnalyzer_:
> 
> ```
> [The] [quick] [brown] [fox] [jumped] [over] [the] [lazy] [dogs]
> 
> ```
> 
> _SimpleAnalyzer_:
> 
> ```
> [the] [quick] [brown] [fox] [jumped] [over] [the] [lazy] [dogs]
> 
> ```
> 
> _StopAnalyzer_:
> 
> ```
> [quick] [brown] [fox] [jumped] [over] [lazy] [dogs]
> 
> ```
> 
> _StandardAnalyzer_:
> 
> ```
> [quick] [brown] [fox] [jumped] [over] [lazy] [dogs]
> 
> ```
> 
> "\*XY&Z Corporation - \*_[xyz@example.com](mailto:xyz@example.com)_ [xyz@example.com](mailto:xyz@example.com)"
> 
> _WhitespaceAnalyzer_:
> 
> ```
> [XY&Z] [Corporation] [-] [xyz@example.com]
> 
> ```
> 
> _SimpleAnalyzer_:
> 
> ```
> [xy] [z] [corporation] [xyz] [example] [com]
> 
> ```
> 
> _StopAnalyzer_:
> 
> ```
> [xy] [z] [corporation] [xyz] [example] [com]
> 
> ```
> 
> _StandardAnalyzer_:
> 
> ```
> [xy&z] [corporation] [xyz@example.com]
> 
> ```
> 
> As Shay stated, we use a keyword tokenizer with a lowercase filter for  
> username lookups. I suppose based on the analyzer example above, we could  
> use the SimpleAnalyzer to achieve a similar result?
> 
> index.analysis.analyzer.lowercase\_keyword.type=custom  
> index.analysis.analyzer.lowercase\_keyword.tokenizer=keyword  
> index.analysis.analyzer.lowercase\_keyword.filter.0=lowercase

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 13, 2011, 7:52pm UTC](https://discuss.elastic.co/t/summary-of-index-types/5142/6 "2011-08-13T19:52:37Z")

</div>

Great information. Even though I've been using ES for 18 months, I am still  
learning the very basics it seems.

Can't wait for the book! 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:57am UTC](https://discuss.elastic.co/t/summary-of-index-types/5142/7 "2017-07-06T03:57:06Z")

</div>


