# Setting up a custom analyzer

**URL:** <https://discuss.elastic.co/t/setting-up-a-custom-analyzer/101480>\
**Category:** Elasticsearch\
**Created:** [September 22, 2017, 12:18pm UTC](https://discuss.elastic.co/t/setting-up-a-custom-analyzer/101480 "2017-09-22T12:18:58Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Karolinebryn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/karolinebryn/32/17095_2.png) [@Karolinebryn](https://discuss.elastic.co/u/Karolinebryn)\
**Post date:** [September 22, 2017, 12:18pm UTC](https://discuss.elastic.co/t/setting-up-a-custom-analyzer/101480/1 "2017-09-22T12:18:58Z")

</div>

Hi!  
I am creating a custom analyzer for one of my indexes, and I have some questions about the lowercase and standard token filters.

In the documentation it says this about the [lowecase tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-lowercase-tokenizer.html)

> The lowercase tokenizer, like the letter tokenizer breaks text into terms whenever it encounters a character which is not a letter, but it also lowercases all terms.

While this is said for the [standard tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-standard-tokenizer.html)

> The standard tokenizer provides grammar based tokenization

Does this mean that there is no point in using them both? Does the lowecase tokenizer overlap the standard tokenizer?

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [September 22, 2017, 5:21pm UTC](https://discuss.elastic.co/t/setting-up-a-custom-analyzer/101480/2 "2017-09-22T17:21:50Z")

</div>

Only one tokenizer can be defined per analyzer. Keep in mind that  
tokenizers and token filters are different items, with the former being  
executed first (of the two) in the analysis chain.

The lowercase tokenizer [1] is based on the letter tokenizer [2], which  
simply breaks on non-letter characters. The standard tokenizer [3] is far  
more complex, with various rules mostly based on the English language. It  
all depends on your corpus and use cases. Data such as names and titles  
could use a simpler letter tokenizer, but free form text that might  
included urls or email address is probably best tokenized by the standard  
tokenizer.

[1]

> <https://github.com/apache/lucene-solr/blob/branch_6x/lucene/analysis/common/src/java/org/apache/lucene/analysis/core/LowerCaseTokenizer.java>

  
[2]  

> <https://github.com/apache/lucene-solr/blob/branch_6x/lucene/analysis/common/src/java/org/apache/lucene/analysis/core/LetterTokenizer.java>

  
[3]  

> <https://github.com/apache/lucene-solr/blob/branch_6x/lucene/core/src/java/org/apache/lucene/analysis/standard/StandardTokenizer.java>

Cheers,

Ivan

---

<div class="post-metadata">

**Author:** ![rjernst](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rjernst/32/6363_2.png) [@rjernst](https://discuss.elastic.co/u/rjernst)\
**Post date:** [September 22, 2017, 7:09pm UTC](https://discuss.elastic.co/t/setting-up-a-custom-analyzer/101480/3 "2017-09-22T19:09:02Z")

</div>

> [@Ivan](#):
>
> The standard tokenizer [3] is far more complex, with various rules mostly based on the English language

As an aside (unrelated to the original question), the English part of this statement is not true. It is based on the Unicode Text Segmentation algorithm. See [UAX #29: Unicode Text Segmentation](http://unicode.org/reports/tr29/). The standard _analyzer_ has some English stuff, specifically the default set of English stop words.

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [September 22, 2017, 8:57pm UTC](https://discuss.elastic.co/t/setting-up-a-custom-analyzer/101480/4 "2017-09-22T20:57:21Z")

</div>

Very true Ryan. I meant to say based on Latin character set languages, but  
even that is false. I hope that the OP sees the difference between  
tokenizers and token filters, especially for the standard tokenizer/token  
filter. The former does tons, the latter does nothing!

---

<div class="post-metadata">

**Author:** ![Karolinebryn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/karolinebryn/32/17095_2.png) [@Karolinebryn](https://discuss.elastic.co/u/Karolinebryn)\
**Post date:** [September 25, 2017, 7:36am UTC](https://discuss.elastic.co/t/setting-up-a-custom-analyzer/101480/5 "2017-09-25T07:36:01Z")

</div>

Okay, then I am messing up the terms (I am really confused now). I thought a token filter was made up by one or more tokenizers (that's at least what I made of [this text](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-tokenfilters.html)).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [September 25, 2017, 2:50pm UTC](https://discuss.elastic.co/t/setting-up-a-custom-analyzer/101480/6 "2017-09-25T14:50:49Z")

</div>

Analyzers are made up of filters and tokenizers as described here  
[https://www.elastic.co/guide/en/elasticsearch/reference/current/analyzer-anatomy.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/analyzer-anatomy.html)

A diagram can be found here:  
[https://www.elastic.co/blog/found-text-analysis-part-1](https://www.elastic.co/blog/found-text-analysis-part-1) The concepts come  
straight from Lucene, so any informations sources regarding analysis in  
Lucene/Solr will apply to Elasticsearch if you care to read more.

That diagram does not highlight the fact that you can have several  
character filters and token filters, but only one tokenizer. In general,  
character filters are seldom used (mainly for pattern removal or  
substitution), then a simple tokenizer, followed by several token filters  
which work on the tokens generated by the tokenizer. Chances are you want  
to focus on the token filters.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 23, 2017, 2:50pm UTC](https://discuss.elastic.co/t/setting-up-a-custom-analyzer/101480/7 "2017-10-23T14:50:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
