# Listing Analyzers

**URL:** <https://discuss.elastic.co/t/listing-analyzers/4566>\
**Category:** Elasticsearch\
**Created:** [June 7, 2011, 7:57pm UTC](https://discuss.elastic.co/t/listing-analyzers/4566 "2011-06-07T19:57:29Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![phobos182](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/phobos182/32/3011_2.png) [@phobos182](https://discuss.elastic.co/u/phobos182)\
**Post date:** [June 7, 2011, 7:57pm UTC](https://discuss.elastic.co/t/listing-analyzers/4566/1 "2011-06-07T19:57:29Z")

</div>

I know that ElasticSearch has a lot of built in analyzers. Basically i'm looking to perform specific analyzers based upon the language identification of a field. I know that I can use the build in "analyzer" field to specify which analyzer I wish based on a field name.

My initial thought was going to be to use my "language" field to determine which analyzer I want to use. So if the "Language" field is "English", I would want to use the english analyzer.

Which brings me to my point. Instead of re-inventing the wheel and creating a lot of custom analyzers for each language, I would like to use the built-in tokenizers / stop words / etc.. for each language. I cannot find a list of built in analyzers that elasticsearch uses so I can just specify as an example "analyzer: english". I would like to know how what each analyzers stopword list is, etc..

Any documentation regarding this?

Thanks,

---

<div class="post-metadata">

**Author:** ![Paul\_Loy](https://avatars.discourse-cdn.com/v4/letter/p/ad7895/32.png) [@Paul\_Loy](https://discuss.elastic.co/u/Paul_Loy)\
**Post date:** [June 7, 2011, 7:59pm UTC](https://discuss.elastic.co/t/listing-analyzers/4566/2 "2011-06-07T19:59:19Z")

</div>

> **[Elastic — The Search AI Company](https://www.elastic.co)**
>
> Power insights and outcomes with The Elastic Search AI Platform. See into your data and find answers that matter with enterprise solutions designed to help you accelerate time to insight. Try Elastic ...

On Tue, Jun 7, 2011 at 8:57 PM, phobos182 [phobos182@gmail.com](mailto:phobos182@gmail.com) wrote:

> I know that Elasticsearch has a lot of built in analyzers. Basically i'm  
> looking to perform specific analyzers based upon the language  
> identification  
> of a field. I know that I can use the build in "analyzer" field to specify  
> which analyzer I wish based on a field name.
> 
> My initial thought was going to be to use my "language" field to determine  
> which analyzer I want to use. So if the "Language" field is "English", I  
> would want to use the english analyzer.
> 
> Which brings me to my point. Instead of re-inventing the wheel and creating  
> a lot of custom analyzers for each language, I would like to use the  
> built-in tokenizers / stop words / etc.. for each language. I cannot find a  
> list of built in analyzers that elasticsearch uses so I can just specify as  
> an example "analyzer: english". I would like to know how what each  
> analyzers  
> stopword list is, etc..
> 
> Any documentation regarding this?
> 
> Thanks,
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Listing-Analyzers-tp3036342p3036342.html](http://elasticsearch-users.115913.n3.nabble.com/Listing-Analyzers-tp3036342p3036342.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

## --

Paul Loy  
[paul@keteracel.com](mailto:paul@keteracel.com)  
[http://uk.linkedin.com/in/paulloy](http://uk.linkedin.com/in/paulloy)

---

<div class="post-metadata">

**Author:** ![phobos182](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/phobos182/32/3011_2.png) [@phobos182](https://discuss.elastic.co/u/phobos182)\
**Post date:** [June 7, 2011, 8:29pm UTC](https://discuss.elastic.co/t/listing-analyzers/4566/3 "2011-06-07T20:29:19Z")

</div>

Thanks. I did not see the "Language" analyzer on the right side.

Any idea what stopwords comprise these analyzers? Any way to look deeper into them to find out how they are constructed?

---

<div class="post-metadata">

**Author:** ![Paul\_Loy](https://avatars.discourse-cdn.com/v4/letter/p/ad7895/32.png) [@Paul\_Loy](https://discuss.elastic.co/u/Paul_Loy)\
**Post date:** [June 7, 2011, 9:13pm UTC](https://discuss.elastic.co/t/listing-analyzers/4566/4 "2011-06-07T21:13:33Z")

</div>

They use the Lucene standard stopwords. Someone on this mailing list posted  
a link but I can't find it...

On Tue, Jun 7, 2011 at 9:29 PM, phobos182 [phobos182@gmail.com](mailto:phobos182@gmail.com) wrote:

> Thanks. I did not see the "Language" analyzer on the right side.
> 
> Any idea what stopwords comprise these analyzers? Any way to look deeper  
> into them to find out how they are constructed?
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Listing-Analyzers-tp3036342p3036572.html](http://elasticsearch-users.115913.n3.nabble.com/Listing-Analyzers-tp3036342p3036572.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

## --

Paul Loy  
[paul@keteracel.com](mailto:paul@keteracel.com)  
[http://uk.linkedin.com/in/paulloy](http://uk.linkedin.com/in/paulloy)

---

<div class="post-metadata">

**Author:** ![Paul\_Loy](https://avatars.discourse-cdn.com/v4/letter/p/ad7895/32.png) [@Paul\_Loy](https://discuss.elastic.co/u/Paul_Loy)\
**Post date:** [June 7, 2011, 9:14pm UTC](https://discuss.elastic.co/t/listing-analyzers/4566/5 "2011-06-07T21:14:36Z")

</div>

here we go, Solr has a good reference:

> **[LanguageAnalysis - Solr - Apache Software Foundation](https://cwiki.apache.org/confluence/display/solr/LanguageAnalysis)**

On Tue, Jun 7, 2011 at 10:13 PM, Paul Loy [keteracel@gmail.com](mailto:keteracel@gmail.com) wrote:

> They use the Lucene standard stopwords. Someone on this mailing list posted  
> a link but I can't find it...
> 
> On Tue, Jun 7, 2011 at 9:29 PM, phobos182 [phobos182@gmail.com](mailto:phobos182@gmail.com) wrote:
> 
> > Thanks. I did not see the "Language" analyzer on the right side.
> > 
> > Any idea what stopwords comprise these analyzers? Any way to look deeper  
> > into them to find out how they are constructed?
> > 
> > --  
> > View this message in context:  
> > [http://elasticsearch-users.115913.n3.nabble.com/Listing-Analyzers-tp3036342p3036572.html](http://elasticsearch-users.115913.n3.nabble.com/Listing-Analyzers-tp3036342p3036572.html)  
> > Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> 
> ## --
> 
> Paul Loy  
> [paul@keteracel.com](mailto:paul@keteracel.com)  
> [Paul Loy - Amihan Entertainment | LinkedIn](http://uk.linkedin.com/in/paulloy)

## --

Paul Loy  
[paul@keteracel.com](mailto:paul@keteracel.com)  
[http://uk.linkedin.com/in/paulloy](http://uk.linkedin.com/in/paulloy)

---

<div class="post-metadata">

**Author:** ![phobos182](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/phobos182/32/3011_2.png) [@phobos182](https://discuss.elastic.co/u/phobos182)\
**Post date:** [June 7, 2011, 11:48pm UTC](https://discuss.elastic.co/t/listing-analyzers/4566/6 "2011-06-07T23:48:01Z")

</div>

I see the stopwords for each language. It seems that they use the Snowball stemmer for each type with the language identifier.

For some fields i'm looking for more precision, and less recall. So I will have to use some custom analyzers for them, but for the others this looks good.

Thanks again,

---

<div class="post-metadata">

**Author:** ![fashionalwallet](https://avatars.discourse-cdn.com/v4/letter/f/839c29/32.png) [@fashionalwallet](https://discuss.elastic.co/u/fashionalwallet)\
**Post date:** [June 10, 2011, 12:30am UTC](https://discuss.elastic.co/t/listing-analyzers/4566/7 "2011-06-10T00:30:14Z")

</div>

- deleted -

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:04am UTC](https://discuss.elastic.co/t/listing-analyzers/4566/8 "2017-07-06T04:04:08Z")

</div>


