# Indexing non-English text

**URL:** <https://discuss.elastic.co/t/indexing-non-english-text/3251>\
**Category:** Elasticsearch\
**Created:** [August 24, 2010, 6:16pm UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251 "2010-08-24T18:16:16Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Andrei](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrei/32/2856_2.png) [@Andrei](https://discuss.elastic.co/u/Andrei)\
**Post date:** [August 24, 2010, 6:16pm UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251/1 "2010-08-24T18:16:16Z")

</div>

I have two questions that related to indexing non-English text.

1. Does ES support accented character folding, i.e. indexing "café",  
but if the search term is "cafe" the doc is still found?

2. If I understand correctly, the analyzers only support English text,  
so indexing Russian, German, etc won't work?

-Andrei

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 24, 2010, 11:45pm UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251/2 "2010-08-24T23:45:13Z")

</div>

On Tue, Aug 24, 2010 at 9:16 PM, Andrei [andrei@zmievski.org](mailto:andrei@zmievski.org) wrote:

> I have two questions that related to indexing non-English text.
> 
> 1. Does ES support accented character folding, i.e. indexing "café",  
> but if the search term is "cafe" the doc is still found?

Yes, you can create your own analyzer and add to it the asciifolding filter.  
The ICU plugin might also be interesting for this.

> 1. If I understand correctly, the analyzers only support English text,  
> so indexing Russian, German, etc won't work?

It depends how far you want to take it. There are specific analyzers for  
different languages. I updated the docs to reflect that.

> -Andrei

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 25, 2010, 12:33pm UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251/3 "2010-08-25T12:33:35Z")

</div>

We have to search text where Arabic and English are both used. I don't  
foresee fields where Arabic and English are contained in the same document,  
but we will definitely have many Arabic and English documents in our index.

Can someone provide configuration options for this scenario?

On Tue, Aug 24, 2010 at 7:45 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> On Tue, Aug 24, 2010 at 9:16 PM, Andrei [andrei@zmievski.org](mailto:andrei@zmievski.org) wrote:
> 
> > I have two questions that related to indexing non-English text.
> > 
> > 1. Does ES support accented character folding, i.e. indexing "café",  
> > but if the search term is "cafe" the doc is still found?
> 
> Yes, you can create your own analyzer and add to it the asciifolding  
> filter. The ICU plugin might also be interesting for this.
> 
> > 1. If I understand correctly, the analyzers only support English text,  
> > so indexing Russian, German, etc won't work?
> 
> It depends how far you want to take it. There are specific analyzers for  
> different languages. I updated the docs to reflect that.
> 
> > -Andrei

---

<div class="post-metadata">

**Author:** ![Andrei](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrei/32/2856_2.png) [@Andrei](https://discuss.elastic.co/u/Andrei)\
**Post date:** [August 25, 2010, 6:13pm UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251/4 "2010-08-25T18:13:17Z")

</div>

On Aug 24, 4:45 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Yes, you can create your own analyzer and add to it the asciifolding filter.  
> The ICU plugin might also be interesting for this.

Do you mean to create one in Java or in the configuration file?

> It depends how far you want to take it. There are specific analyzers for  
> different languages. I updated the docs to reflect that.

Could you link to the page that you updated? I couldn't find the  
references to non-English languages there.

-Andrei

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 25, 2010, 6:57pm UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251/5 "2010-08-25T18:57:58Z")

</div>

On Wed, Aug 25, 2010 at 9:13 PM, Andrei [andrei@zmievski.org](mailto:andrei@zmievski.org) wrote:

> On Aug 24, 4:45 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Yes, you can create your own analyzer and add to it the asciifolding  
> > filter.  
> > The ICU plugin might also be interesting for this.
> 
> Do you mean to create one in Java or in the configuration file?

Its in a configuration file. You create a custom analyzer that include it.

> > It depends how far you want to take it. There are specific analyzers for  
> > different languages. I updated the docs to reflect that.
> 
> Could you link to the page that you updated? I couldn't find the  
> references to non-English languages there.

Here it is:  
[http://www.elasticsearch.com/docs/elasticsearch/index\_modules/analysis/analyzer/lang/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/analyzer/lang/)

> -Andrei

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 26, 2010, 1:31pm UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251/6 "2010-08-26T13:31:55Z")

</div>

Circling around to my earlier question, can I have an English _and_ Arabic  
analyzer specified on the same fields across documents?

On Wed, Aug 25, 2010 at 2:57 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> On Wed, Aug 25, 2010 at 9:13 PM, Andrei [andrei@zmievski.org](mailto:andrei@zmievski.org) wrote:
> 
> > On Aug 24, 4:45 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > Yes, you can create your own analyzer and add to it the asciifolding  
> > > filter.  
> > > The ICU plugin might also be interesting for this.
> > 
> > Do you mean to create one in Java or in the configuration file?
> 
> Its in a configuration file. You create a custom analyzer that include it.
> 
> > > It depends how far you want to take it. There are specific analyzers for  
> > > different languages. I updated the docs to reflect that.
> > 
> > Could you link to the page that you updated? I couldn't find the  
> > references to non-English languages there.
> 
> Here it is:  
> [http://www.elasticsearch.com/docs/elasticsearch/index\_modules/analysis/analyzer/lang/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/analyzer/lang/)
> 
> > -Andrei

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 27, 2010, 10:33am UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251/7 "2010-08-27T10:33:01Z")

</div>

No, you can't specify different analyzers on the same field.

On Thu, Aug 26, 2010 at 4:31 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:

> Circling around to my earlier question, can I have an English _and_ Arabic  
> analyzer specified on the same fields across documents?
> 
> On Wed, Aug 25, 2010 at 2:57 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:
> 
> > On Wed, Aug 25, 2010 at 9:13 PM, Andrei [andrei@zmievski.org](mailto:andrei@zmievski.org) wrote:
> > 
> > > On Aug 24, 4:45 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > 
> > > > Yes, you can create your own analyzer and add to it the asciifolding  
> > > > filter.  
> > > > The ICU plugin might also be interesting for this.
> > > 
> > > Do you mean to create one in Java or in the configuration file?
> > 
> > Its in a configuration file. You create a custom analyzer that include it.
> > 
> > > > It depends how far you want to take it. There are specific analyzers  
> > > > for  
> > > > different languages. I updated the docs to reflect that.
> > > 
> > > Could you link to the page that you updated? I couldn't find the  
> > > references to non-English languages there.
> > 
> > Here it is:  
> > [http://www.elasticsearch.com/docs/elasticsearch/index\_modules/analysis/analyzer/lang/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/analyzer/lang/)
> > 
> > > -Andrei

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [August 27, 2010, 11:34am UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251/8 "2010-08-27T11:34:09Z")

</div>

On Fri, 2010-08-27 at 13:33 +0300, Shay Banon wrote:

> No, you can't specify different analyzers on the same field.

But you can index the same field twice, as a multi field, with different  
analysers:

[http://www.elasticsearch.com/docs/elasticsearch/mapping/multi\_field/](http://www.elasticsearch.com/docs/elasticsearch/mapping/multi_field/)

clint

> On Thu, Aug 26, 2010 at 4:31 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com)  
> wrote:  
> Circling around to my earlier question, can I have an English  
> _and_ Arabic analyzer specified on the same fields across  
> documents?
> 
> ```
> On Wed, Aug 25, 2010 at 2:57 PM, Shay Banon
> <shay.banon@elasticsearch.com> wrote:
> On Wed, Aug 25, 2010 at 9:13 PM, Andrei
> <andrei@zmievski.org> wrote:
>             
> On Aug 24, 4:45 pm, Shay Banon
> <shay.ba...@elasticsearch.com> wrote:
> > Yes, you can create your own analyzer and
> add to it the asciifolding filter.
> > The ICU plugin might also be interesting for
> this.
>                     
>                     
> Do you mean to create one in Java or in the
> configuration file?
>             
>             
> Its in a configuration file. You create a custom
> analyzer that include it.
>              
>                     
> > It depends how far you want to take it.
> There are specific analyzers for
> > different languages. I updated the docs to
> reflect that.
>                     
>                     
> Could you link to the page that you updated? I
> couldn't find the
> references to non-English languages there.
>             
>             
> Here it
> is: http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/analyzer/lang/
>              
>                     
> -Andrei
> 
> ```

--  
Web Announcements Limited is a company registered in England and Wales,  
with company number 05608868, with registered address at 10 Arvon Road,  
London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 27, 2010, 12:57pm UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251/9 "2010-08-27T12:57:24Z")

</div>

In my example, I have an object I am indexing which is similar to a  
discussion thread. The 'content' proeprty will contain text which may be in  
English or Arabic.

If the JSON document I am indexing can determine which language it is using,  
can an analyzer be chosen at index and search time?

I don't know much about mappings yet, but the multi-type approach worries me  
because the 'content' field will be knowingly indexed once with the correct  
analyzer and once with the incorrect analyzer.

It appears from the doc entry that the query is then performed only against  
the 'default' entry in the multi-type instead of applying against all  
multi-type entries. This makes it a bit harder to manage queries I think. If  
multi-type is the only way to be able to search for multilingual text in a  
field, I suppose I will have to adapt. 🙂

A quick search shows there are some analyzers out there that have been  
developed for this problem. (i.e.  
[Cloud Monitoring Tools & Services | Sematext](http://www.sematext.com/products/multilingual-indexer/index.html)) In the  
docs there is a list of built in analyzers. Is it straightforward to include  
and configure other analyzers? Any pointers to docs?

Thanks

On Fri, Aug 27, 2010 at 7:34 AM, Clinton Gormley [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:

> On Fri, 2010-08-27 at 13:33 +0300, Shay Banon wrote:
> 
> > No, you can't specify different analyzers on the same field.
> 
> But you can index the same field twice, as a multi field, with different  
> analysers:
> 
> [http://www.elasticsearch.com/docs/elasticsearch/mapping/multi\_field/](http://www.elasticsearch.com/docs/elasticsearch/mapping/multi_field/)
> 
> clint
> 
> > On Thu, Aug 26, 2010 at 4:31 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com)  
> > wrote:  
> > Circling around to my earlier question, can I have an English  
> > _and_ Arabic analyzer specified on the same fields across  
> > documents?
> > 
> > ```
> > On Wed, Aug 25, 2010 at 2:57 PM, Shay Banon
> > <shay.banon@elasticsearch.com> wrote:
> > On Wed, Aug 25, 2010 at 9:13 PM, Andrei
> > <andrei@zmievski.org> wrote:
> > 
> > On Aug 24, 4:45 pm, Shay Banon
> > <shay.ba...@elasticsearch.com> wrote:
> > > Yes, you can create your own analyzer and
> > add to it the asciifolding filter.
> > > The ICU plugin might also be interesting for
> > this.
> > 
> > Do you mean to create one in Java or in the
> > configuration file?
> > 
> > Its in a configuration file. You create a custom
> > analyzer that include it.
> > 
> > > It depends how far you want to take it.
> > There are specific analyzers for
> > > different languages. I updated the docs to
> > reflect that.
> > 
> > Could you link to the page that you updated? I
> > couldn't find the
> > references to non-English languages there.
> > 
> > Here it
> > is:
> > 
> > ```
> 
> [http://www.elasticsearch.com/docs/elasticsearch/index\_modules/analysis/analyzer/lang/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/analyzer/lang/)
> 
> > ```
> > -Andrei
> > 
> > ```
> 
> --  
> Web Announcements Limited is a company registered in England and Wales,  
> with company number 05608868, with registered address at 10 Arvon Road,  
> London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [September 9, 2010, 11:47pm UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251/10 "2010-09-09T23:47:40Z")

</div>

Hello,

I spotted this reference to Sematext's Multilingual Indexer (MI):

> A quick search shows there are some analyzers out there that have been  
> developed for this problem. (i.e.[http://www.sematext.com/products/multilingual-indexer/index.html](http://www.sematext.com/products/multilingual-indexer/index.html)) In the  
> docs there is a list of built in analyzers. Is it straightforward to include  
> and configure other analyzers? Any pointers to docs?

Not sure if you are asking for MI docs or some other docs. MI comes  
with good docs, but they are not public. Adding it to Solr is well  
documented and easy to do. Let us know if you need it for Elastic  
Search.

## Otis

Sematext :: [http://sematext.com/](http://sematext.com/) :: Solr - Lucene - Nutch  
Lucene ecosystem search :: [http://search-lucene.com/](http://search-lucene.com/)

On Aug 27, 8:57 am, James Cook [jc...@tracermedia.com](mailto:jc...@tracermedia.com) wrote:

> In my example, I have an object I am indexing which is similar to a  
> discussion thread. The 'content' proeprty will contain text which may be in  
> English or Arabic.
> 
> If the JSON document I am indexing can determine which language it is using,  
> can an analyzer be chosen at index and search time?
> 
> I don't know much about mappings yet, but the multi-type approach worries me  
> because the 'content' field will be knowingly indexed once with the correct  
> analyzer and once with the incorrect analyzer.
> 
> It appears from the doc entry that the query is then performed only against  
> the 'default' entry in the multi-type instead of applying against all  
> multi-type entries. This makes it a bit harder to manage queries I think. If  
> multi-type is the only way to be able to search for multilingual text in a  
> field, I suppose I will have to adapt. 🙂
> 
> A quick search shows there are some analyzers out there that have been  
> developed for this problem. (i.e.[http://www.sematext.com/products/multilingual-indexer/index.html](http://www.sematext.com/products/multilingual-indexer/index.html)) In the  
> docs there is a list of built in analyzers. Is it straightforward to include  
> and configure other analyzers? Any pointers to docs?
> 
> Thanks
> 
> On Fri, Aug 27, 2010 at 7:34 AM, Clinton Gormley [clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk)wrote:
> 
> > On Fri, 2010-08-27 at 13:33 +0300, Shay Banon wrote:
> > 
> > > No, you can't specify different analyzers on the same field.
> 
> > But you can index the same field twice, as a multi field, with different  
> > analysers:
> 
> > [http://www.elasticsearch.com/docs/elasticsearch/mapping/multi\_field/](http://www.elasticsearch.com/docs/elasticsearch/mapping/multi_field/)
> 
> > clint
> 
> > > On Thu, Aug 26, 2010 at 4:31 PM, James Cook [jc...@tracermedia.com](mailto:jc...@tracermedia.com)  
> > > wrote:  
> > > Circling around to my earlier question, can I have an English  
> > > _and_ Arabic analyzer specified on the same fields across  
> > > documents?
> 
> > > ```
> > > On Wed, Aug 25, 2010 at 2:57 PM, Shay Banon
> > > <shay.ba...@elasticsearch.com> wrote:
> > > On Wed, Aug 25, 2010 at 9:13 PM, Andrei
> > > <and...@zmievski.org> wrote:
> > > 
> > > ```
> 
> > > ```
> > > On Aug 24, 4:45 pm, Shay Banon
> > > <shay.ba...@elasticsearch.com> wrote:
> > > > Yes, you can create your own analyzer and
> > > add to it the asciifolding filter.
> > > > The ICU plugin might also be interesting for
> > > this.
> > > 
> > > ```
> 
> > > ```
> > > Do you mean to create one in Java or in the
> > > configuration file?
> > > 
> > > ```
> 
> > > ```
> > > Its in a configuration file. You create a custom
> > > analyzer that include it.
> > > 
> > > ```
> 
> > > ```
> > > > It depends how far you want to take it.
> > > There are specific analyzers for
> > > > different languages. I updated the docs to
> > > reflect that.
> > > 
> > > ```
> 
> > > ```
> > > Could you link to the page that you updated? I
> > > couldn't find the
> > > references to non-English languages there.
> > > 
> > > ```
> 
> > > ```
> > > Here it
> > > is:
> > > 
> > > ```
> > 
> > [http://www.elasticsearch.com/docs/elasticsearch/index\_modules/analysi](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysi)...
> 
> > > ```
> > > -Andrei
> > > 
> > > ```
> 
> > --  
> > Web Announcements Limited is a company registered in England and Wales,  
> > with company number 05608868, with registered address at 10 Arvon Road,  
> > London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [September 10, 2010, 5:13pm UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251/11 "2010-09-10T17:13:40Z")

</div>

Hi,

We are definitely in need of a way to search a single field for content that  
can be in a variety of languages.

If that requires a product that needs to be licensed, then I am willing to  
go down that road.

Please feel free to contact me at jcook at tracermedia dot c-o-m.

On Thu, Sep 9, 2010 at 7:47 PM, Otis [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com) wrote:

> Hello,
> 
> I spotted this reference to Sematext's Multilingual Indexer (MI):
> 
> > A quick search shows there are some analyzers out there that have been  
> > developed for this problem. (i.e.  
> > [Cloud Monitoring Tools & Services | Sematext](http://www.sematext.com/products/multilingual-indexer/index.html)) In the  
> > docs there is a list of built in analyzers. Is it straightforward to  
> > include  
> > and configure other analyzers? Any pointers to docs?
> 
> Not sure if you are asking for MI docs or some other docs. MI comes  
> with good docs, but they are not public. Adding it to Solr is well  
> documented and easy to do. Let us know if you need it for Elastic  
> Search.
> 
> ## Otis
> 
> Sematext :: [http://sematext.com/](http://sematext.com/) :: Solr - Lucene - Nutch  
> Lucene ecosystem search :: [http://search-lucene.com/](http://search-lucene.com/)
> 
> On Aug 27, 8:57 am, James Cook [jc...@tracermedia.com](mailto:jc...@tracermedia.com) wrote:
> 
> > In my example, I have an object I am indexing which is similar to a  
> > discussion thread. The 'content' proeprty will contain text which may be  
> > in  
> > English or Arabic.
> > 
> > If the JSON document I am indexing can determine which language it is  
> > using,  
> > can an analyzer be chosen at index and search time?
> > 
> > I don't know much about mappings yet, but the multi-type approach worries  
> > me  
> > because the 'content' field will be knowingly indexed once with the  
> > correct  
> > analyzer and once with the incorrect analyzer.
> > 
> > It appears from the doc entry that the query is then performed only  
> > against  
> > the 'default' entry in the multi-type instead of applying against all  
> > multi-type entries. This makes it a bit harder to manage queries I think.  
> > If  
> > multi-type is the only way to be able to search for multilingual text in  
> > a  
> > field, I suppose I will have to adapt. 🙂
> > 
> > A quick search shows there are some analyzers out there that have been  
> > developed for this problem. (i.e.  
> > [Cloud Monitoring Tools & Services | Sematext](http://www.sematext.com/products/multilingual-indexer/index.html)) In the  
> > docs there is a list of built in analyzers. Is it straightforward to  
> > include  
> > and configure other analyzers? Any pointers to docs?
> > 
> > Thanks
> > 
> > On Fri, Aug 27, 2010 at 7:34 AM, Clinton Gormley \<  
> > [clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk)\>wrote:
> > 
> > > On Fri, 2010-08-27 at 13:33 +0300, Shay Banon wrote:
> > > 
> > > > No, you can't specify different analyzers on the same field.
> > 
> > > But you can index the same field twice, as a multi field, with  
> > > different  
> > > analysers:
> > 
> > > [http://www.elasticsearch.com/docs/elasticsearch/mapping/multi\_field/](http://www.elasticsearch.com/docs/elasticsearch/mapping/multi_field/)
> > 
> > > clint
> > 
> > > > On Thu, Aug 26, 2010 at 4:31 PM, James Cook [jc...@tracermedia.com](mailto:jc...@tracermedia.com)  
> > > > wrote:  
> > > > Circling around to my earlier question, can I have an English  
> > > > _and_ Arabic analyzer specified on the same fields across  
> > > > documents?
> > 
> > > > ```
> > > > On Wed, Aug 25, 2010 at 2:57 PM, Shay Banon
> > > > <shay.ba...@elasticsearch.com> wrote:
> > > > On Wed, Aug 25, 2010 at 9:13 PM, Andrei
> > > > <and...@zmievski.org> wrote:
> > > > 
> > > > ```
> > 
> > > > ```
> > > > On Aug 24, 4:45 pm, Shay Banon
> > > > <shay.ba...@elasticsearch.com> wrote:
> > > > > Yes, you can create your own analyzer and
> > > > add to it the asciifolding filter.
> > > > > The ICU plugin might also be interesting
> > > > 
> > > > ```
> 
> for
> 
> > > > ```
> > > > this.
> > > > 
> > > > ```
> > 
> > > > ```
> > > > Do you mean to create one in Java or in the
> > > > configuration file?
> > > > 
> > > > ```
> > 
> > > > ```
> > > > Its in a configuration file. You create a custom
> > > > analyzer that include it.
> > > > 
> > > > ```
> > 
> > > > ```
> > > > > It depends how far you want to take it.
> > > > There are specific analyzers for
> > > > > different languages. I updated the docs to
> > > > reflect that.
> > > > 
> > > > ```
> > 
> > > > ```
> > > > Could you link to the page that you updated?
> > > > 
> > > > ```
> 
> I
> 
> > > > ```
> > > > couldn't find the
> > > > references to non-English languages there.
> > > > 
> > > > ```
> > 
> > > > ```
> > > > Here it
> > > > is:
> > > > 
> > > > ```
> > > 
> > > [http://www.elasticsearch.com/docs/elasticsearch/index\_modules/analysi](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysi).  
> > > ..
> > 
> > > > ```
> > > > -Andrei
> > > > 
> > > > ```
> > 
> > > --  
> > > Web Announcements Limited is a company registered in England and Wales,  
> > > with company number 05608868, with registered address at 10 Arvon Road,  
> > > London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:19am UTC](https://discuss.elastic.co/t/indexing-non-english-text/3251/12 "2017-07-06T04:19:30Z")

</div>


