# Language and HTML analyzer

**URL:** <https://discuss.elastic.co/t/language-and-html-analyzer/47059>\
**Category:** Elasticsearch\
**Created:** [April 12, 2016, 1:44am UTC](https://discuss.elastic.co/t/language-and-html-analyzer/47059 "2016-04-12T01:44:44Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Karthik\_Ramachandran](https://avatars.discourse-cdn.com/v4/letter/k/f19dbf/32.png) [@Karthik\_Ramachandran](https://discuss.elastic.co/u/Karthik_Ramachandran)\
**Post date:** [April 12, 2016, 1:44am UTC](https://discuss.elastic.co/t/language-and-html-analyzer/47059/1 "2016-04-12T01:44:44Z")

</div>

I need more clarity on language analyzer and html filtering. My content sometimes come within html tags, that I need to strip out during indexing. Also, it varies by language. I create mapping for each language and have to use appropriate analyzer. How do I combine these?

For ex. I get English Content with or without HTML tags, I get Spanish Content with or without HTML tags. I need to index only the actual content. I also assume language specific analyzer do consider English tokens by default. Because, my content do contain English sentences though classified to be some other language...

Thanks for help..

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [April 12, 2016, 7:00am UTC](https://discuss.elastic.co/t/language-and-html-analyzer/47059/2 "2016-04-12T07:00:07Z")

</div>

It sounds like you need to do a bit of filtering before hand and send different languages into different indices, with different analysers.

Once they are in (eg) english and spanish language indices you can then just run your analysers.

---

<div class="post-metadata">

**Author:** ![Karthik\_Ramachandran](https://avatars.discourse-cdn.com/v4/letter/k/f19dbf/32.png) [@Karthik\_Ramachandran](https://discuss.elastic.co/u/Karthik_Ramachandran)\
**Post date:** [April 12, 2016, 7:42pm UTC](https://discuss.elastic.co/t/language-and-html-analyzer/47059/3 "2016-04-12T19:42:48Z")

</div>

Thanks Mark.

Should I send different languages to different Indices? I thought of using type - mappings for each language withing same index? Can't I have the analyzers with types?

Also, w.r.t html tags, i thought of using html\_strip charfilter. Won't it help.  
URL Referred: [http://stackoverflow.com/questions/18780346/html-strip-in-elastic-search](http://stackoverflow.com/questions/18780346/html-strip-in-elastic-search)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [April 12, 2016, 10:09pm UTC](https://discuss.elastic.co/t/language-and-html-analyzer/47059/4 "2016-04-12T22:09:43Z")

</div>

> [@Karthik\_Ramachandran](#):
>
> Should I send different languages to different Indices? I thought of using type - mappings for each language withing same index?

I would, it just keeps the logical domains cleaner and lets you play with analysis on a per language basis.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:59pm UTC](https://discuss.elastic.co/t/language-and-html-analyzer/47059/5 "2017-07-05T22:59:49Z")

</div>


