# Can language analyzers be configured to use char\_filters and token\_filters?

**URL:** <https://discuss.elastic.co/t/can-language-analyzers-be-configured-to-use-char-filters-and-token-filters/7929>\
**Category:** Elasticsearch\
**Created:** [May 31, 2012, 2:13pm UTC](https://discuss.elastic.co/t/can-language-analyzers-be-configured-to-use-char-filters-and-token-filters/7929 "2012-05-31T14:13:57Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Robin\_Hughes](https://avatars.discourse-cdn.com/v4/letter/r/e36b37/32.png) [@Robin\_Hughes](https://discuss.elastic.co/u/Robin_Hughes)\
**Post date:** [May 31, 2012, 2:13pm UTC](https://discuss.elastic.co/t/can-language-analyzers-be-configured-to-use-char-filters-and-token-filters/7929/1 "2012-05-31T14:13:57Z")

</div>

Hi

I have documents in many languages containing basic html, that need to be  
searched in a case insensitive, ascii-folded manner.

Is it possible to use the standard language analyzers from  
[http://www.elasticsearch.org/guide/reference/index-modules/analysis/lang-analyzer.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/lang-analyzer.html)  
(in addition to plugins such as the smart chinese and stempel analyzers) in  
conjunction with the html\_strip char\_filter, lowercase and asciifolding  
token\_filters?

As far as I can tell this isn't possible by config alone, but would love to  
be proved wrong.

Thanks,  
Robin

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [June 1, 2012, 6:53am UTC](https://discuss.elastic.co/t/can-language-analyzers-be-configured-to-use-char-filters-and-token-filters/7929/2 "2012-06-01T06:53:26Z")

</div>

Hi Robin,

You can always re-create the analyzer from scratch using a custom  
analyzer. Language analyzers are analyzers with a language specific  
stemmer filter. Not hard to do in Elasticsearch.

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

I have never used a language analyzer, but I would assume it does  
lowercase and asciifolding already. At least the former.

Ivan

On Thu, May 31, 2012 at 7:13 AM, Robin Hughes [robinhughes@fastmail.fm](mailto:robinhughes@fastmail.fm) wrote:

> Hi
> 
> I have documents in many languages containing basic html, that need to be  
> searched in a case insensitive, ascii-folded manner.
> 
> Is it possible to use the standard language analyzers from  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/lang-analyzer.html)  
> (in addition to plugins such as the smart chinese and stempel analyzers) in  
> conjunction with the html\_strip char\_filter, lowercase and asciifolding  
> token\_filters?
> 
> As far as I can tell this isn't possible by config alone, but would love to  
> be proved wrong.
> 
> Thanks,  
> Robin

---

<div class="post-metadata">

**Author:** ![Robin\_Hughes](https://avatars.discourse-cdn.com/v4/letter/r/e36b37/32.png) [@Robin\_Hughes](https://discuss.elastic.co/u/Robin_Hughes)\
**Post date:** [June 1, 2012, 3:34pm UTC](https://discuss.elastic.co/t/can-language-analyzers-be-configured-to-use-char-filters-and-token-filters/7929/3 "2012-06-01T15:34:09Z")

</div>

Thanks for your help.

That certainly covers a lot of languages. It looks like some (Polish, Smart  
Chinese, Thai) will need a bit of extra work.

Thanks again,

Robin.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:26am UTC](https://discuss.elastic.co/t/can-language-analyzers-be-configured-to-use-char-filters-and-token-filters/7929/4 "2017-07-06T03:26:02Z")

</div>


