# Analizer with stop words removal by language

**URL:** <https://discuss.elastic.co/t/analizer-with-stop-words-removal-by-language/3531>\
**Category:** Elasticsearch\
**Created:** [November 7, 2010, 5:34am UTC](https://discuss.elastic.co/t/analizer-with-stop-words-removal-by-language/3531 "2010-11-07T05:34:19Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sebastian\_Gavarini](https://avatars.discourse-cdn.com/v4/letter/s/db5fbb/32.png) [@Sebastian\_Gavarini](https://discuss.elastic.co/u/Sebastian_Gavarini)\
**Post date:** [November 7, 2010, 5:34am UTC](https://discuss.elastic.co/t/analizer-with-stop-words-removal-by-language/3531/1 "2010-11-07T05:34:19Z")

</div>

Hi all,

I am facing an issue with stop words removal for one of my fields. I  
have like 10 to 15 fields that are analyzed without the stop filter,  
but I have a long field called "description", that needs stop words  
removal.  
The problem I have is that I need a solution for many languages, I  
added in elasticsearch.yml definitions for my analyzers, for example  
"default", "en\_stop\_analyzer" "es\_stop\_analyzer", ...  
(en, es being English and Spanish).  
So far so good, I have also some custom mappings explicitly defined,  
and the dynamic features off.  
The problem is I can't use for my types the "analyzer" setting in the  
JSON mapping, because I use the same mapping for all the languages,  
let's say I have a mapping for "myDocument", which has a field  
"description", that I know must be indexed with stop filter, but I  
will only know the language at indexing time.

I could create many mappings, one for each language, but I have 5  
different object types already, multiplied by the languages I must  
support, it's not nice to maintain for just the stop word list of a  
single field.  
Is there a way to use at least some "include/import" feature to  
minimize the differences among files?  
Could the analyzer be passed with the index and bulk apis?  
any other ideas?

Thanks,  
Sebastian.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 7, 2010, 11:57am UTC](https://discuss.elastic.co/t/analizer-with-stop-words-removal-by-language/3531/2 "2010-11-07T11:57:17Z")

</div>

Hi,

Yea, one of the features on my list is to have the ability to drive the  
analyzer used based on a field in the json doc. I got most of it implemented  
on a local branch, open a feature for it?

-shay.bnaon

On Sun, Nov 7, 2010 at 7:34 AM, Sebastian [sgavarini@gmail.com](mailto:sgavarini@gmail.com) wrote:

> Hi all,
> 
> I am facing an issue with stop words removal for one of my fields. I  
> have like 10 to 15 fields that are analyzed without the stop filter,  
> but I have a long field called "description", that needs stop words  
> removal.  
> The problem I have is that I need a solution for many languages, I  
> added in elasticsearch.yml definitions for my analyzers, for example  
> "default", "en\_stop\_analyzer" "es\_stop\_analyzer", ...  
> (en, es being English and Spanish).  
> So far so good, I have also some custom mappings explicitly defined,  
> and the dynamic features off.  
> The problem is I can't use for my types the "analyzer" setting in the  
> JSON mapping, because I use the same mapping for all the languages,  
> let's say I have a mapping for "myDocument", which has a field  
> "description", that I know must be indexed with stop filter, but I  
> will only know the language at indexing time.
> 
> I could create many mappings, one for each language, but I have 5  
> different object types already, multiplied by the languages I must  
> support, it's not nice to maintain for just the stop word list of a  
> single field.  
> Is there a way to use at least some "include/import" feature to  
> minimize the differences among files?  
> Could the analyzer be passed with the index and bulk apis?  
> any other ideas?
> 
> Thanks,  
> Sebastian.

---

<div class="post-metadata">

**Author:** ![Sebastian\_Gavarini](https://avatars.discourse-cdn.com/v4/letter/s/db5fbb/32.png) [@Sebastian\_Gavarini](https://discuss.elastic.co/u/Sebastian_Gavarini)\
**Post date:** [November 7, 2010, 6:28pm UTC](https://discuss.elastic.co/t/analizer-with-stop-words-removal-by-language/3531/3 "2010-11-07T18:28:52Z")

</div>

Hi Shay,

I think that is a very good idea.

Sure, I have just opened it: [Issues · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/#issue/487)

I don't want to rush you, but I would like to know more or less when  
do you expect that to be implemented? For now I can go with my plan to  
create many files, maybe with a template generation.

Thanks,  
Sebastian.

On Nov 7, 8:57 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Hi,
> 
> Yea, one of the features on my list is to have the ability to drive the  
> analyzer used based on a field in the json doc. I got most of it implemented  
> on a local branch, open a feature for it?
> 
> -shay.bnaon
> 
> On Sun, Nov 7, 2010 at 7:34 AM, Sebastian [sgavar...@gmail.com](mailto:sgavar...@gmail.com) wrote:
> 
> > Hi all,
> 
> > I am facing an issue with stop words removal for one of my fields. I  
> > have like 10 to 15 fields that are analyzed without the stop filter,  
> > but I have a long field called "description", that needs stop words  
> > removal.  
> > The problem I have is that I need a solution for many languages, I  
> > added in elasticsearch.yml definitions for my analyzers, for example  
> > "default", "en\_stop\_analyzer" "es\_stop\_analyzer", ...  
> > (en, es being English and Spanish).  
> > So far so good, I have also some custom mappings explicitly defined,  
> > and the dynamic features off.  
> > The problem is I can't use for my types the "analyzer" setting in the  
> > JSON mapping, because I use the same mapping for all the languages,  
> > let's say I have a mapping for "myDocument", which has a field  
> > "description", that I know must be indexed with stop filter, but I  
> > will only know the language at indexing time.
> 
> > I could create many mappings, one for each language, but I have 5  
> > different object types already, multiplied by the languages I must  
> > support, it's not nice to maintain for just the stop word list of a  
> > single field.  
> > Is there a way to use at least some "include/import" feature to  
> > minimize the differences among files?  
> > Could the analyzer be passed with the index and bulk apis?  
> > any other ideas?
> 
> > Thanks,  
> > Sebastian.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 7, 2010, 7:02pm UTC](https://discuss.elastic.co/t/analizer-with-stop-words-removal-by-language/3531/4 "2010-11-07T19:02:59Z")

</div>

Already implemented and pushed to master:  
[Issues · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/closed#issue/485).

On Sun, Nov 7, 2010 at 8:28 PM, Sebastian [sgavarini@gmail.com](mailto:sgavarini@gmail.com) wrote:

> Hi Shay,
> 
> I think that is a very good idea.
> 
> Sure, I have just opened it:  
> [Issues · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/#issue/487)
> 
> I don't want to rush you, but I would like to know more or less when  
> do you expect that to be implemented? For now I can go with my plan to  
> create many files, maybe with a template generation.
> 
> Thanks,  
> Sebastian.
> 
> On Nov 7, 8:57 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Hi,
> > 
> > Yea, one of the features on my list is to have the ability to drive the  
> > analyzer used based on a field in the json doc. I got most of it  
> > implemented  
> > on a local branch, open a feature for it?
> > 
> > -shay.bnaon
> > 
> > On Sun, Nov 7, 2010 at 7:34 AM, Sebastian [sgavar...@gmail.com](mailto:sgavar...@gmail.com) wrote:
> > 
> > > Hi all,
> > 
> > > I am facing an issue with stop words removal for one of my fields. I  
> > > have like 10 to 15 fields that are analyzed without the stop filter,  
> > > but I have a long field called "description", that needs stop words  
> > > removal.  
> > > The problem I have is that I need a solution for many languages, I  
> > > added in elasticsearch.yml definitions for my analyzers, for example  
> > > "default", "en\_stop\_analyzer" "es\_stop\_analyzer", ...  
> > > (en, es being English and Spanish).  
> > > So far so good, I have also some custom mappings explicitly defined,  
> > > and the dynamic features off.  
> > > The problem is I can't use for my types the "analyzer" setting in the  
> > > JSON mapping, because I use the same mapping for all the languages,  
> > > let's say I have a mapping for "myDocument", which has a field  
> > > "description", that I know must be indexed with stop filter, but I  
> > > will only know the language at indexing time.
> > 
> > > I could create many mappings, one for each language, but I have 5  
> > > different object types already, multiplied by the languages I must  
> > > support, it's not nice to maintain for just the stop word list of a  
> > > single field.  
> > > Is there a way to use at least some "include/import" feature to  
> > > minimize the differences among files?  
> > > Could the analyzer be passed with the index and bulk apis?  
> > > any other ideas?
> > 
> > > Thanks,  
> > > Sebastian.

---

<div class="post-metadata">

**Author:** ![Sebastian\_Gavarini](https://avatars.discourse-cdn.com/v4/letter/s/db5fbb/32.png) [@Sebastian\_Gavarini](https://discuss.elastic.co/u/Sebastian_Gavarini)\
**Post date:** [November 9, 2010, 3:41am UTC](https://discuss.elastic.co/t/analizer-with-stop-words-removal-by-language/3531/5 "2010-11-09T03:41:26Z")

</div>

Hi Shay,

I posted an update in issue [Mapper: An analyzer mapper allowing to control the index analyzer of a document based on a document field · Issue #485 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/issue/485)

It isn't working for me, I followed the example but the analyzer field  
is not found by AnalyzerMapper, line 85.  
The document where the mapper tries to find the analyzer field doesn't  
contain it yet.

Sebastian.

On Nov 7, 4:02 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Already implemented and pushed to master:[Issues · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/closed#issue/485).
> 
> On Sun, Nov 7, 2010 at 8:28 PM, Sebastian [sgavar...@gmail.com](mailto:sgavar...@gmail.com) wrote:
> 
> > Hi Shay,
> 
> > I think that is a very good idea.
> 
> > Sure, I have just opened it:  
> > [Issues · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/#issue/487)
> 
> > I don't want to rush you, but I would like to know more or less when  
> > do you expect that to be implemented? For now I can go with my plan to  
> > create many files, maybe with a template generation.
> 
> > Thanks,  
> > Sebastian.
> 
> > On Nov 7, 8:57 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > Hi,
> 
> > > Yea, one of the features on my list is to have the ability to drive the  
> > > analyzer used based on a field in the json doc. I got most of it  
> > > implemented  
> > > on a local branch, open a feature for it?
> 
> > > -shay.bnaon
> 
> > > On Sun, Nov 7, 2010 at 7:34 AM, Sebastian [sgavar...@gmail.com](mailto:sgavar...@gmail.com) wrote:
> > > 
> > > > Hi all,
> 
> > > > I am facing an issue with stop words removal for one of my fields. I  
> > > > have like 10 to 15 fields that are analyzed without the stop filter,  
> > > > but I have a long field called "description", that needs stop words  
> > > > removal.  
> > > > The problem I have is that I need a solution for many languages, I  
> > > > added in elasticsearch.yml definitions for my analyzers, for example  
> > > > "default", "en\_stop\_analyzer" "es\_stop\_analyzer", ...  
> > > > (en, es being English and Spanish).  
> > > > So far so good, I have also some custom mappings explicitly defined,  
> > > > and the dynamic features off.  
> > > > The problem is I can't use for my types the "analyzer" setting in the  
> > > > JSON mapping, because I use the same mapping for all the languages,  
> > > > let's say I have a mapping for "myDocument", which has a field  
> > > > "description", that I know must be indexed with stop filter, but I  
> > > > will only know the language at indexing time.
> 
> > > > I could create many mappings, one for each language, but I have 5  
> > > > different object types already, multiplied by the languages I must  
> > > > support, it's not nice to maintain for just the stop word list of a  
> > > > single field.  
> > > > Is there a way to use at least some "include/import" feature to  
> > > > minimize the differences among files?  
> > > > Could the analyzer be passed with the index and bulk apis?  
> > > > any other ideas?
> 
> > > > Thanks,  
> > > > Sebastian.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:16am UTC](https://discuss.elastic.co/t/analizer-with-stop-words-removal-by-language/3531/6 "2017-07-06T04:16:54Z")

</div>


