# Protect some words when tokenizing

**URL:** <https://discuss.elastic.co/t/protect-some-words-when-tokenizing/10710>\
**Category:** Elasticsearch\
**Created:** [February 12, 2013, 1:19pm UTC](https://discuss.elastic.co/t/protect-some-words-when-tokenizing/10710 "2013-02-12T13:19:57Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![pdesoyres](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pdesoyres/32/509_2.png) [@pdesoyres](https://discuss.elastic.co/u/pdesoyres)\
**Post date:** [February 12, 2013, 1:19pm UTC](https://discuss.elastic.co/t/protect-some-words-when-tokenizing/10710/1 "2013-02-12T13:19:57Z")

</div>

Hi,

I use for some fields the standard tokenizer and I would like to know if  
there is a way to prevent strings such as "c++", "c#" or ".net" to be  
tokenized as "c", "c" or "net" but to be kept unmodified.

Thanks in advance

Pierre

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![egaumer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/egaumer/32/2365_2.png) [@egaumer](https://discuss.elastic.co/u/egaumer)\
**Post date:** [February 12, 2013, 1:47pm UTC](https://discuss.elastic.co/t/protect-some-words-when-tokenizing/10710/2 "2013-02-12T13:47:23Z")

</div>

You could use a whitespace tokenizer instead to preserve punctuation on  
this field...

curl -XGET 'localhost:9200/\_analyze?tokenizer=whitespace&pretty=1' -d 'I  
write C++ code.'

On Tuesday, February 12, 2013 8:19:57 AM UTC-5, Pierre de Soyres wrote:

> Hi,
> 
> I use for some fields the standard tokenizer and I would like to know if  
> there is a way to prevent strings such as "c++", "c#" or ".net" to be  
> tokenized as "c", "c" or "net" but to be kept unmodified.
> 
> Thanks in advance
> 
> Pierre

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Pierre\_De\_Soyres](https://avatars.discourse-cdn.com/v4/letter/p/f14d63/32.png) [@Pierre\_De\_Soyres](https://discuss.elastic.co/u/Pierre_De_Soyres)\
**Post date:** [February 12, 2013, 2:09pm UTC](https://discuss.elastic.co/t/protect-some-words-when-tokenizing/10710/3 "2013-02-12T14:09:10Z")

</div>

Thank you for response,

but using 'whitespace' is not an option for me because I need comma, dot,  
dash, etc. to be delimiters as well

Pierre.

Le mardi 12 février 2013 14:47:23 UTC+1, egaumer a écrit :

> You could use a whitespace tokenizer instead to preserve punctuation on  
> this field...
> 
> curl -XGET 'localhost:9200/\_analyze?tokenizer=whitespace&pretty=1' -d 'I  
> write C++ code.'
> 
> On Tuesday, February 12, 2013 8:19:57 AM UTC-5, Pierre de Soyres wrote:
> 
> > Hi,
> > 
> > I use for some fields the standard tokenizer and I would like to know if  
> > there is a way to prevent strings such as "c++", "c#" or ".net" to be  
> > tokenized as "c", "c" or "net" but to be kept unmodified.
> > 
> > Thanks in advance
> > 
> > Pierre

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![egaumer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/egaumer/32/2365_2.png) [@egaumer](https://discuss.elastic.co/u/egaumer)\
**Post date:** [February 12, 2013, 3:11pm UTC](https://discuss.elastic.co/t/protect-some-words-when-tokenizing/10710/4 "2013-02-12T15:11:37Z")

</div>

You should be able to use a custom tuned word\_delimeter to clean up  
unwanted punctuation...

egaumer@ares:(src)$ curl -XPUT '[http://localhost:9200/test](http://localhost:9200/test)' -d '{  
"settings" : {  
"index" : {  
"number\_of\_shards" : 1,  
"number\_of\_replicas" : 1  
},  
"analysis" : {  
"filter" : {  
"my\_delimiter" : {  
"type" : "word\_delimiter",  
"split\_on\_numerics" : true,  
"split\_on\_case\_change" : true,  
"my\_delimiter.catenate\_numbers" : true,  
"generate\_word\_parts" : true,  
"protected\_words": ["C++", "C#"]  
}  
}  
}  
}  
}'

curl -XGET  
'localhost:9200/test/\_analyze?tokenizer=whitespace&filters=my\_delimiter&pretty=1'  
-d 'Hello, I write C++ code for wi-fi.'

Test that out and see if it does what you need. You can tweak other  
settings on the word\_delimeter to meet your needs.

On Tuesday, February 12, 2013 9:09:10 AM UTC-5, Pierre De Soyres wrote:

> Thank you for response,
> 
> but using 'whitespace' is not an option for me because I need comma, dot,  
> dash, etc. to be delimiters as well
> 
> Pierre.
> 
> Le mardi 12 février 2013 14:47:23 UTC+1, egaumer a écrit :
> 
> > You could use a whitespace tokenizer instead to preserve punctuation on  
> > this field...
> > 
> > curl -XGET 'localhost:9200/\_analyze?tokenizer=whitespace&pretty=1' -d 'I  
> > write C++ code.'
> > 
> > On Tuesday, February 12, 2013 8:19:57 AM UTC-5, Pierre de Soyres wrote:
> > 
> > > Hi,
> > > 
> > > I use for some fields the standard tokenizer and I would like to know if  
> > > there is a way to prevent strings such as "c++", "c#" or ".net" to be  
> > > tokenized as "c", "c" or "net" but to be kept unmodified.
> > > 
> > > Thanks in advance
> > > 
> > > Pierre

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Pierre\_De\_Soyres](https://avatars.discourse-cdn.com/v4/letter/p/f14d63/32.png) [@Pierre\_De\_Soyres](https://discuss.elastic.co/u/Pierre_De_Soyres)\
**Post date:** [February 12, 2013, 3:44pm UTC](https://discuss.elastic.co/t/protect-some-words-when-tokenizing/10710/5 "2013-02-12T15:44:14Z")

</div>

thank you, this fits my needs

Le mardi 12 février 2013 16:11:37 UTC+1, egaumer a écrit :

> You should be able to use a custom tuned word\_delimeter to clean up  
> unwanted punctuation...
> 
> egaumer@ares:(src)$ curl -XPUT '[http://localhost:9200/test](http://localhost:9200/test)' -d '{  
> "settings" : {  
> "index" : {  
> "number\_of\_shards" : 1,  
> "number\_of\_replicas" : 1  
> },  
> "analysis" : {  
> "filter" : {  
> "my\_delimiter" : {  
> "type" : "word\_delimiter",  
> "split\_on\_numerics" : true,  
> "split\_on\_case\_change" : true,  
> "my\_delimiter.catenate\_numbers" : true,  
> "generate\_word\_parts" : true,  
> "protected\_words": ["C++", "C#"]  
> }  
> }  
> }  
> }  
> }'
> 
> curl -XGET  
> 'localhost:9200/test/\_analyze?tokenizer=whitespace&filters=my\_delimiter&pretty=1'  
> -d 'Hello, I write C++ code for wi-fi.'
> 
> Test that out and see if it does what you need. You can tweak other  
> settings on the word\_delimeter to meet your needs.
> 
> On Tuesday, February 12, 2013 9:09:10 AM UTC-5, Pierre De Soyres wrote:
> 
> > Thank you for response,
> > 
> > but using 'whitespace' is not an option for me because I need comma, dot,  
> > dash, etc. to be delimiters as well
> > 
> > Pierre.
> > 
> > Le mardi 12 février 2013 14:47:23 UTC+1, egaumer a écrit :
> > 
> > > You could use a whitespace tokenizer instead to preserve punctuation on  
> > > this field...
> > > 
> > > curl -XGET 'localhost:9200/\_analyze?tokenizer=whitespace&pretty=1' -d 'I  
> > > write C++ code.'
> > > 
> > > On Tuesday, February 12, 2013 8:19:57 AM UTC-5, Pierre de Soyres wrote:
> > > 
> > > > Hi,
> > > > 
> > > > I use for some fields the standard tokenizer and I would like to know  
> > > > if there is a way to prevent strings such as "c++", "c#" or ".net" to be  
> > > > tokenized as "c", "c" or "net" but to be kept unmodified.
> > > > 
> > > > Thanks in advance
> > > > 
> > > > Pierre

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [February 15, 2013, 4:19pm UTC](https://discuss.elastic.co/t/protect-some-words-when-tokenizing/10710/6 "2013-02-15T16:19:23Z")

</div>

If you know the list of keywords to protect, you can also use a Keyword  
Marker Token Filter.

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

On Tue, Feb 12, 2013 at 7:44 AM, Pierre De Soyres \<  
[pierre.de-soyres@eptica.com](mailto:pierre.de-soyres@eptica.com)\> wrote:

> thank you, this fits my needs
> 
> Le mardi 12 février 2013 16:11:37 UTC+1, egaumer a écrit :
> 
> > You should be able to use a custom tuned word\_delimeter to clean up  
> > unwanted punctuation...
> > 
> > egaumer@ares:(src)$ curl -XPUT '[http://localhost:9200/test](http://localhost:9200/test)' -d '{  
> > "settings" : {  
> > "index" : {  
> > "number\_of\_shards" : 1,  
> > "number\_of\_replicas" : 1  
> > },  
> > "analysis" : {  
> > "filter" : {  
> > "my\_delimiter" : {  
> > "type" : "word\_delimiter",  
> > "split\_on\_numerics" : true,  
> > "split\_on\_case\_change" : true,  
> > "my\_delimiter.catenate\_\*\*numbers" : true,  
> > "generate\_word\_parts" : true,  
> > "protected\_words": ["C++", "C#"]  
> > }  
> > }  
> > }  
> > }  
> > }'
> > 
> > curl -XGET 'localhost:9200/test/\_analyze?\*_tokenizer=whitespace&filters=_  
> > \*my\_delimiter&pretty=1' -d 'Hello, I write C++ code for wi-fi.'
> > 
> > Test that out and see if it does what you need. You can tweak other  
> > settings on the word\_delimeter to meet your needs.
> > 
> > On Tuesday, February 12, 2013 9:09:10 AM UTC-5, Pierre De Soyres wrote:
> > 
> > > Thank you for response,
> > > 
> > > but using 'whitespace' is not an option for me because I need comma,  
> > > dot, dash, etc. to be delimiters as well
> > > 
> > > Pierre.
> > > 
> > > Le mardi 12 février 2013 14:47:23 UTC+1, egaumer a écrit :
> > > 
> > > > You could use a whitespace tokenizer instead to preserve punctuation on  
> > > > this field...
> > > > 
> > > > curl -XGET 'localhost:9200/\_analyze?\*\*tokenizer=whitespace&pretty=1'  
> > > > -d 'I write C++ code.'
> > > > 
> > > > On Tuesday, February 12, 2013 8:19:57 AM UTC-5, Pierre de Soyres wrote:
> > > > 
> > > > > Hi,
> > > > > 
> > > > > I use for some fields the standard tokenizer and I would like to know  
> > > > > if there is a way to prevent strings such as "c++", "c#" or ".net" to be  
> > > > > tokenized as "c", "c" or "net" but to be kept unmodified.
> > > > > 
> > > > > Thanks in advance
> > > > > 
> > > > > Pierre
> > > > 
> > > > --  
> > > > You received this message because you are subscribed to the Google Groups  
> > > > "elasticsearch" group.  
> > > > To unsubscribe from this group and stop receiving emails from it, send an  
> > > > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)\
**Post date:** [February 16, 2013, 3:15pm UTC](https://discuss.elastic.co/t/protect-some-words-when-tokenizing/10710/7 "2013-02-16T15:15:44Z")

</div>

Ivan, unfortunately the keywordMarkerFilter only works for in combination  
with stemmers.I added the keyword attribute years ago to prevent some  
stemmers from running the stemming alg on terms that are known to be names  
etc. I don't think this would help here.  
In general I would recommend to use a simple tokenizer like whitespace and  
then use synonym filter to transform these kind of token (c++ / c#) to a  
text represenations (cPLUSPLUS / CSHARP) then you can go wild with  
WordDelimiterFilter etc. once you did this mapping.

simon

On Friday, February 15, 2013 5:19:23 PM UTC+1, Ivan Brusic wrote:

> If you know the list of keywords to protect, you can also use a Keyword  
> Marker Token Filter.
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/keyword-marker-tokenfilter.html)
> 
> On Tue, Feb 12, 2013 at 7:44 AM, Pierre De Soyres \<[pierre.d...@eptica.com](mailto:pierre.d...@eptica.com)\<javascript:\>
> 
> > wrote:
> 
> > thank you, this fits my needs
> > 
> > Le mardi 12 février 2013 16:11:37 UTC+1, egaumer a écrit :
> > 
> > > You should be able to use a custom tuned word\_delimeter to clean up  
> > > unwanted punctuation...
> > > 
> > > egaumer@ares:(src)$ curl -XPUT '[http://localhost:9200/test](http://localhost:9200/test)' -d '{  
> > > "settings" : {  
> > > "index" : {  
> > > "number\_of\_shards" : 1,  
> > > "number\_of\_replicas" : 1  
> > > },  
> > > "analysis" : {  
> > > "filter" : {  
> > > "my\_delimiter" : {  
> > > "type" : "word\_delimiter",  
> > > "split\_on\_numerics" : true,  
> > > "split\_on\_case\_change" : true,  
> > > "my\_delimiter.catenate\_\*\*numbers" : true,  
> > > "generate\_word\_parts" : true,  
> > > "protected\_words": ["C++", "C#"]  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }'
> > > 
> > > curl -XGET 'localhost:9200/test/\_analyze?\*\*tokenizer=whitespace&filters=  
> > > \*\*my\_delimiter&pretty=1' -d 'Hello, I write C++ code for wi-fi.'
> > > 
> > > Test that out and see if it does what you need. You can tweak other  
> > > settings on the word\_delimeter to meet your needs.
> > > 
> > > On Tuesday, February 12, 2013 9:09:10 AM UTC-5, Pierre De Soyres wrote:
> > > 
> > > > Thank you for response,
> > > > 
> > > > but using 'whitespace' is not an option for me because I need comma,  
> > > > dot, dash, etc. to be delimiters as well
> > > > 
> > > > Pierre.
> > > > 
> > > > Le mardi 12 février 2013 14:47:23 UTC+1, egaumer a écrit :
> > > > 
> > > > > You could use a whitespace tokenizer instead to preserve punctuation  
> > > > > on this field...
> > > > > 
> > > > > curl -XGET 'localhost:9200/\_analyze?\*\*tokenizer=whitespace&pretty=1'  
> > > > > -d 'I write C++ code.'
> > > > > 
> > > > > On Tuesday, February 12, 2013 8:19:57 AM UTC-5, Pierre de Soyres wrote:
> > > > > 
> > > > > > Hi,
> > > > > > 
> > > > > > I use for some fields the standard tokenizer and I would like to know  
> > > > > > if there is a way to prevent strings such as "c++", "c#" or ".net" to be  
> > > > > > tokenized as "c", "c" or "net" but to be kept unmodified.
> > > > > > 
> > > > > > Thanks in advance
> > > > > > 
> > > > > > Pierre
> > > > > 
> > > > > --  
> > > > > You received this message because you are subscribed to the Google Groups  
> > > > > "elasticsearch" group.  
> > > > > To unsubscribe from this group and stop receiving emails from it, send an  
> > > > > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > > > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [February 18, 2013, 4:43pm UTC](https://discuss.elastic.co/t/protect-some-words-when-tokenizing/10710/8 "2013-02-18T16:43:33Z")

</div>

My mistake! I read the word "protect" and thought of the keyword marker  
filter. I once wrote a custom token filter on a Lucene project I was on,  
not related to stemming, that used the keyword attributes. Useful  
attribute, but it is post tokenization and not what the OP is looking for.

Nowadays in Lucene I use a pattern tokenizer since the whitespace tokenizer  
is too lenient, plus a word\_delimiter filter (and stemmer overrides).

--  
Ivan

On Sat, Feb 16, 2013 at 7:15 AM, simonw  
[simon.willnauer@elasticsearch.com](mailto:simon.willnauer@elasticsearch.com)wrote:

> Ivan, unfortunately the keywordMarkerFilter only works for in combination  
> with stemmers.I added the keyword attribute years ago to prevent some  
> stemmers from running the stemming alg on terms that are known to be names  
> etc. I don't think this would help here.  
> In general I would recommend to use a simple tokenizer like whitespace and  
> then use synonym filter to transform these kind of token (c++ / c#) to a  
> text represenations (cPLUSPLUS / CSHARP) then you can go wild with  
> WordDelimiterFilter etc. once you did this mapping.
> 
> simon
> 
> On Friday, February 15, 2013 5:19:23 PM UTC+1, Ivan Brusic wrote:
> 
> > If you know the list of keywords to protect, you can also use a Keyword  
> > Marker Token Filter.
> > 
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/**guide/reference/index-modules/)\*\*  
> > analysis/keyword-marker-\*\*tokenfilter.html[http://www.elasticsearch.org/guide/reference/index-modules/analysis/keyword-marker-tokenfilter.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/keyword-marker-tokenfilter.html)
> > 
> > On Tue, Feb 12, 2013 at 7:44 AM, Pierre De Soyres \<[pierre.d...@eptica.com](mailto:pierre.d...@eptica.com)
> > 
> > > wrote:
> > 
> > > thank you, this fits my needs
> > > 
> > > Le mardi 12 février 2013 16:11:37 UTC+1, egaumer a écrit :
> > > 
> > > > You should be able to use a custom tuned word\_delimeter to clean up  
> > > > unwanted punctuation...
> > > > 
> > > > egaumer@ares:(src)$ curl -XPUT '[http://localhost:9200/test](http://localhost:9200/test)' -d '{  
> > > > "settings" : {  
> > > > "index" : {  
> > > > "number\_of\_shards" : 1,  
> > > > "number\_of\_replicas" : 1  
> > > > },  
> > > > "analysis" : {  
> > > > "filter" : {  
> > > > "my\_delimiter" : {  
> > > > "type" : "word\_delimiter",  
> > > > "split\_on\_numerics" : true,  
> > > > "split\_on\_case\_change" : true,  
> > > > "my\_delimiter.catenate\_ **numbers**" : true,  
> > > > "generate\_word\_parts" : true,  
> > > > "protected\_words": ["C++", "C#"]  
> > > > }  
> > > > }  
> > > > }  
> > > > }  
> > > > }'
> > > > 
> > > > curl -XGET 'localhost:9200/test/\_analyze?\*\*\*\*  
> > > > tokenizer=whitespace&filters= **m** y\_delimiter&pretty=1' -d 'Hello, I  
> > > > write C++ code for wi-fi.'
> > > > 
> > > > Test that out and see if it does what you need. You can tweak other  
> > > > settings on the word\_delimeter to meet your needs.
> > > > 
> > > > On Tuesday, February 12, 2013 9:09:10 AM UTC-5, Pierre De Soyres wrote:
> > > > 
> > > > > Thank you for response,
> > > > > 
> > > > > but using 'whitespace' is not an option for me because I need comma,  
> > > > > dot, dash, etc. to be delimiters as well
> > > > > 
> > > > > Pierre.
> > > > > 
> > > > > Le mardi 12 février 2013 14:47:23 UTC+1, egaumer a écrit :
> > > > > 
> > > > > > You could use a whitespace tokenizer instead to preserve punctuation  
> > > > > > on this field...
> > > > > > 
> > > > > > curl -XGET 'localhost:9200/\_analyze? **token** izer=whitespace&pretty=1'  
> > > > > > -d 'I write C++ code.'
> > > > > > 
> > > > > > On Tuesday, February 12, 2013 8:19:57 AM UTC-5, Pierre de Soyres  
> > > > > > wrote:
> > > > > > 
> > > > > > > Hi,
> > > > > > > 
> > > > > > > I use for some fields the standard tokenizer and I would like to  
> > > > > > > know if there is a way to prevent strings such as "c++", "c#" or ".net" to  
> > > > > > > be tokenized as "c", "c" or "net" but to be kept unmodified.
> > > > > > > 
> > > > > > > Thanks in advance
> > > > > > > 
> > > > > > > Pierre
> > > > > > 
> > > > > > --  
> > > > > > You received this message because you are subscribed to the Google  
> > > > > > Groups "elasticsearch" group.  
> > > > > > To unsubscribe from this group and stop receiving emails from it, send  
> > > > > > an email to elasticsearc...@\*\*[googlegroups.com](http://googlegroups.com).
> > > 
> > > For more options, visit [https://groups.google.com/\*\*groups/opt\_out](https://groups.google.com/**groups/opt_out)[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > > .
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:50am UTC](https://discuss.elastic.co/t/protect-some-words-when-tokenizing/10710/9 "2017-07-06T02:50:54Z")

</div>


