# N edge gram analyzer's behave not as expected

**URL:** <https://discuss.elastic.co/t/n-edge-gram-analyzers-behave-not-as-expected/22884>\
**Category:** Elasticsearch\
**Created:** [March 25, 2015, 8:49am UTC](https://discuss.elastic.co/t/n-edge-gram-analyzers-behave-not-as-expected/22884 "2015-03-25T08:49:14Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![narinder\_izap](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/narinder_izap/32/803_2.png) [@narinder\_izap](https://discuss.elastic.co/u/narinder_izap)\
**Post date:** [March 25, 2015, 8:49am UTC](https://discuss.elastic.co/t/n-edge-gram-analyzers-behave-not-as-expected/22884/1 "2015-03-25T08:49:14Z")

</div>

Hi All,

I have an custom analyzer based on n edge gram analyzer, The expectation  
was it should analyze the text based on edges. So as per my understanding,  
the analysis of a multi word like (Narinder Kaur)term will give  
N  
Na  
Nar  
Nari  
Narin  
Narind  
Narinde  
Narinder  
K  
Ka  
Kau  
Kaur

So now if search for narinder or kaur by using the following query:

{  
"query": {  
"constant\_score": {  
"query": {  
"match\_phrase\_prefix": {  
"primary\_search\_new": {  
"query": "narinder",  
"analyzer": "ys\_search\_analyzer\_long"  
}  
}  
}  
}  
}  
}

OR

{  
"query": {  
"constant\_score": {  
"query": {  
"match\_phrase\_prefix": {  
"primary\_search\_new": {  
"query": "kaur",  
"analyzer": "ys\_search\_analyzer\_long"  
}  
}  
}  
}  
}  
}

both should have searched for the documents containing "Narinder Kaur". But  
currently I can not search for kaur. Its working only for first term match.  
The analyzer's used are as followed:

analysis: {  
analyzer: {  
ys\_search\_analyzer: {  
type: custom  
filter: [  
ys\_word\_delimiter  
trim  
lowercase  
]  
tokenizer: ys\_edge\_ngram\_tokenizer  
}  
ys\_search\_analyzer\_long: {  
type: custom  
filter: [  
ys\_word\_delimiter  
trim  
lowercase  
]  
tokenizer: ys\_edge\_ngram\_tokenizer\_long  
}  
}  
filter: {  
ys\_word\_delimiter: {  
type: word\_delimiter  
stem\_english\_possessive: False  
}  
}  
tokenizer: {  
ys\_edge\_ngram\_tokenizer\_long: {  
type: edgeNGram  
min\_gram: 1  
max\_gram: 60  
}  
ys\_edge\_ngram\_tokenizer: {  
min\_gram: 1  
type: edgeNGram  
max\_gram: 20  
}  
}  
}

Please elaborate how its not working as expected? and what should I do to  
make my requirement work without re-indexing the data.

All help is appreciated.  
thanks

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/76753cd7-7a47-4ca3-ba7b-90be025386b4%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/76753cd7-7a47-4ca3-ba7b-90be025386b4%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![masaru](https://avatars.discourse-cdn.com/v4/letter/m/a9a28c/32.png) [@masaru](https://discuss.elastic.co/u/masaru)\
**Post date:** [March 26, 2015, 11:36pm UTC](https://discuss.elastic.co/t/n-edge-gram-analyzers-behave-not-as-expected/22884/2 "2015-03-26T23:36:15Z")

</div>

Hi,

You'd need to specify token\_chars when you configure edge ngram tokenizer([http://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-edgengram-tokenizer.html](http://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-edgengram-tokenizer.html)). Unless, all characters are kept. Which means, words are not split on white spaces.  
You can see how the analyzer works by \_analyze API([http://www.elastic.co/guide/en/elasticsearch/reference/current/indices-analyze.html](http://www.elastic.co/guide/en/elasticsearch/reference/current/indices-analyze.html))

You need to fix analyzer and re-index all documents.

Masaru

On March 25, 2015 at 17:49:24, Narinder Kaur (narinder.kaur@izap.in) wrote:

Hi All,

I have an custom analyzer based on n edge gram analyzer, The expectation was it should analyze the text based on edges. So as per my understanding, the analysis of a multi word like (Narinder Kaur)term will give  
N  
Na  
Nar  
Nari  
Narin  
Narind  
Narinde  
Narinder  
K  
Ka  
Kau  
Kaur

So now if search for narinder or kaur by using the following query:

{  
"query": {  
"constant\_score": {  
"query": {  
"match\_phrase\_prefix": {  
"primary\_search\_new": {  
"query": "narinder",  
"analyzer": "ys\_search\_analyzer\_long"  
}  
}  
}  
}  
}  
}

OR

{  
"query": {  
"constant\_score": {  
"query": {  
"match\_phrase\_prefix": {  
"primary\_search\_new": {  
"query": "kaur",  
"analyzer": "ys\_search\_analyzer\_long"  
}  
}  
}  
}  
}  
}

both should have searched for the documents containing "Narinder Kaur". But currently I can not search for kaur. Its working only for first term match. The analyzer's used are as followed:

analysis: {  
analyzer: {  
ys\_search\_analyzer: {  
type: custom  
filter: [  
ys\_word\_delimiter  
trim  
lowercase  
]  
tokenizer: ys\_edge\_ngram\_tokenizer  
}  
ys\_search\_analyzer\_long: {  
type: custom  
filter: [  
ys\_word\_delimiter  
trim  
lowercase  
]  
tokenizer: ys\_edge\_ngram\_tokenizer\_long  
}  
}  
filter: {  
ys\_word\_delimiter: {  
type: word\_delimiter  
stem\_english\_possessive: False  
}  
}  
tokenizer: {  
ys\_edge\_ngram\_tokenizer\_long: {  
type: edgeNGram  
min\_gram: 1  
max\_gram: 60  
}  
ys\_edge\_ngram\_tokenizer: {  
min\_gram: 1  
type: edgeNGram  
max\_gram: 20  
}  
}  
}

Please elaborate how its not working as expected? and what should I do to make my requirement work without re-indexing the data.

## All help is appreciated. thanks

You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/76753cd7-7a47-4ca3-ba7b-90be025386b4%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/76753cd7-7a47-4ca3-ba7b-90be025386b4%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/etPan.5514982b.79e2a9e3.166%40citra-2.local](https://groups.google.com/d/msgid/elasticsearch/etPan.5514982b.79e2a9e3.166%40citra-2.local).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![narinder\_izap](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/narinder_izap/32/803_2.png) [@narinder\_izap](https://discuss.elastic.co/u/narinder_izap)\
**Post date:** [March 27, 2015, 4:38am UTC](https://discuss.elastic.co/t/n-edge-gram-analyzers-behave-not-as-expected/22884/3 "2015-03-27T04:38:34Z")

</div>

Thanks for your reply. It much better clear now how to

On Friday, 27 March 2015 05:07:34 UTC+5:30, Masaru Hasegawa wrote:

> Hi,
> 
> You'd need to specify token\_chars when you configure edge ngram tokenizer(  
> [Edge n-gram tokenizer | Elasticsearch Guide [8.11] | Elastic](http://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-edgengram-tokenizer.html)).  
> Unless, all characters are kept. Which means, words are not split on white  
> spaces.  
> You can see how the analyzer works by \_analyze API(  
> [Analyze API | Elasticsearch Guide [8.11] | Elastic](http://www.elastic.co/guide/en/elasticsearch/reference/current/indices-analyze.html)  
> )
> 
> You need to fix analyzer and re-index all documents.
> 
> Masaru
> 
> On March 25, 2015 at 17:49:24, Narinder Kaur (narind...@izap.in  
> \<javascript:\>) wrote:
> 
> Hi All,
> 
> I have an custom analyzer based on n edge gram analyzer, The expectation  
> was it should analyze the text based on edges. So as per my understanding,  
> the analysis of a multi word like (Narinder Kaur)term will give  
> N  
> Na  
> Nar  
> Nari  
> Narin  
> Narind  
> Narinde  
> Narinder  
> K  
> Ka  
> Kau  
> Kaur
> 
> So now if search for narinder or kaur by using the following query:
> 
> {  
> "query": {  
> "constant\_score": {  
> "query": {  
> "match\_phrase\_prefix": {  
> "primary\_search\_new": {  
> "query": "narinder",  
> "analyzer": "ys\_search\_analyzer\_long"  
> }  
> }  
> }  
> }  
> }  
> }
> 
> OR
> 
> {  
> "query": {  
> "constant\_score": {  
> "query": {  
> "match\_phrase\_prefix": {  
> "primary\_search\_new": {  
> "query": "kaur",  
> "analyzer": "ys\_search\_analyzer\_long"  
> }  
> }  
> }  
> }  
> }  
> }
> 
> both should have searched for the documents containing "Narinder Kaur".  
> But currently I can not search for kaur. Its working only for first term  
> match. The analyzer's used are as followed:
> 
> analysis: {  
> analyzer: {  
> ys\_search\_analyzer: {  
> type: custom  
> filter: [  
> ys\_word\_delimiter  
> trim  
> lowercase  
> ]  
> tokenizer: ys\_edge\_ngram\_tokenizer  
> }  
> ys\_search\_analyzer\_long: {  
> type: custom  
> filter: [  
> ys\_word\_delimiter  
> trim  
> lowercase  
> ]  
> tokenizer: ys\_edge\_ngram\_tokenizer\_long  
> }  
> }  
> filter: {  
> ys\_word\_delimiter: {  
> type: word\_delimiter  
> stem\_english\_possessive: False  
> }  
> }  
> tokenizer: {  
> ys\_edge\_ngram\_tokenizer\_long: {  
> type: edgeNGram  
> min\_gram: 1  
> max\_gram: 60  
> }  
> ys\_edge\_ngram\_tokenizer: {  
> min\_gram: 1  
> type: edgeNGram  
> max\_gram: 20  
> }  
> }  
> }
> 
> Please elaborate how its not working as expected? and what should I do to  
> make my requirement work without re-indexing the data.
> 
> ## All help is appreciated. thanks
> 
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/76753cd7-7a47-4ca3-ba7b-90be025386b4%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/76753cd7-7a47-4ca3-ba7b-90be025386b4%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/76753cd7-7a47-4ca3-ba7b-90be025386b4%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/76753cd7-7a47-4ca3-ba7b-90be025386b4%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/54f7657f-ecb2-459f-8947-913a678745b0%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/54f7657f-ecb2-459f-8947-913a678745b0%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![narinder\_izap](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/narinder_izap/32/803_2.png) [@narinder\_izap](https://discuss.elastic.co/u/narinder_izap)\
**Post date:** [March 27, 2015, 4:39am UTC](https://discuss.elastic.co/t/n-edge-gram-analyzers-behave-not-as-expected/22884/4 "2015-03-27T04:39:25Z")

</div>

thanks for reply. I will try it.

On Friday, 27 March 2015 05:07:34 UTC+5:30, Masaru Hasegawa wrote:

> Hi,
> 
> You'd need to specify token\_chars when you configure edge ngram tokenizer(  
> [Edge n-gram tokenizer | Elasticsearch Guide [8.11] | Elastic](http://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-edgengram-tokenizer.html)).  
> Unless, all characters are kept. Which means, words are not split on white  
> spaces.  
> You can see how the analyzer works by \_analyze API(  
> [Analyze API | Elasticsearch Guide [8.11] | Elastic](http://www.elastic.co/guide/en/elasticsearch/reference/current/indices-analyze.html)  
> )
> 
> You need to fix analyzer and re-index all documents.
> 
> Masaru
> 
> On March 25, 2015 at 17:49:24, Narinder Kaur (narind...@izap.in  
> \<javascript:\>) wrote:
> 
> Hi All,
> 
> I have an custom analyzer based on n edge gram analyzer, The expectation  
> was it should analyze the text based on edges. So as per my understanding,  
> the analysis of a multi word like (Narinder Kaur)term will give  
> N  
> Na  
> Nar  
> Nari  
> Narin  
> Narind  
> Narinde  
> Narinder  
> K  
> Ka  
> Kau  
> Kaur
> 
> So now if search for narinder or kaur by using the following query:
> 
> {  
> "query": {  
> "constant\_score": {  
> "query": {  
> "match\_phrase\_prefix": {  
> "primary\_search\_new": {  
> "query": "narinder",  
> "analyzer": "ys\_search\_analyzer\_long"  
> }  
> }  
> }  
> }  
> }  
> }
> 
> OR
> 
> {  
> "query": {  
> "constant\_score": {  
> "query": {  
> "match\_phrase\_prefix": {  
> "primary\_search\_new": {  
> "query": "kaur",  
> "analyzer": "ys\_search\_analyzer\_long"  
> }  
> }  
> }  
> }  
> }  
> }
> 
> both should have searched for the documents containing "Narinder Kaur".  
> But currently I can not search for kaur. Its working only for first term  
> match. The analyzer's used are as followed:
> 
> analysis: {  
> analyzer: {  
> ys\_search\_analyzer: {  
> type: custom  
> filter: [  
> ys\_word\_delimiter  
> trim  
> lowercase  
> ]  
> tokenizer: ys\_edge\_ngram\_tokenizer  
> }  
> ys\_search\_analyzer\_long: {  
> type: custom  
> filter: [  
> ys\_word\_delimiter  
> trim  
> lowercase  
> ]  
> tokenizer: ys\_edge\_ngram\_tokenizer\_long  
> }  
> }  
> filter: {  
> ys\_word\_delimiter: {  
> type: word\_delimiter  
> stem\_english\_possessive: False  
> }  
> }  
> tokenizer: {  
> ys\_edge\_ngram\_tokenizer\_long: {  
> type: edgeNGram  
> min\_gram: 1  
> max\_gram: 60  
> }  
> ys\_edge\_ngram\_tokenizer: {  
> min\_gram: 1  
> type: edgeNGram  
> max\_gram: 20  
> }  
> }  
> }
> 
> Please elaborate how its not working as expected? and what should I do to  
> make my requirement work without re-indexing the data.
> 
> ## All help is appreciated. thanks
> 
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/76753cd7-7a47-4ca3-ba7b-90be025386b4%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/76753cd7-7a47-4ca3-ba7b-90be025386b4%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/76753cd7-7a47-4ca3-ba7b-90be025386b4%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/76753cd7-7a47-4ca3-ba7b-90be025386b4%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/3ceb9a8f-9846-4779-81ed-8cfc4bb07847%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/3ceb9a8f-9846-4779-81ed-8cfc4bb07847%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:23am UTC](https://discuss.elastic.co/t/n-edge-gram-analyzers-behave-not-as-expected/22884/5 "2017-07-06T00:23:32Z")

</div>


