# Phrase matching using query\_string on nGram analyzed data

**URL:** <https://discuss.elastic.co/t/phrase-matching-using-query-string-on-ngram-analyzed-data/9029>\
**Category:** Elasticsearch\
**Created:** [September 14, 2012, 10:11pm UTC](https://discuss.elastic.co/t/phrase-matching-using-query-string-on-ngram-analyzed-data/9029 "2012-09-14T22:11:50Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mike](https://avatars.discourse-cdn.com/v4/letter/m/58f4c7/32.png) [@Mike](https://discuss.elastic.co/u/Mike)\
**Post date:** [September 14, 2012, 10:11pm UTC](https://discuss.elastic.co/t/phrase-matching-using-query-string-on-ngram-analyzed-data/9029/1 "2012-09-14T22:11:50Z")

</div>

I have my string field index\_analyzed with nGrams, and I can't seem to get  
phrase matching using " " in my search text to work. Other things like  
fuzzy matching with ~, combining words with && and ||, boosting with ^ work  
fine though. Am I doing something wrong, or does phrase matching not work  
with ngrams?

My mapping:  
"properties" : {

```
                "myquery" : {                                     
                     
                    "type" : "multi_field",                             
                  
                    "fields" : {                                       
                   
                        "myquery" : { "type" : "string", 

```

"index\_analyzer" : "myAnalyzer", "search\_analyzer" : "myAnalyzer2" },  
"myqueryUntouched" : { "type" : "string",  
"index" : "not\_analyzed" }  
}

```
                },
                ...                

```

My settings:  
"analysis" : {

```
            "analyzer" : {                                             
                   
                "myAnalyzer" : {                                       
                   
                    "tokenizer" : "standard",                           
                  
                    "filter" : ["standard", "lowercase", "stop", 

```

"myNGram"]  
},

```
                "myAnalyzer2" : {                                       
                  
                    "tokenizer" : "standard",                           
                  
                    "filter" : ["standard", "lowercase", "stop"]       
                   
                }                                                       
                  
            },                                                         
                   
            "filter" : {                                               
                   
                "myNGram" : {                                           
                  
                    "type" : "nGram",                                   
                  
                    "min_gram" : 1,                                     
                  
                    "max_gram" : 8                                     
                   
                }                                                      
                   
            }                                                           

```

My query:  
"query":{  
"query\_string":{  
"default\_field":"myquery",  
"default\_operator":"AND",  
"query":""ibm eps""  
}  
}

If I remove the escaped " ", I get many results as I expect, like:  
ibm eps  
ibm q2 eps  
ibm 2001 eps

If someone adds " " though I want only the ibm eps results.

--

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [September 17, 2012, 9:15am UTC](https://discuss.elastic.co/t/phrase-matching-using-query-string-on-ngram-analyzed-data/9029/2 "2012-09-17T09:15:12Z")

</div>

Hi Mike

On Fri, 2012-09-14 at 15:11 -0700, Mike wrote:

> I have my string field index\_analyzed with nGrams, and I can't seem to  
> get phrase matching using " " in my search text to work. Other things  
> like fuzzy matching with ~, combining words with && and ||, boosting  
> with ^ work fine though. Am I doing something wrong, or does phrase  
> matching not work with ngrams?

Phrase matching does work with ngrams, but: there is a long-standing bug  
in the edge-ngram analyzer in lucene which outputs different token  
positions to the standard tokenizer.

So if you analyze the field with edge-ngrams and you do a phrase-search  
on the field using the SAME analyzer, then it will work. But you are  
using the standard tokenizer at search time, not the edge-ngram  
tokenizer.

clint

> My mapping:  
> "properties" : {
> 
> ```
> "myquery" : {
>                            
> "type" : "multi_field",
>                         
> "fields" : {
>                          
> "myquery" : { "type" : "string",
> 
> ```
> 
> "index\_analyzer" : "myAnalyzer", "search\_analyzer" : "myAnalyzer2" },
> 
> ```
> "myqueryUntouched" : { "type" : "string",
> 
> ```
> 
> "index" : "not\_analyzed" }  
> }
> 
> ```
> },
> ...                
> 
> ```
> 
> My settings:  
> "analysis" : {
> 
> ```
> "analyzer" : {
>                          
> "myAnalyzer" : {
>                          
> "tokenizer" : "standard",
>                         
> "filter" : ["standard", "lowercase", "stop",
> 
> ```
> 
> "myNGram"]  
> },
> 
> ```
> "myAnalyzer2" : {
>                         
> "tokenizer" : "standard",
>                         
> "filter" : ["standard", "lowercase", "stop"]
>                          
> }
>                         
> },
>                          
> "filter" : {
>                          
> "myNGram" : {
>                         
> "type" : "nGram",
>                         
> "min_gram" : 1,
>                         
> "max_gram" : 8
>                          
> }
>                          
> }
> 
> ```
> 
> My query:  
> "query":{  
> "query\_string":{  
> "default\_field":"myquery",  
> "default\_operator":"AND",  
> "query":""ibm eps""  
> }  
> }
> 
> If I remove the escaped " ", I get many results as I expect, like:  
> ibm eps  
> ibm q2 eps  
> ibm 2001 eps
> 
> If someone adds " " though I want only the ibm eps results.
> 
> --

--

---

<div class="post-metadata">

**Author:** ![Mike](https://avatars.discourse-cdn.com/v4/letter/m/58f4c7/32.png) [@Mike](https://discuss.elastic.co/u/Mike)\
**Post date:** [September 17, 2012, 2:40pm UTC](https://discuss.elastic.co/t/phrase-matching-using-query-string-on-ngram-analyzed-data/9029/3 "2012-09-17T14:40:48Z")

</div>

> Thanks for the response Clint! I assume what you said applies to both the  
> edge-nGram and regular nGram filters, since I am only using the regular  
> nGrams filter in my index analyzer.

> You mentioned that I should use the ngram tokenizer not the standard  
> tokenizer, does this mean that I should not use the ngram filter? I was  
> hoping to get partial search matches, which is why I used the ngram filter  
> only during index time and not during query time as well (national should  
> find a match with international).

--

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [September 17, 2012, 5:23pm UTC](https://discuss.elastic.co/t/phrase-matching-using-query-string-on-ngram-analyzed-data/9029/4 "2012-09-17T17:23:57Z")

</div>

On Mon, 2012-09-17 at 07:40 -0700, Mike wrote:

> ```
> Thanks for the response Clint! I assume what you said applies
> to both the edge-nGram and regular nGram filters, since I am
> only using the regular nGrams filter in my index analyzer.  
> 
> ```

Yes, it affects the ngrams as well:

[https://issues.apache.org/jira/browse/LUCENE-1224](https://issues.apache.org/jira/browse/LUCENE-1224)

> ```
> You mentioned that I should use the ngram tokenizer not the
> standard tokenizer, does this mean that I should not use the
> ngram filter? I was hoping to get partial search matches,
> which is why I used the ngram filter only during index time
> and not during query time as well (national should find a
> match with international).
> 
> ```

No, you can use the ngram tokenizer or token filter. The important  
thing is to use the same analyzer at index and search time. This is  
almost a golden rule, unless you really understand what you're doing.

clint

> --

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:12am UTC](https://discuss.elastic.co/t/phrase-matching-using-query-string-on-ngram-analyzed-data/9029/5 "2017-07-06T03:12:35Z")

</div>


