# Analyzer issues using query\_string in ES 1.5.2?

**URL:** <https://discuss.elastic.co/t/analyzer-issues-using-query-string-in-es-1-5-2/28944>\
**Category:** Elasticsearch\
**Created:** [September 9, 2015, 3:39pm UTC](https://discuss.elastic.co/t/analyzer-issues-using-query-string-in-es-1-5-2/28944 "2015-09-09T15:39:31Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![dunnlow](https://avatars.discourse-cdn.com/v4/letter/d/8797f3/32.png) [@dunnlow](https://discuss.elastic.co/u/dunnlow)\
**Post date:** [September 9, 2015, 3:39pm UTC](https://discuss.elastic.co/t/analyzer-issues-using-query-string-in-es-1-5-2/28944/1 "2015-09-09T15:39:31Z")

</div>

I'm a noob using ES 1.5.2 I want to ngram analyze a field on index, but do no analysis on search. Why? I want a user to be able to search for "group" and match the field "aged grouper" (no wildcards required - but still supported). However, if the user enters "aged grouper" I only want to match documents where my search field contains (at least) that entire phase.

I created an ngram analyzer that I map to the field for index, and a "dummy analyzer" (to keep the whole phrase together) that I map to the field for search. I can test both analyzers using the analyze api, and see that they are getting tokenized correctly.

Everything seems correct. However, when I do my query\_string search, the search text still gets tokenized into words. So, searching for "group" DOES find "aged grouper" but searching for "the group" finds all documents that have EITHER "the" OR "group" in them. I want the whole phrase to be used in the search.

I'm confused that when I use the analyze api and the validate api I seem to get two different answers (I think):

If I use the analyze api: \_analyze/analyzer=dummy\_analyzer&text=Hello there  
..  
`<token>Hello there</token>` \<== looks correct  
..

However, If I use the validate api:  
`_validate/query?pretty=true&explain=true&analyzer=dummy_analyzer`

```
{ "query" : {
     "query_string" : {
         "query" : "Hello there",
         "default_field" : "tfield",
         "analyzer" : "dummy_analyzer"
      }
   }
}

```

results in:  
`<explanation>props.tfield:Hello props.tfield:there</explanation>` \<== looks INCORRECT (breaking phrase apart)

My config is below. My questions:

- Can someone explain the differences between the api results?
- Why isn't the search using the dummy\_analyzer (would you expect this approach to work)?
- Is there a better way to have a field not analyzed on search only rather than using my kludged dummy\_analyzer)

Thanks very much for any insight! -J

```
"analysis":{
  "analyzer":{
      "ngram_analyzer":{
          "type":"custom",
          "tokenizer":"ngram_tokenizer"
      },
      "dummy_analyzer":{
        "type":"pattern",
        "pattern":"00xyzzy00" <-- a dummy string trying to never separate words
      }
   },
   "tokenizer":{
       "ngram_tokenizer": {
           "type":"nGram",
           "min_gram":"4",
           "max_gram":"500"
        }
    }
}

"mapping":{
 ....
   "tfield":{
       "index_analyzer":"ngram_analyzer",
       "search_analyzer":"dummy:analyzer",
       "type": "string",
       "index","analyzed"
    }
....
```

---

<div class="post-metadata">

**Author:** ![javanna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javanna/32/4698_2.png) [@javanna](https://discuss.elastic.co/u/javanna)\
**Post date:** [September 10, 2015, 4:29pm UTC](https://discuss.elastic.co/t/analyzer-issues-using-query-string-in-es-1-5-2/28944/2 "2015-09-10T16:29:18Z")

</div>

I think these differences may just have to do with using the query\_string. May I ask if you tried the match query instead? Or are there features that you need out of the query\_string query?

---

<div class="post-metadata">

**Author:** ![dunnlow](https://avatars.discourse-cdn.com/v4/letter/d/8797f3/32.png) [@dunnlow](https://discuss.elastic.co/u/dunnlow)\
**Post date:** [September 10, 2015, 6:39pm UTC](https://discuss.elastic.co/t/analyzer-issues-using-query-string-in-es-1-5-2/28944/3 "2015-09-10T18:39:34Z")

</div>

The syntax of the query\_string is ideal for my users. I suppose I could give up on providing wild card. It was suggested that I abandon query\_string and use match (working on that now)

My goal:

- query `"fort"` should match "unfortunately" (as if `"*fort*"` was entered)
- query "is unfortunate" should only match fields with (at least) that whole phrase
- query "my fort??e" should match "my fortune" (in the best of all possible worlds)
- queries should allow for simple AND, OR, NOT, and () grouping logic
- there can be no fuzziness (only allow exact phase matches with constraints above)

I get the leading/trailing wildcard simulation by indexing with an ngram analyzer (I can test with the analyze api and verify it is working correctly).

Does this seem feasible with a match (minus the mid-term wildcard support)?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:51pm UTC](https://discuss.elastic.co/t/analyzer-issues-using-query-string-in-es-1-5-2/28944/4 "2017-07-05T23:51:05Z")

</div>


