# Avoiding analysis of query strings

**URL:** <https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255>\
**Category:** Elasticsearch\
**Created:** [May 25, 2015, 1:02pm UTC](https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255 "2015-05-25T13:02:07Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [May 25, 2015, 1:02pm UTC](https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255/1 "2015-05-25T13:02:07Z")

</div>

I’m using a path hierarchy analyzer to analyze fields containing Java class names:

```
"analysis": {
  "analyzer": {
    "java_classname_analyzer": {
      "tokenizer": "java_classname_tokenizer",
      "type": "custom"
    }
  },
  "tokenizer": {
    "java_classname_tokenizer": {
      "type": "PathHierarchy",
      "delimiter": ".",
      "reverse": false
    }
  }
}

```

This results in a correct tokenization of input strings, i.e. java.io.File gets tokenized into (java, [java.io](http://java.io), java.io.File) and I expected to be able to search for [java.io](http://java.io) and get back e.g. java.io.File, java.io.Reader, and java.io.Writer. However, I’ve realized that when including a java\_classname\_analyzer field in a query string, e.g. using the query

```
class:java.io

```

in Kibana I get many more hits than I asked for since the search term itself is tokenized into (java, [java.io](http://java.io)) and I’m actually getting hits for everything matching java.\*.

Is there a way to I avoid this? With the query DSL I guess I can use a term query rather than e.g. a match query but since the use case is Kibana it needs to be a query string.

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [May 25, 2015, 2:10pm UTC](https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255/2 "2015-05-25T14:10:11Z")

</div>

You could configure your string field to have a `keyword` analyzer as a `search_analyzer`.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [May 27, 2015, 8:56am UTC](https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255/3 "2015-05-27T08:56:57Z")

</div>

Thanks @jpountz, that looks exactly like what I'm looking for (could've sworn I looked at that documentation page the other day). However, I don't get the result I'm looking for. I updated my index template so that the logstash-2015.05.27 index uses the keyword analyzer for a number of fields and verified that the actual mapping of the index looks okay:

```
$ curl --silent hostname:9200/logstash-2015.05.27/_mapping/dotnet | \
    jq '."logstash-2015.05.27".mappings.dotnet.properties.class'
{
  "type": "string",
  "analyzer": "java_classname_analyzer",
  "fields": {
    "raw": {
      "type": "string",
      "index": "not_analyzed",
      "ignore_above": 256
    }
  },
  "search_quote_analyzer": "keyword"
}

```

Well, I expected "search\_analyzer" rather than "search\_quote\_analyzer" but I suppose that's okay.

However, a Kibana query for `type:dotnet AND class:TestApp.MainForm.Foo` still returns the following:

```
{
  "_index": "logstash-2015.05.27",
  "_type": "dotnet",
  "_id": "AU2Uf1duWK9xLBfC4G3m",
  "_score": null,
  "_source": {
    "@timestamp": "2015-05-27T10:31:22.126+02:00",
    "message": "another test message",
    "type": "dotnet",
    "class": "TestApp.MainForm",
    "@version": "1",
  },
  "sort": [
    1432715482126,
    1432715482126
  ]
}

```

Is this by any chance because ES doesn't analyze the query for each index being searched but in this case uses the analyzer specified in the mappings of the dozens of other indexes that don't use the keyword analyzer for that field?

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [May 27, 2015, 11:41am UTC](https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255/4 "2015-05-27T11:41:12Z")

</div>

> [@magnusbaeck](#):
>
> Well, I expected "search\_analyzer" rather than "search\_quote\_analyzer" but I suppose that's okay.

Hmm, this looks like a bug! How does your mapping template look like, did you actually modify the `search_analyzer` and not the `search_quote_analyzer`?

Regarding your other question, Elasticsearch actually analyzes the query string per shard, so your change to new indices should work on these new indices.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [May 27, 2015, 2:53pm UTC](https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255/5 "2015-05-27T14:53:42Z")

</div>

> Hmm, this looks like a bug! How does your mapping template look like, did you actually modify the search\_analyzer and not the search\_quote\_analyzer?

Yes, this is what I ended up with:

```
    "class": {
      "type": "string",
      "analyzer": "java_classname_analyzer",
      "search_analyzer": "keyword",
      "fields": {
        "raw": {
          "type": "string",
          "index": "not_analyzed",
          "ignore_above": 256
        }
      }
    },

```

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [May 27, 2015, 3:09pm UTC](https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255/6 "2015-05-27T15:09:27Z")

</div>

What version of Elasticsearch are you running?

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [May 27, 2015, 7:58pm UTC](https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255/7 "2015-05-27T19:58:30Z")

</div>

We're running ES 1.5.2.

---

<div class="post-metadata">

**Author:** ![plebedev](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/plebedev/32/6278_2.png) [@plebedev](https://discuss.elastic.co/u/plebedev)\
**Post date:** [November 30, 2015, 10:38pm UTC](https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255/8 "2015-11-30T22:38:33Z")

</div>

We have the same problem with ES 1.7.0. We specify search\_analyzer in the mapping template but the actual mapping has search\_quote\_analyzer and it does not work as expected.

---

<div class="post-metadata">

**Author:** ![adrianocrestani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/adrianocrestani/32/4416_2.png) [@adrianocrestani](https://discuss.elastic.co/u/adrianocrestani)\
**Post date:** [December 16, 2015, 4:46pm UTC](https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255/9 "2015-12-16T16:46:37Z")

</div>

I just saw the same problem in 2.1. In the processing of migrating from 1.5.2 to 2.1, I used the same mapping and in 2.1 my search\_analyzer gets applied as search\_quote\_analyzer.

---

<div class="post-metadata">

**Author:** ![adrianocrestani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/adrianocrestani/32/4416_2.png) [@adrianocrestani](https://discuss.elastic.co/u/adrianocrestani)\
**Post date:** [December 16, 2015, 8:14pm UTC](https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255/10 "2015-12-16T20:14:30Z")

</div>

Actually, I think it is a 1.5.2 bug, probably fixed in a later version.

So, in the process of I changed all my mappings to no longer use "index\_analyzer" and just use "analyzer". So, all my fields would have "search\_analyzer" and "analyzer". However, when I do that on 1.5.2, my index mapping ends up like this:

"search\_quote\_analyzer": "ngram\_search\_analyzer",  
"analyzer": "ngram\_index\_analyzer",

As you can see, there is not search\_analzyer and when you run a search against that field, it seems to default the "search\_analyzer" to "analyzer".

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:30pm UTC](https://discuss.elastic.co/t/avoiding-analysis-of-query-strings/1255/11 "2017-07-05T23:30:25Z")

</div>


