# Multi match query with non fully analyzed fields

**URL:** <https://discuss.elastic.co/t/multi-match-query-with-non-fully-analyzed-fields/57181>\
**Category:** Elasticsearch\
**Created:** [August 4, 2016, 7:11am UTC](https://discuss.elastic.co/t/multi-match-query-with-non-fully-analyzed-fields/57181 "2016-08-04T07:11:32Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![ugolas](https://avatars.discourse-cdn.com/v4/letter/u/7993a0/32.png) [@ugolas](https://discuss.elastic.co/u/ugolas)\
**Post date:** [August 4, 2016, 7:11am UTC](https://discuss.elastic.co/t/multi-match-query-with-non-fully-analyzed-fields/57181/1 "2016-08-04T07:11:32Z")

</div>

Hi,

We need to implement a free text search, so that a user can search a string and we need to return the docs which have this string in one of multiple fields..

So I've written a multi match, cross fields query on the required fields

Lets say the search is for "Alex"

and the query is:

```
        "multi_match": {
          "query": "Alex",
          "type": "cross_fields",
          "fields": [
            "name",
            "company",
            "email"
          ],

```

Now, name field is standard fully analyzed, company is not\_analyzed (only exact matches should return) and email field is analyzed with tokenizer keyword and filter lowercase (I need exact non case sensitive match of email).

The weird behaviour is when I search for multiple terms which do not exist:

If I search for "Alex Facebook" - I expect to get all docs that the 3 fields above contain either "Alex" or "Facebook", and it there are no matches for "Facebook", I still expect to get the docs which match to "Alex".

But, if I search for a value which matches one of the not fully analyzed fields with another value which does not exists - I get no results.

Example:

Query: "Amazon James"  
There is a doc which company = "Amazon", but there isn't any match for "James" - no result returns.

Can someone explain this behavior and how I can overcome it?

---

<div class="post-metadata">

**Author:** ![davidbkemp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidbkemp/32/594_2.png) [@davidbkemp](https://discuss.elastic.co/u/davidbkemp)\
**Post date:** [August 8, 2016, 9:39am UTC](https://discuss.elastic.co/t/multi-match-query-with-non-fully-analyzed-fields/57181/2 "2016-08-08T09:39:08Z")

</div>

cross\_fields don't work all that well when the fields have different analysers.

"If you include fields with a different analysis chain, they will be added to the query in the same way as for best\_fields"  
[https://www.elastic.co/guide/en/elasticsearch/guide/2.x/\_cross\_fields\_queries.html](https://www.elastic.co/guide/en/elasticsearch/guide/2.x/_cross_fields_queries.html)

Given you have specified "not\_analaysed" for "company", it will be using the full untokenized query and only match companies actually named "Amazon James".

---

<div class="post-metadata">

**Author:** ![ugolas](https://avatars.discourse-cdn.com/v4/letter/u/7993a0/32.png) [@ugolas](https://discuss.elastic.co/u/ugolas)\
**Post date:** [August 8, 2016, 10:25am UTC](https://discuss.elastic.co/t/multi-match-query-with-non-fully-analyzed-fields/57181/3 "2016-08-08T10:25:29Z")

</div>

Hi,

Thanks for the reply. Is there a nice way to achieve what I aiming for in one query?

I've thought about splitting the 'multi\_match' queries between the different analysed fields and assembling them together under 'should'.

---

<div class="post-metadata">

**Author:** ![davidbkemp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidbkemp/32/594_2.png) [@davidbkemp](https://discuss.elastic.co/u/davidbkemp)\
**Post date:** [August 8, 2016, 10:54am UTC](https://discuss.elastic.co/t/multi-match-query-with-non-fully-analyzed-fields/57181/4 "2016-08-08T10:54:29Z")

</div>

Yes, in a similar situation, I've seen a bool/should query work OK. But a problem you need to address first is that you will have trouble getting a string like "Amazon James" to match "Amazon" on a "not\_analysed" field. You might get away with specifying a query time analyser consisting of a standard tokenizer and no token filters, but then no queries would match multi-word company names like "Mercedes Benze". There might be some clever things you could do using shingle token filters in the query analyser to work around this, but it depends on what your requirements really are.

---

<div class="post-metadata">

**Author:** ![ugolas](https://avatars.discourse-cdn.com/v4/letter/u/7993a0/32.png) [@ugolas](https://discuss.elastic.co/u/ugolas)\
**Post date:** [August 9, 2016, 10:33am UTC](https://discuss.elastic.co/t/multi-match-query-with-non-fully-analyzed-fields/57181/5 "2016-08-09T10:33:14Z")

</div>

One last thing, I've did some testing and with seems like using query\_string instead of multi\_match works on the search ''Amazon James" on a non-analyzed field where only Amazon matches..

Why is the difference in behavior between query\_string and multi\_match?

---

<div class="post-metadata">

**Author:** ![davidbkemp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidbkemp/32/594_2.png) [@davidbkemp](https://discuss.elastic.co/u/davidbkemp)\
**Post date:** [August 9, 2016, 12:11pm UTC](https://discuss.elastic.co/t/multi-match-query-with-non-fully-analyzed-fields/57181/6 "2016-08-09T12:11:41Z")

</div>

I haven't used query\_string much, but I'd be interested in seeing some examples that you have managed to get working.

---

<div class="post-metadata">

**Author:** ![ugolas](https://avatars.discourse-cdn.com/v4/letter/u/7993a0/32.png) [@ugolas](https://discuss.elastic.co/u/ugolas)\
**Post date:** [August 9, 2016, 3:52pm UTC](https://discuss.elastic.co/t/multi-match-query-with-non-fully-analyzed-fields/57181/7 "2016-08-09T15:52:25Z")

</div>

From what I understood, query\_string by default splits the entire query to multiple terms by spaces and applies OR operator between them. So basically "Amazon James" is analyzed separately as "Amazon" and "James", for companies as Mercedes Benz they should be passed to the query as "Mercedes Benz" and then this will be the whole term.

---

<div class="post-metadata">

**Author:** ![davidbkemp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidbkemp/32/594_2.png) [@davidbkemp](https://discuss.elastic.co/u/davidbkemp)\
**Post date:** [August 9, 2016, 10:03pm UTC](https://discuss.elastic.co/t/multi-match-query-with-non-fully-analyzed-fields/57181/8 "2016-08-09T22:03:29Z")

</div>

Are you specifying specific fields in the query, or are you leaving it as the default "\_all" field? I ask this because I suspect that everything in the "\_all" field is analysed using the "standard" analyser, and so "Benze Mercedes" would also match "Mercedes Benze".

---

<div class="post-metadata">

**Author:** ![ugolas](https://avatars.discourse-cdn.com/v4/letter/u/7993a0/32.png) [@ugolas](https://discuss.elastic.co/u/ugolas)\
**Post date:** [August 10, 2016, 5:27am UTC](https://discuss.elastic.co/t/multi-match-query-with-non-fully-analyzed-fields/57181/9 "2016-08-10T05:27:19Z")

</div>

Specifying a closed list of fields.. In this example it would be ["name", "company", "email"]

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:28pm UTC](https://discuss.elastic.co/t/multi-match-query-with-non-fully-analyzed-fields/57181/10 "2017-07-05T22:28:54Z")

</div>


