# Multimatch with CROSS\_FIELD query and decompounder

**URL:** <https://discuss.elastic.co/t/multimatch-with-cross-field-query-and-decompounder/296852>\
**Category:** Elasticsearch\
**Created:** [February 10, 2022, 2:11pm UTC](https://discuss.elastic.co/t/multimatch-with-cross-field-query-and-decompounder/296852 "2022-02-10T14:11:35Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![thaarbach](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thaarbach/32/87115_2.png) [@thaarbach](https://discuss.elastic.co/u/thaarbach)\
**Post date:** [February 10, 2022, 2:11pm UTC](https://discuss.elastic.co/t/multimatch-with-cross-field-query-and-decompounder/296852/1 "2022-02-10T14:11:35Z")

</div>

The problem was also discussed by @singer and @hbruch

> [@Why does hyphenation\_decompounder require word\_list?](https://discuss.elastic.co/t/why-does-hyphenation-decompounder-require-word-list/114567):
>
> HyphenationCompoundWordTokenFilterFactory inherits from AbstractCompoundWordTokenFilterFactory , which performs a mandatory check for a supplied word\_list. As the underlying lucene HyphenationCompoundWordTokenFilter does not require a word\_list, is there a specific requirement, why it must be supplied for elasticsearch? In my use case, I'd like to avoid specifying in advance all possible matching subwords. Regards, Holger

> [@Decompounder in query\_string analyzer](https://discuss.elastic.co/t/decompounder-in-query-string-analyzer/11749):
>
> Hi everyone, I'm building a search engine for a German website and therefore have to deal with compound word filters... The main problem currently are compound nouns that are sometimes written as one word and sometimes divided by dashes, e.g. "Schlossbergtunnel" and "Schlossberg-Tunnel". A query for either "Schlossbergtunnel" or "Schlossberg-Tunnel" (or "Schlossberg Tunnel") should match both "Schlossbergtunnel" and "Schlossberg-Tunnel". My current approach is to use dictionary\_decompoun…

and

> [@German compound words in an e-commerce search](https://discuss.elastic.co/t/german-compound-words-in-an-e-commerce-search/270280):
>
> Hi elastic! We are developing an e-commerce application and have been using elasticsearch 2.0 since 2014. We are prevented from upgrading by a problem that we simply cannot solve with the current ES version. The german language has the concept of compound words. To fit our needs in ES 2.0 we use the plugin elasticsearch-analysis-decompound which works fine. Typical words in our domain are: [Kabelkanal, Antennenkabel, Antennenhalterung, Tablethalterung, Tablethalter] =\> [cable duct, aerial w…

Here a example, that the search works not as expected:

Create a index

```auto
PUT example/
{
  "settings": {
    "number_of_shards": 1,
    "number_of_replicas": 0,
    "analysis": {
      "filter": {
        "decomp_de": {
          "type": "hyphenation_decompounder",
          "word_list": ["kaffee", "tasse", "tüte"],
          "hyphenation_patterns_path": "hyph/de_DR.xml",
          "min_subword_size": 3,
          "only_longest_match": true
        }
      },
      "analyzer": {
        "german_analyzer": {
          "filter": [
            "lowercase",
            "decomp_de",
            "unique"
          ],
          "type": "custom",
          "tokenizer": "standard"
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "text": {
        "type": "text",
        "analyzer": "german_analyzer",
        "norms": false
      },
      "type": {
        "type": "text",
        "analyzer": "german_analyzer",
        "norms": false
      }
    }
  }
}

```

Index documents

```POST
{"index":{"_id":1}}
{"text": "Kaffeetasse", "type":"Tasse"}
{"index":{"_id":2}}
{"text": "Kaffeetüte", "type": "Tüte"}

```

Serach for `Kaffeetasse`

```GET
{
  "query": {
    "multi_match": {
      "query": "Kaffeetasse",
      "fields": [
        "text",
        "type"
      ],
      "type": "cross_fields",
      "operator": "and",
      "slop": 1,
      "prefix_length": 0,
      "max_expansions": 50,
      "zero_terms_query": "none",
      "auto_generate_synonyms_phrase_query": "true",
      "fuzzy_transpositions": false,
      "boost": 1
    }
  }
}

```

The unexpected result and it doesn't matter which type, operator or minmum should match is given is

```auto
  "took" : 1,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 2,
      "relation" : "eq"
    },
    "max_score" : 0.6931471,
    "hits" : [
      {
        "_index" : "example",
        "_type" : "_doc",
        "_id" : "1",
        "_score" : 0.6931471,
        "_source" : {
          "text" : "Kaffeetasse",
          "type" : "Tasse"
        }
      },
      {
        "_index" : "example",
        "_type" : "_doc",
        "_id" : "2",
        "_score" : 0.25069216,
        "_source" : {
          "text" : "Kaffeetüte",
          "type" : "Tüte"
        }
      }
    ]
  }
}

```

The expected result is only the document containing `Kaffeetasse`.

Analyzing:

```auto
GET example/_analyze
{
  "analyzer": "german_analyzer"
  , "text": "Kaffeetasse"
}

```

produces

```auto
{
  "tokens" : [
    {
      "token" : "kaffeetasse",
      "start_offset" : 0,
      "end_offset" : 11,
      "type" : "<ALPHANUM>",
      "position" : 0
    },
    {
      "token" : "kaffee",
      "start_offset" : 0,
      "end_offset" : 11,
      "type" : "<ALPHANUM>",
      "position" : 0
    },
    {
      "token" : "tasse",
      "start_offset" : 0,
      "end_offset" : 11,
      "type" : "<ALPHANUM>",
      "position" : 0
    }
  ]
}

```

And IMHO the query should be rewritten to `(text:kaffeetasse OR (text:kaffee AND text: tasse)) OR (type:kaffeetasse OR (type:kaffee AND type: tasse)) ` and not to `(text:kaffeetasse OR text:kaffee OR text: tasse OR type:kaffeetasse OR type:kaffee OR type: tasse)`.

With the `_validate/query`

```auto
GET example/_validate/query?explain=true&rewrite=true
{
  "query": {
    "multi_match": {
      "query": "Kaffeetasse",
      "fields": [
        "text",
        "type"
      ],
      "type": "cross_fields",
      "operator": "and",
      "slop": 0,
      "prefix_length": 0,
      "max_expansions": 50,
      "minimum_should_match": "-45%",
      "zero_terms_query": "NONE",
      "auto_generate_synonyms_phrase_query": "false",
      "fuzzy_transpositions": false,
      "boost": 1
    }
  }
}

```

i got

```auto
{
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "failed" : 0
  },
  "valid" : true,
  "explanations" : [
    {
      "index" : "example",
      "valid" : true,
      "explanation" : "(text:kaffeetasse | text:kaffee | text:tasse | type:kaffeetasse | type:kaffee | type:tasse)"
    }
  ]
}

```

or without parameter rewrite:

```auto
{
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "failed" : 0
  },
  "valid" : true,
  "explanations" : [
    {
      "index" : "example",
      "valid" : true,
      "explanation" : "blended(terms:[text:kaffeetasse, text:kaffee, text:tasse, type:kaffeetasse, type:kaffee, type:tasse])"
    }
  ]
}

```

Searching for `Kaffee` should return both documents, searching for `Kaffeetasse` or `tasse` the document id `1` and searching for `Kaffeetüte` or `tüte` document with the id `2` or what am I misunderstanding?

---

<div class="post-metadata">

**Author:** ![thaarbach](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thaarbach/32/87115_2.png) [@thaarbach](https://discuss.elastic.co/u/thaarbach)\
**Post date:** [February 14, 2022, 10:33am UTC](https://discuss.elastic.co/t/multimatch-with-cross-field-query-and-decompounder/296852/2 "2022-02-14T10:33:40Z")

</div>

Based on the example shown above, searching for `kaffee tasse` returns the expected result:

```auto
GET example/_search
{
  "query": {
    "multi_match": {
      "query": "Kaffee tasse",
      "fields": [
      "text", 
        "type",
        "text.raw"
      ],
      "type": "cross_fields",
      "operator": "and",
      "slop": 1,
      "prefix_length": 0,
      "max_expansions": 50,
      "zero_terms_query": "none",
      "auto_generate_synonyms_phrase_query": "false",
      "minimum_should_match": "-66%", 
      "fuzzy_transpositions": false,
      "boost": 1
    }
  }
}

```

```auto
{
  "took" : 6,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 1,
      "relation" : "eq"
    },
    "max_score" : 1.2037694,
    "hits" : [
      {
        "_index" : "example",
        "_type" : "_doc",
        "_id" : "1",
        "_score" : 1.2037694,
        "_source" : {
          "text" : "Kaffeetasse",
          "type" : "Tasse"
        }
      }
    ]
  }
}

```

If I understand it correctly, decompounding is intended to ensure that searching for `kaffeetasse` or `kaffee tasse` returns the same results or am I wrong?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 14, 2022, 10:34am UTC](https://discuss.elastic.co/t/multimatch-with-cross-field-query-and-decompounder/296852/3 "2022-03-14T10:34:06Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
