# Custom analyzer on match\_phrase

**URL:** <https://discuss.elastic.co/t/custom-analyzer-on-match-phrase/123846>\
**Category:** Elasticsearch\
**Created:** [March 14, 2018, 6:51am UTC](https://discuss.elastic.co/t/custom-analyzer-on-match-phrase/123846 "2018-03-14T06:51:23Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![bunch\_of\_bytes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bunch_of_bytes/32/12771_2.png) [@bunch\_of\_bytes](https://discuss.elastic.co/u/bunch_of_bytes)\
**Post date:** [March 14, 2018, 6:51am UTC](https://discuss.elastic.co/t/custom-analyzer-on-match-phrase/123846/1 "2018-03-14T06:51:23Z")

</div>

Hi,

I am quite puzzled by using analyzer on one of my search fields. Here is the mapping:

```
{"settings": {
    "analysis": {
      "filter": {
        "filter_shingle":{
               "type":"shingle",
               "max_shingle_size":3,
               "min_shingle_size":2,
               "output_unigrams":"true"
        },
        "tf_eng_stop": {
                "type": "stop",
                "stopwords": "_english_"
              },
        "tf_title_stop": {
                    "type": "stop",
                    "stopwords": ["intern", "internship", "senior", "Sr.", "Sr"]
                },      
        "tf_synonym": {
          "type": "synonym",
          "synonyms_path" : "synonyms.txt"
        }
      },
      "analyzer": {
        "tf_synonym_analyzer": {
          "tokenizer": "whitespace",
          "filter": [
              "lowercase",
              "tf_eng_stop",
            "tf_synonym"
             
            
          ]
        },
        "tf_title_analyzer": {
          "tokenizer": "standard",
          "filter": [
              "lowercase",
              "tf_title_stop",
              "standard",
              "filter_shingle"
            
          ]
        },
        "tf_synonym_analyzer_keyword_only":{
            "tokenizer": "whitespace",
          "filter": [
            "lowercase",
            "tf_eng_stop",
            "tf_synonym"
          ]
        }
      }
    }
  },
  
       "mappings":{  
          "job":{  
             "properties":{  
                "name":{  
                   "type":"text"
                },
                "keywords":{  
                    "type":"text",
                    "analyzer":"tf_synonym_analyzer"
                }, 
                "alias":{  
                    "type":"text"
                },
                "color":{  
                    "type":"text"
                },
                "id":{  
                   "type":"long"
                }
             }
          }
       }
   
}

```

Here is the query:

\_search

```
{  
   "query":{  
      "match_phrase":{  
         "alias":{  
            "query":"senior staff engineer/ manager",
            "analyzer":"tf_title_analyzer",
            "boost":1.5
         }
      }
   },
   "_source":{  
      "includes":[  
         "name",
         "color"
      ]
   },
   "highlight":{  
      "fields":{  
         "alias":{  

         }
      }
   }
}

```

I noticed if the query is "senior staff engineer", nothing comes up. If I use "staff engineer", it returns a result. I am not sure why since I specified the query to use a stop word token filter already. Can someone help?

Thanks a lot!

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [March 14, 2018, 10:12am UTC](https://discuss.elastic.co/t/custom-analyzer-on-match-phrase/123846/2 "2018-03-14T10:12:26Z")

</div>

I suspect shingles are confusing the query parser a bit. Can you share the output of the `validate` API on your query with `rewrite` equal to true? [https://www.elastic.co/guide/en/elasticsearch/reference/current/search-validate.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-validate.html)

---

<div class="post-metadata">

**Author:** ![bunch\_of\_bytes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bunch_of_bytes/32/12771_2.png) [@bunch\_of\_bytes](https://discuss.elastic.co/u/bunch_of_bytes)\
**Post date:** [March 14, 2018, 4:38pm UTC](https://discuss.elastic.co/t/custom-analyzer-on-match-phrase/123846/3 "2018-03-14T16:38:21Z")

</div>

Thank you very much for your quick response.

Here is the explain:

```
"explanations": [
      {
        "index": "jobs",
        "valid": true,
        "explanation": "(alias:\"(_ staff _ staff engineer _ staff engineer manager) (staff staff engineer staff engineer manager) (engineer engineer manager) manager\")^1.5"
      }
    ]

```

It looks like the shingle is being funny. Why would it break down words like that?

Senior Staff Engineer Manager should be something like

staff engineer, engineer manager.....

One more thing, what's the relationship among the words inside the bracket (staff staff engineer staff engineer manager). Are they OR or AND? or this is just a long string that is taken as a phrase?

What should I do to make the shingle behave correctly?

Thanks a lot!

UPDATE  
If I use the same set up, but remove "senior" in the query, here is the explaination:

```
"explanations": [
      {
        "index": "jobs",
        "valid": true,
        "explanation": "(alias:\"(staff staff engineer staff engineer manager) (engineer engineer manager) manager\")^1.5"
      }
    ]

```

It does have a hit. But I cannot understand is the difference

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [March 20, 2018, 8:39am UTC](https://discuss.elastic.co/t/custom-analyzer-on-match-phrase/123846/4 "2018-03-20T08:39:15Z")

</div>

Unfortunately I think there are multiple issues here, some of them being hard to fix:

- shingles do not work well with synonyms at the moment [https://issues.apache.org/jira/browse/LUCENE-3475](https://issues.apache.org/jira/browse/LUCENE-3475)
- you should only use one shingle size at search time
- shingles don't make it easy to integrate correctly with match\_phrase.

I'd probably recommend to remove shingles from the analyzers.

---

<div class="post-metadata">

**Author:** ![bunch\_of\_bytes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bunch_of_bytes/32/12771_2.png) [@bunch\_of\_bytes](https://discuss.elastic.co/u/bunch_of_bytes)\
**Post date:** [March 20, 2018, 6:39pm UTC](https://discuss.elastic.co/t/custom-analyzer-on-match-phrase/123846/5 "2018-03-20T18:39:52Z")

</div>

Thank you very much for your response. Another a different note, in this case, the documents consisit mostly of phrases.

for example:

"Senior Software Developer"  
"Data Analysts"

They are not really documents.

I found if the search query is "Data Engineer"

It may include "Data Analyst" as result since they both contain "data".

Is there any way to index these phrases as they are?

I also looked at Span search. I found Span Near maybe the best query. However, the issue is that it has to include all span terms. But in this case, we don't necessarily need ALL span terms. For example:

```
{
    "query": {
        "span_near" : {
            "clauses" : [
                { "span_term" : { "field" : "Senior" } },
                { "span_term" : { "field" : "Engineering" } },
                { "span_term" : { "field" : "Manager" } }
            ],
            "slop" : 12,
            "in_order" : true
        }
    }
}

```

What if the document contains a phrase "Engineering Manager"? This search would not come up since it also looks for "senior". span\_or on the other hand does not support in\_order or slope. Any suggestion?

Thaks!

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [March 26, 2018, 1:13pm UTC](https://discuss.elastic.co/t/custom-analyzer-on-match-phrase/123846/6 "2018-03-26T13:13:00Z")

</div>

> [@bunch\_of\_bytes](#):
>
> It may include "Data Analyst" as result since they both contain "data".

This is true, but at the same time matches that contain both `data` and `analyst` should rank higher than those that only contain `data`.

Query parsers also have a way to make all terms required, have a look at the `operator` or `minimum_should_match` options. [minimum\_should\_match parameter | Elasticsearch Guide [8.11] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-minimum-should-match.html)

> [@bunch\_of\_bytes](#):
>
> I also looked at Span search. I found Span Near maybe the best query. However, the issue is that it has to include all span terms. But in this case, we don't necessarily need ALL span terms. For example:

Right. There is no easy answer to this problem. One workaround would be to put your regular query in a MUST clause and a phrase (or span) query in a should clause. This way the phrase query is not required for matching, but if it matches then it will boost scores.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 23, 2018, 1:13pm UTC](https://discuss.elastic.co/t/custom-analyzer-on-match-phrase/123846/7 "2018-04-23T13:13:03Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
