# Organize search by priority with/without spaces

**URL:** <https://discuss.elastic.co/t/organize-search-by-priority-with-without-spaces/294276>\
**Category:** Elasticsearch\
**Created:** [January 13, 2022, 1:12pm UTC](https://discuss.elastic.co/t/organize-search-by-priority-with-without-spaces/294276 "2022-01-13T13:12:57Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![gagasusrock](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gagasusrock/32/100209_2.png) [@gagasusrock](https://discuss.elastic.co/u/gagasusrock)\
**Post date:** [January 13, 2022, 1:12pm UTC](https://discuss.elastic.co/t/organize-search-by-priority-with-without-spaces/294276/1 "2022-01-13T13:12:57Z")

</div>

Hello. I am using elastic 7.15.0.  
I want to organize search by priority with/without spaces.  
What does it mean?

query - "surf coff"

1. I want to see records contains or startswith "surf coff" - (SURF COFFEE, SURF CAFFETERIA, SURFCOFFEE MAN)
2. When records contains or startswith "surf" - (SURF, SURF LOVE, ENDLESS SURF)
3. When records contains or startswith "coff" - (LOVE COFFEE, COFFEE MAN)

query - "surfcoff"

1. i want to see records which contains or startswith "surfcoff" - (SURF COFFEE, SURF CAFFETERIA, SURFCOFFEE MAN) only.

I created the analyzer with filters:

- lowercase
- word\_delimiter\_graph
- shingle
- edge n gram
- pattern replace for spaces

```auto
{
   "settings":{
       "index": {
            "max_shingle_diff" : 9,
            "max_ngram_diff": 9
       },
      "analysis":{
         "analyzer":{
            "word_join_analyzer":{
               "tokenizer":"standard",
               "filter":[
                  "lowercase",
                  "word_delimiter_graph",
                  "my_shingle",
                  "my_edge_ngram",
                  "my_char_filter"
               ]
            }
         },
         "filter":{
            "my_shingle":{
               "type":"shingle",
               "min_shingle_size": 2,
                "max_shingle_size": 10
            },
            "my_edge_ngram": { 
                "type": "edge_ngram",
                "min_gram": 2,
                "max_gram": 10,
                "token_chars": ["letter", "digit"]
            },
            "my_char_filter": {
                "type": "pattern_replace",
                "pattern": " ",
                "replacement": ""
            }
         }
      }
   }
}

```

So when i analyzed text = "SURF COFFEE", i got this result

```auto
{
    "tokens": [
        {
            "token": "su",
            "start_offset": 0,
            "end_offset": 4,
            "type": "<ALPHANUM>",
            "position": 0
        },
        {
            "token": "sur",
            "start_offset": 0,
            "end_offset": 4,
            "type": "<ALPHANUM>",
            "position": 0
        },
        {
            "token": "surf",
            "start_offset": 0,
            "end_offset": 4,
            "type": "<ALPHANUM>",
            "position": 0
        },
        {
            "token": "su",
            "start_offset": 0,
            "end_offset": 11,
            "type": "shingle",
            "position": 0,
            "positionLength": 2
        },
        {
            "token": "sur",
            "start_offset": 0,
            "end_offset": 11,
            "type": "shingle",
            "position": 0,
            "positionLength": 2
        },
        {
            "token": "surf",
            "start_offset": 0,
            "end_offset": 11,
            "type": "shingle",
            "position": 0,
            "positionLength": 2
        },
        {
            "token": "surf",
            "start_offset": 0,
            "end_offset": 11,
            "type": "shingle",
            "position": 0,
            "positionLength": 2
        },
        {
            "token": "surfc",
            "start_offset": 0,
            "end_offset": 11,
            "type": "shingle",
            "position": 0,
            "positionLength": 2
        },
        {
            "token": "surfco",
            "start_offset": 0,
            "end_offset": 11,
            "type": "shingle",
            "position": 0,
            "positionLength": 2
        },
        {
            "token": "surfcof",
            "start_offset": 0,
            "end_offset": 11,
            "type": "shingle",
            "position": 0,
            "positionLength": 2
        },
        {
            "token": "surfcoff",
            "start_offset": 0,
            "end_offset": 11,
            "type": "shingle",
            "position": 0,
            "positionLength": 2
        },
        {
            "token": "surfcoffe",
            "start_offset": 0,
            "end_offset": 11,
            "type": "shingle",
            "position": 0,
            "positionLength": 2
        },
        {
            "token": "co",
            "start_offset": 5,
            "end_offset": 11,
            "type": "<ALPHANUM>",
            "position": 1
        },
        {
            "token": "cof",
            "start_offset": 5,
            "end_offset": 11,
            "type": "<ALPHANUM>",
            "position": 1
        },
        {
            "token": "coff",
            "start_offset": 5,
            "end_offset": 11,
            "type": "<ALPHANUM>",
            "position": 1
        },
        {
            "token": "coffe",
            "start_offset": 5,
            "end_offset": 11,
            "type": "<ALPHANUM>",
            "position": 1
        },
        {
            "token": "coffee",
            "start_offset": 5,
            "end_offset": 11,
            "type": "<ALPHANUM>",
            "position": 1
        }
    ]
}

```

As you can see, there is token "surfcoff".

How my search should be organized?

I've tried to combine approaches by bool should query with -  
query\_string, match\_phrase\_prefix, match\_prefix and others.

But none of them gave correct results.

Can you please, help me.

How my query should be built?  
Or maybe i should try other analyzer filters.

For example query

```auto
{
  "query": {
    "bool": {
      "should": [
        {
          "query_string": {
                "query": "surf coff",
                "default_field": "text",
                "default_operator": "AND"
            }
        },
        {
          "query_string": {
                "query": "surf",
                "default_field": "text"
            }
        },
        {
          "query_string": {
                "query": "coff",
                "default_field": "text"
            }
        }
      ]
    }
  }
}

```

or this query

```auto
{
  "query": {
    "bool": {
      "should": [
        {
          "query_string": {
                "query": "(surf coff) OR (surf) OR (coff)",
                "default_field": "text"
            }
        }
      ]
    }
  }
}

```

or this query

```auto
{
  "query": {
    "bool": {
      "should": [
        {
          "query_string": {
                "query": "((surf AND coff)^3 OR (surf)^2 OR (coff)^1)",
                "default_field": "text"
            }
        }
      ]
    }
  }
}

```

or

```auto
{
  "query": {
    "match_bool_prefix" : {
      "text" : "surf coff"
    }
  }
}

```

gives

1. SURF COFFEE SURFING NEVER ALONE
2. CONOSUR COLCHAGUA CONO SUR
3. SUNRISE CONCHA TORO SUNRISE 300 DAYS
4. SUN COFFEE
5. SURF COFFEE PROPAGANDA  
....

but its strange for me, i think i misunderstand something.

---

<div class="post-metadata">

**Author:** ![gagasusrock](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gagasusrock/32/100209_2.png) [@gagasusrock](https://discuss.elastic.co/u/gagasusrock)\
**Post date:** [January 14, 2022, 10:16am UTC](https://discuss.elastic.co/t/organize-search-by-priority-with-without-spaces/294276/2 "2022-01-14T10:16:58Z")

</div>

```auto
{
  "query": {
    "bool": {
      "should": [
        {
          "query_string": {
                "query": "(surf* AND coff*)^3 OR (surf*)^2 OR (coff*)^1",
                "default_field": "text"
            }
        }
      ]
    }
  }
}

```

```auto
{
   "settings":{
       "index": {
            "max_shingle_diff" : 9,
            "max_ngram_diff": 9
       },
      "analysis":{
         "analyzer":{
            "word_join_analyzer":{
               "tokenizer":"standard",
               "filter":[
                  "lowercase",
                  "word_delimiter_graph",
                   "my_shingle",
                   "my_char_filter"
               ]
            }
         },
         "filter":{
            "my_shingle":{
               "type":"shingle",
               "min_shingle_size": 2,
                "max_shingle_size": 10
            },
            "my_char_filter": {
                "type": "pattern_replace",
                "pattern": " ",
                "replacement": ""
            }
         }
      }
   }
}

```

removing edge-n-gram and adding wilcard query with priority resolved my question.  
But I still don't understand why edge n gram didn't work.

---

<div class="post-metadata">

**Author:** ![gagasusrock](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gagasusrock/32/100209_2.png) [@gagasusrock](https://discuss.elastic.co/u/gagasusrock)\
**Post date:** [January 14, 2022, 2:05pm UTC](https://discuss.elastic.co/t/organize-search-by-priority-with-without-spaces/294276/3 "2022-01-14T14:05:42Z")

</div>

Finally resolved with

```auto
"filter":[
                  "lowercase",
                  "word_delimiter_graph",
                  "my_shingle",
                   "my_edge_ngram",
                  "my_char_filter"
               ]

```

the problem was with search\_analyzer, because docs said " Sometimes, though, it can make sense to use a different analyzer at search time, such as when using the [`edge_ngram`](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-edgengram-tokenizer.html) tokenizer for autocomplete or when using search-time synonyms."

So i added standard search\_analyzer to my text field:

```auto
"text": { "type": "text", "analyzer": "word_join_analyzer", "search_analyzer": "standard" }

```

Search query:

```auto
{
  "query": {
    "bool": {
      "should": [
        {
          "query_string": {
                "query": "surf coff",
                "default_field": "text"
            }
        }
      ]
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 11, 2022, 2:06pm UTC](https://discuss.elastic.co/t/organize-search-by-priority-with-without-spaces/294276/4 "2022-02-11T14:06:36Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
