# Tokenizer: whitespace not working with edge\_ngram

**URL:** https://discuss.elastic.co/t/tokenizer-whitespace-not-working-with-edge-ngram/118397
**Category:** Elasticsearch
**Created:** [February 5, 2018, 7:29am UTC](https://discuss.elastic.co/t/tokenizer-whitespace-not-working-with-edge-ngram/118397 "2018-02-05T07:29:02Z")
**Posts on this page:** 10
**Page:** 1

<div class="post-metadata">

### Author: ![anoopvalluthadam](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anoopvalluthadam/32/30286_2.png) [@anoopvalluthadam](https://discuss.elastic.co/u/anoopvalluthadam)
#### Post date: [February 5, 2018, 7:29am UTC](https://discuss.elastic.co/t/tokenizer-whitespace-not-working-with-edge-ngram/118397/1 "2018-02-05T07:29:02Z")

</div>

Trying to include special characters in ngram tokeniser

```
 DELETE test
    PUT test
    {
      "settings": {
        "analysis": {
          "analyzer": {
            "my_analyzer": {
              "type": "custom",
              "tokenizer": "whitespace",
              "filter": [
                "lowercase", "ngram", "asciifolding", "stop"
              ]
            }
          },
          "filter": {
            "ngram": {
              "type": "edge_ngram",
              "min_gram": 1,
              "max_gram": 20,
              "token_chars": [
                "letter",
                "digit",
                "punctuation",
                "symbol"
              ]
            }
          }
        }
      },
      "mappings": {
        "doc": {
          "properties": {
            "text": {
              "type": "text",
              "analyzer": "my_analyzer",
              "search_analyzer": "simple"
            }
          }
        }
      }
    }
    PUT test/doc/1
    {
      "text": "2 #Quick Foxes lived and died"
    }
    PUT test/doc/2
    {
      "text": "2 #Quick Foxes lived died"
    }
    PUT test/doc/3
    {
      "text": "2 #Quick Foxes lived died and resurrected their wys "
    }
    
    PUT test/doc/6
    {
      "text": "$100 dollars manga #thenga @trump"
    }

```

When we try the query

```
POST test/_refresh
GET test/_search
GET test/doc/_search
{
  "query": {
    "match_phrase": {
      "text": "#Qui"
    }
  }
}

```

Result is

```
   {
      "took": 0,
      "timed_out": false,
      "_shards": {
        "total": 5,
        "successful": 5,
        "skipped": 0,
        "failed": 0
      },
      "hits": {
        "total": 0,
        "max_score": null,
        "hits": []
      }
    }

```

But, When we try this

```
GET test/_search
{
  "query": {
    "match_phrase": {
      "text": "fo"
    }
  }
}

```

Result is

```
{
  "took": 0,
  "timed_out": false,
  "_shards": {
    "total": 5,
    "successful": 5,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": 3,
    "max_score": 1.0247581,
    "hits": [
      {
        "_index": "test",
        "_type": "doc",
        "_id": "2",
        "_score": 1.0247581,
        "_source": {
          "text": "2 #Quick Foxes lived died"
        }
      },
      {
        "_index": "test",
        "_type": "doc",
        "_id": "1",
        "_score": 0.41531453,
        "_source": {
          "text": "2 #Quick Foxes lived and died"
        }
      },
      {
        "_index": "test",
        "_type": "doc",
        "_id": "3",
        "_score": 0.41030136,
        "_source": {
          "text": "2 #Quick Foxes lived died and resurrected their wys "
        }
      }
    ]
  }

```

verifying the analyzer

```
GET test/_analyze
{
  "analyzer": "my_analyzer",
  "text": "2 #Quick Foxes lived and died"
}

```

Result

```
{
  "tokens": [
    {
      "token": "2",
      "start_offset": 0,
      "end_offset": 1,
      "type": "word",
      "position": 0
    },
    {
      "token": "#",
      "start_offset": 2,
      "end_offset": 8,
      "type": "word",
      "position": 1
    },
    {
      "token": "#q",
      "start_offset": 2,
      "end_offset": 8,
      "type": "word",
      "position": 1
    },
    {
      "token": "#qu",
      "start_offset": 2,
      "end_offset": 8,
      "type": "word",
      "position": 1
    },
    {
      "token": "#qui",
      "start_offset": 2,
      "end_offset": 8,
      "type": "word",
      "position": 1
    },
    {
      "token": "#quic",
      "start_offset": 2,
      "end_offset": 8,
      "type": "word",
      "position": 1
    },
    {
      "token": "#quick",
      "start_offset": 2,
      "end_offset": 8,
      "type": "word",
      "position": 1
    },
    {
      "token": "f",
      "start_offset": 9,
      "end_offset": 14,
      "type": "word",
      "position": 2
    },
    {
      "token": "fo",
      "start_offset": 9,
      "end_offset": 14,
      "type": "word",
      "position": 2
    },
    {
      "token": "fox",
      "start_offset": 9,
      "end_offset": 14,
      "type": "word",
      "position": 2
    },
    {
      "token": "foxe",
      "start_offset": 9,
      "end_offset": 14,
      "type": "word",
      "position": 2
    },
    {
      "token": "foxes",
      "start_offset": 9,
      "end_offset": 14,
      "type": "word",
      "position": 2
    },
    {
      "token": "l",
      "start_offset": 15,
      "end_offset": 20,
      "type": "word",
      "position": 3
    },
    {
      "token": "li",
      "start_offset": 15,
      "end_offset": 20,
      "type": "word",
      "position": 3
    },
    {
      "token": "liv",
      "start_offset": 15,
      "end_offset": 20,
      "type": "word",
      "position": 3
    },
    {
      "token": "live",
      "start_offset": 15,
      "end_offset": 20,
      "type": "word",
      "position": 3
    },
    {
      "token": "lived",
      "start_offset": 15,
      "end_offset": 20,
      "type": "word",
      "position": 3
    },
    {
      "token": "d",
      "start_offset": 25,
      "end_offset": 29,
      "type": "word",
      "position": 5
    },
    {
      "token": "di",
      "start_offset": 25,
      "end_offset": 29,
      "type": "word",
      "position": 5
    },
    {
      "token": "die",
      "start_offset": 25,
      "end_offset": 29,
      "type": "word",
      "position": 5
    },
    {
      "token": "died",
      "start_offset": 25,
      "end_offset": 29,
      "type": "word",
      "position": 5
    }
  ]
}

```

How do I include special characters in the search?

---

<div class="post-metadata">

### Author: ![anoopvalluthadam](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anoopvalluthadam/32/30286_2.png) [@anoopvalluthadam](https://discuss.elastic.co/u/anoopvalluthadam)
#### Post date: [February 5, 2018, 7:31am UTC](https://discuss.elastic.co/t/tokenizer-whitespace-not-working-with-edge-ngram/118397/2 "2018-02-05T07:31:46Z")

</div>

@dadoonet any thoughts?

---

<div class="post-metadata">

### Author: ![johtani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/johtani/32/44956_2.png) [@johtani](https://discuss.elastic.co/u/johtani)
#### Post date: [February 5, 2018, 8:23am UTC](https://discuss.elastic.co/t/tokenizer-whitespace-not-working-with-edge-ngram/118397/3 "2018-02-05T08:23:25Z")

</div>

You specified "search\_analyzer" in your settings.  
So, query uses "simple" analyzer for your query.

```auto
GET test/_analyze
{
  "analyzer": "simple",
  "text": "#Qui"
}

```

The index has "#qui", but the query uses "qui".  
Then, you cannot get the result you are expected.

---

<div class="post-metadata">

### Author: ![anoopvalluthadam](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anoopvalluthadam/32/30286_2.png) [@anoopvalluthadam](https://discuss.elastic.co/u/anoopvalluthadam)
#### Post date: [February 5, 2018, 8:35am UTC](https://discuss.elastic.co/t/tokenizer-whitespace-not-working-with-edge-ngram/118397/4 "2018-02-05T08:35:02Z")

</div>

I used my\_analyzer as well, but extra results are getting. Which analyzer can I use?

---

<div class="post-metadata">

### Author: ![anoopvalluthadam](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anoopvalluthadam/32/30286_2.png) [@anoopvalluthadam](https://discuss.elastic.co/u/anoopvalluthadam)
#### Post date: [February 5, 2018, 8:36am UTC](https://discuss.elastic.co/t/tokenizer-whitespace-not-working-with-edge-ngram/118397/5 "2018-02-05T08:36:29Z")

</div>

```
DELETE test
PUT test
{
  "settings": {
    "analysis": {
      "analyzer": {
        "my_analyzer": {
          "type": "custom",
          "tokenizer": "whitespace",
          "filter": [
            "lowercase", "ngram", "asciifolding", "stop"
          ]
        }
      },
      "filter": {
        "ngram": {
          "type": "edge_ngram",
          "min_gram": 1,
          "max_gram": 20,
          "token_chars": [
            "letter",
            "digit",
            "punctuation",
            "symbol"
          ]
        }
      }
    }
  },
  "mappings": {
    "doc": {
      "properties": {
        "text": {
          "type": "text",
          "analyzer": "my_analyzer",
          "search_analyzer": "my_analyzer"
        }
      }
    }
  }
}
PUT test/doc/1
{
  "text": "2 #quick Foxes lived and died"
}
PUT test/doc/2
{
  "text": "2 #Quick Foxes lived died"
}
PUT test/doc/3
{
  "text": "2 #Quick Foxes lived died and resurrected their wys "
}
PUT test/doc/6
{
  "text": "$100 dollars manga #thenga @trump"
}
POST test/_refresh
GET test/_search
GET test/doc/_search
{
  "query": {
    "match_phrase": {
      "text": "#qu"
    }
  }
}

```

Result is

```
{
  "took": 0,
  "timed_out": false,
  "_shards": {
    "total": 5,
    "successful": 5,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": 4,
    "max_score": 0.9152459,
    "hits": [
      {
        "_index": "test",
        "_type": "doc",
        "_id": "2",
        "_score": 0.9152459,
        "_source": {
          "text": "2 #Quick Foxes lived died"
        }
      },
      {
        "_index": "test",
        "_type": "doc",
        "_id": "6",
        "_score": 0.7442766,
        "_source": {
          "text": "$100 dollars manga #thenga @trump"
        }
      },
      {
        "_index": "test",
        "_type": "doc",
        "_id": "1",
        "_score": 0.5388059,
        "_source": {
          "text": "2 #quick Foxes lived and died"
        }
      },
      {
        "_index": "test",
        "_type": "doc",
        "_id": "3",
        "_score": 0.53597397,
        "_source": {
          "text": "2 #Quick Foxes lived died and resurrected their wys "
        }
      }
    ]
  }
}

```

In this

```
{
        "_index": "test",
        "_type": "doc",
        "_id": "6",
        "_score": 0.7442766,
        "_source": {
          "text": "$100 dollars manga #thenga @trump"
        }
      }

```

is wrong, isn't it?

---

<div class="post-metadata">

### Author: ![johtani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/johtani/32/44956_2.png) [@johtani](https://discuss.elastic.co/u/johtani)
#### Post date: [February 5, 2018, 8:55am UTC](https://discuss.elastic.co/t/tokenizer-whitespace-not-working-with-edge-ngram/118397/6 "2018-02-05T08:55:11Z")

</div>

because, you use ngram from 1 to 20.

you can see what your query is with "explain" param.

```auto
GET test/doc/_search?explain=true
{
  "query": {
    "match_phrase": {
      "text": "#qu"
    }
  }
}

```

Using "my\_analyzer" with query, your query is "#" or "#q" or "#qu".  
I'm not sure what your requirement in your query...  
How about using "whitespace" tokenizer + "lowercase" for search\_analyzer?

---

<div class="post-metadata">

### Author: ![anoopvalluthadam](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anoopvalluthadam/32/30286_2.png) [@anoopvalluthadam](https://discuss.elastic.co/u/anoopvalluthadam)
#### Post date: [February 5, 2018, 9:01am UTC](https://discuss.elastic.co/t/tokenizer-whitespace-not-working-with-edge-ngram/118397/7 "2018-02-05T09:01:21Z")

</div>

Requirements is something like this:

text will be

> $100 dollars manga #thenga  
> 2 #Quick Foxes lived died

and when I search `$10`, result should be

> $100 dollars manga #thenga

and when I search `ied` , result should be

> 2 #Quick Foxes lived died

you can treat it like a replacement of wildcard 🙂

---

<div class="post-metadata">

### Author: ![johtani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/johtani/32/44956_2.png) [@johtani](https://discuss.elastic.co/u/johtani)
#### Post date: [February 5, 2018, 9:20am UTC](https://discuss.elastic.co/t/tokenizer-whitespace-not-working-with-edge-ngram/118397/8 "2018-02-05T09:20:03Z")

</div>

you should read [https://www.elastic.co/guide/en/elasticsearch/guide/current/full-text-search.html](https://www.elastic.co/guide/en/elasticsearch/guide/current/full-text-search.html) first.

And `ied` does not work with edge\_ngram.  
You can not see `ied` in the following result:

```auto
GET test/_analyze
{
  "field": "text",
  "text": "2 #Quick Foxes lived died"
}

```

And not good for elasticsearch with wildcard especially using a pattern that starts with a wildcard...  
[https://www.elastic.co/guide/en/elasticsearch/guide/2.x/\_wildcard\_and\_regexp\_queries.html](https://www.elastic.co/guide/en/elasticsearch/guide/2.x/_wildcard_and_regexp_queries.html)

---

<div class="post-metadata">

### Author: ![anoopvalluthadam](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anoopvalluthadam/32/30286_2.png) [@anoopvalluthadam](https://discuss.elastic.co/u/anoopvalluthadam)
#### Post date: [February 5, 2018, 9:20am UTC](https://discuss.elastic.co/t/tokenizer-whitespace-not-working-with-edge-ngram/118397/9 "2018-02-05T09:20:56Z")

</div>

Yeah

> How about using "whitespace" tokenizer + "lowercase" for search\_analyzer?

is the solution.

i am making the \* concept using ngram and edge\_ngram

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [March 5, 2018, 9:20am UTC](https://discuss.elastic.co/t/tokenizer-whitespace-not-working-with-edge-ngram/118397/10 "2018-03-05T09:20:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
