# Generate\_number\_parts not working as expected

**URL:** <https://discuss.elastic.co/t/generate-number-parts-not-working-as-expected/123468>\
**Category:** Elasticsearch\
**Created:** [March 12, 2018, 2:21am UTC](https://discuss.elastic.co/t/generate-number-parts-not-working-as-expected/123468 "2018-03-12T02:21:53Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![askids](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/askids/32/22252_2.png) [@askids](https://discuss.elastic.co/u/askids)\
**Post date:** [March 12, 2018, 2:21am UTC](https://discuss.elastic.co/t/generate-number-parts-not-working-as-expected/123468/1 "2018-03-12T02:21:53Z")

</div>

hi,

I have a requirement to include special characters in search. So for both intake and search, I am trying to create a custom analyzer so that I can retain original character as is. I have set the tokenizer to whitespace and generate\_number\_parts to false. But when I check the analyzer output, its not working as expected.

```
PUT /testing
{
  "settings": {
    "index": {
      "analysis": {
        "filter": {
          "word_delimiter_v3_filter": {
            "type": "word_delimiter",
            "generate_number_parts ": false,
            "split_on_numerics": false,
            "preserve_original": true
          }
        },
        "analyzer": {
          "searchword_v3_analyzer": {
            "filter": [
              "lowercase",
              "word_delimiter_v3_filter"
            ],
            "type": "custom",
            "tokenizer": "whitespace"
          }
        }
      }
    }
  },
  "mappings": {
    "testmap": {
      "properties": {
        "fullname": {
          "type": "text",
          "analyzer": "searchword_v3_analyzer",
          "search_analyzer": "searchword_v3_analyzer"
        }
      }
    }
  }
}

```

Using above analyzer, if I enter "2-10" as search text, I want it to be searched as is.

```
GET testing/_analyze 
{
  "analyzer": "searchword_v3_analyzer", 
  "text": "2-10"
}

```

But when I check the analyzer output, its still splitting 2 and 10 as separate words, even when I am using white space tokenizer and have set generate\_number\_parts to false and split\_on\_numerics to false.

I am running v5.5.1 of ES on Windows 2012.

```
{
  "tokens": [
    {
      "token": "2-10",
      "start_offset": 0,
      "end_offset": 4,
      "type": "word",
      "position": 0
    },
    {
      "token": "2",
      "start_offset": 0,
      "end_offset": 1,
      "type": "word",
      "position": 0
    },
    {
      "token": "10",
      "start_offset": 2,
      "end_offset": 4,
      "type": "word",
      "position": 1
    }
  ]
}

```

Thanks  
askids

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 12, 2018, 8:51am UTC](https://discuss.elastic.co/t/generate-number-parts-not-working-as-expected/123468/2 "2018-03-12T08:51:50Z")

</div>

Why not doing this:

```auto
DELETE testing
PUT /testing
{
  "settings": {
    "index": {
      "analysis": {
        "analyzer": {
          "searchword_v3_analyzer": {
            "filter": [
              "lowercase"
            ],
            "type": "custom",
            "tokenizer": "whitespace"
          }
        }
      }
    }
  },
  "mappings": {
    "testmap": {
      "properties": {
        "fullname": {
          "type": "text",
          "analyzer": "searchword_v3_analyzer",
          "search_analyzer": "searchword_v3_analyzer"
        }
      }
    }
  }
}
GET testing/_analyze 
{
  "analyzer": "searchword_v3_analyzer", 
  "text": "2-10"
}

```

It gives:

```auto
{
  "tokens": [
    {
      "token": "2-10",
      "start_offset": 0,
      "end_offset": 4,
      "type": "word",
      "position": 0
    }
  ]
}

```

PS: Thanks for the reproduction script. I wish everybody provide this.

---

<div class="post-metadata">

**Author:** ![askids](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/askids/32/22252_2.png) [@askids](https://discuss.elastic.co/u/askids)\
**Post date:** [March 14, 2018, 12:44am UTC](https://discuss.elastic.co/t/generate-number-parts-not-working-as-expected/123468/3 "2018-03-14T00:44:21Z")

</div>

Thanks @dadoonet. Actually I figured it out a day after I posted it that I was unnecessarily getting into all that complexity with word delimiter filter, when simply using whitespace tokenizer would meet my requirement. I came back today to just update my post, but I see that you have suggested the same. Thank you!!!

However, I had a secondary question for you. When I was using the whitespace tokenizer along with word delimiter filter, why would it generate more tokens as even the custom word\_delimiter filter, though redundant, should have given same result. But now, effectively, looks like its overriding whitespace tokenizer, but still not following all the rules set in the filter. Only rule that worked was the preserver original. Looks like a bug to me.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 11, 2018, 12:44am UTC](https://discuss.elastic.co/t/generate-number-parts-not-working-as-expected/123468/4 "2018-04-11T00:44:41Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
