# A number followed by a dot is considered a word break?

**URL:** <https://discuss.elastic.co/t/a-number-followed-by-a-dot-is-considered-a-word-break/368333>\
**Category:** Elasticsearch\
**Created:** [October 7, 2024, 9:04am UTC](https://discuss.elastic.co/t/a-number-followed-by-a-dot-is-considered-a-word-break/368333 "2024-10-07T09:04:17Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Hrusha](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hrusha/32/123392_2.png) [@Hrusha](https://discuss.elastic.co/u/Hrusha)\
**Post date:** [October 7, 2024, 9:04am UTC](https://discuss.elastic.co/t/a-number-followed-by-a-dot-is-considered-a-word-break/368333/1 "2024-10-07T09:04:17Z")

</div>

I am looking at very strange (to me at least) behavior and am unable to find clear documentation for it.

```auto
GET _analyze
{
  "analyzer": "standard",
  "text": "server.mycopany.com"
}

```

Produces

```auto
{
  "tokens": [
    {
      "token": "server.mycopany.com",
      "start_offset": 0,
      "end_offset": 19,
      "type": "<ALPHANUM>",
      "position": 0
    }
  ]
}

```

So far so good. I do not expect the standard tokenzier to break on dots without w/s.

But when a number creeps in before a dot:

```auto
GET _analyze
{
  "analyzer": "standard",
  "text": "server1.mycopany.com"
}

```

This happens:

```auto
{
  "tokens": [
    {
      "token": "server1",
      "start_offset": 0,
      "end_offset": 7,
      "type": "<ALPHANUM>",
      "position": 0
    },
    {
      "token": "mycopany.com",
      "start_offset": 8,
      "end_offset": 20,
      "type": "<ALPHANUM>",
      "position": 1
    }
  ]
}

```

1. Why is this happening?
2. How to overcome this? ==\> I dont want a number followed by a dot to break the word.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 7, 2024, 10:24am UTC](https://discuss.elastic.co/t/a-number-followed-by-a-dot-is-considered-a-word-break/368333/2 "2024-10-07T10:24:44Z")

</div>

Yes that's the [expected behavior](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-standard-tokenizer.html).  
Note that if you analyze `server1. mycopany.`, it will produces 2 tokens `server1` and `mycopany`.

If you want to keep everything as is, you need to use a `keyword` tokenizer.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 7, 2024, 10:24am UTC](https://discuss.elastic.co/t/a-number-followed-by-a-dot-is-considered-a-word-break/368333/3 "2024-10-07T10:24:54Z")

</div>

From #Elastic Search to #Elasticsearch

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 7, 2024, 10:35am UTC](https://discuss.elastic.co/t/a-number-followed-by-a-dot-is-considered-a-word-break/368333/4 "2024-10-07T10:35:47Z")

</div>

You can also try creating a custom analyser that handles this the way you want, e.g. first use a [pattern replace token filter](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-pattern_replace-tokenfilter.html) to replace any full stops (maybe also commas and other special characters?) followed by whitespace with just a white space and then apply a [whitespace tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-whitespace-tokenizer.html). That should solve the issue you are seeing now but may require some tweaking to handle other scenarios/edge cases.

---

<div class="post-metadata">

**Author:** ![Hrusha](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hrusha/32/123392_2.png) [@Hrusha](https://discuss.elastic.co/u/Hrusha)\
**Post date:** [October 7, 2024, 3:47pm UTC](https://discuss.elastic.co/t/a-number-followed-by-a-dot-is-considered-a-word-break/368333/5 "2024-10-07T15:47:28Z")

</div>

I see. Thanks!

According to the spec, dot delimited letter sequences (abra.kadabra) or number sequences (74.75) are not broken down.

But if I have abra1.kadabra, this is no longer considered a letter sequence.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 4, 2024, 3:48pm UTC](https://discuss.elastic.co/t/a-number-followed-by-a-dot-is-considered-a-word-break/368333/6 "2024-11-04T15:48:26Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
