# Tokenizing a Hashtag containing underscores

**URL:** <https://discuss.elastic.co/t/tokenizing-a-hashtag-containing-underscores/64735>\
**Category:** Elasticsearch\
**Created:** [November 2, 2016, 3:38pm UTC](https://discuss.elastic.co/t/tokenizing-a-hashtag-containing-underscores/64735 "2016-11-02T15:38:29Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![winder](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/winder/32/54090_2.png) [@winder](https://discuss.elastic.co/u/winder)\
**Post date:** [November 2, 2016, 3:38pm UTC](https://discuss.elastic.co/t/tokenizing-a-hashtag-containing-underscores/64735/1 "2016-11-02T15:38:29Z")

</div>

My application has been tokenizing twitter style hashtags with a word\_delimiter filter for a long time. For a hashtag '#SomeHashtag' we expect users might search for '#SomeHashtag' or 'SomeHashtag'.

For ' **#SomeHashtag**' our analyzer produces the following tokens:  
`#SomeHashtag`, `Some`, `SomeHashtag`, `Hashtag`

For ' **#Some\_Hashtag**', which is common in some different languages, our analyzer removes the underscore from all tokens except the original:  
`#Some_Hashtag`, `Some`, `SomeHashtag`, `Hashtag`

Here is our analyzer:

```
"analysis": {
  "analyzer": {
    "tweet_test": {
      "type": "custom",
      "char_filter": ["html_strip", "quotes"],
      "tokenizer": "standard_custom",
      "filter": ["custom_text_word_delimiter_query"]
    }
  },
  "filter": {
    "custom_text_word_delimiter_query": {
      "type": "word_delimiter",
      "generate_word_parts": "0",
      "generate_number_parts": "0",
      "catenate_words": "1",
      "catenate_numbers": "1",
      "catenate_all": "0",
      "split_on_case_change": "0",
      "split_on_numerics": "0",
      "preserve_original": "0",
      "type_table": [
          "# => ALPHA",
          "@ => ALPHA",
          "& => ALPHA",
          "- => ALPHA",
          ". => ALPHA",
          "/ => ALPHA",
          "_ => ALPHA"
      ]
    }
  }
}

```

One solution I'm considering is an additional pattern\_capture filter (the actual regular expression twitter uses is MUCH longer than this):

```
"hashtag_filter": {
    "type" : "pattern_capture",
    "preserve_original" : 1,
    "patterns" : ["#([^\\s]*)"]
}

```

This application may index several hundred messages per second, so my questions are:

Should I be concerned about the performance of a regular expression match?  
Is there an alternative approach that might be less expensive than a regular expression match?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:07pm UTC](https://discuss.elastic.co/t/tokenizing-a-hashtag-containing-underscores/64735/2 "2017-07-05T22:07:19Z")

</div>


