# Configuring icu\_tokenizer to keep hashtag in token

**URL:** <https://discuss.elastic.co/t/configuring-icu-tokenizer-to-keep-hashtag-in-token/129529>\
**Category:** Elasticsearch\
**Created:** [April 25, 2018, 3:38pm UTC](https://discuss.elastic.co/t/configuring-icu-tokenizer-to-keep-hashtag-in-token/129529 "2018-04-25T15:38:25Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Robert\_Fiser1](https://avatars.discourse-cdn.com/v4/letter/r/6bbea6/32.png) [@Robert\_Fiser1](https://discuss.elastic.co/u/Robert_Fiser1)\
**Post date:** [April 25, 2018, 3:38pm UTC](https://discuss.elastic.co/t/configuring-icu-tokenizer-to-keep-hashtag-in-token/129529/1 "2018-04-25T15:38:26Z")

</div>

Hi,  
we'r using icu\_tokenizer to analyze text which may be in many languages. The problem is that text contains hastags like #dog #cat etc. and icu\_tokenizer removes the '#' characters from tokens. So we'r not able to find documents which contains exactly the '#cat'.  
Is there a simple way to achieve calling \_analyze text:'#cat' produces 2 tokens: ['#cat', 'cat']?  
Robert

---

<div class="post-metadata">

**Author:** ![Robert\_Fiser1](https://avatars.discourse-cdn.com/v4/letter/r/6bbea6/32.png) [@Robert\_Fiser1](https://discuss.elastic.co/u/Robert_Fiser1)\
**Post date:** [May 16, 2018, 12:15pm UTC](https://discuss.elastic.co/t/configuring-icu-tokenizer-to-keep-hashtag-in-token/129529/2 "2018-05-16T12:15:33Z")

</div>

Any idea?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 13, 2018, 12:15pm UTC](https://discuss.elastic.co/t/configuring-icu-tokenizer-to-keep-hashtag-in-token/129529/3 "2018-06-13T12:15:35Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
