# Indexing Special characters as symbols and searching with unicode values

**URL:** <https://discuss.elastic.co/t/indexing-special-characters-as-symbols-and-searching-with-unicode-values/60711>\
**Category:** Elasticsearch\
**Created:** [September 16, 2016, 1:08pm UTC](https://discuss.elastic.co/t/indexing-special-characters-as-symbols-and-searching-with-unicode-values/60711 "2016-09-16T13:08:23Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![premkumar3438](https://avatars.discourse-cdn.com/v4/letter/p/f19dbf/32.png) [@premkumar3438](https://discuss.elastic.co/u/premkumar3438)\
**Post date:** [September 16, 2016, 1:08pm UTC](https://discuss.elastic.co/t/indexing-special-characters-as-symbols-and-searching-with-unicode-values/60711/1 "2016-09-16T13:08:23Z")

</div>

I am Trying to index and search for Special characters in elastic search  
I used White space tokenizer and i am able to index Special characters and search them fine.

But i have a situation where i need to index the special character Symbol and search it using its equivalent unicode values  
Example:  
I am indexing the below document

```json
{
   "id": 1,
   "documentId": "334567"
   "fieldValue": [
      {
         "fieldId": 175699,
         "textValue": [{
         "paragraph":"@doc"
         },{
         "paragraph":"γcomp"// this a lowercased gamma symbol
         },
         {
         "paragraph":"@Keyboard"
         }
         ],
         "integerValue": "",
         "numericValue": "",
         "modifiedDate": "2010-01-01",
         "modifiedUser": "Tr"
      }
   ]
}

```

Now i want to search "γcomp" Using "γ" unicode value "&#947comp" but it is not working

Can anybody please help with this?

---

<div class="post-metadata">

**Author:** ![polyfractal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/polyfractal/32/48162_2.png) [@polyfractal](https://discuss.elastic.co/u/polyfractal)\
**Post date:** [September 16, 2016, 1:48pm UTC](https://discuss.elastic.co/t/indexing-special-characters-as-symbols-and-searching-with-unicode-values/60711/2 "2016-09-16T13:48:50Z")

</div>

There is not automatic conversion of HTML entities into their corresponding unicode. Elasticsearch expects UTF-8 encoded strings, so encoding schemes like HTML entities and url-encoding just look like valid UTF-8 and are indexed/searched as they are.

If you need to convert HTML entities into their corresponding unicode characters, I'd probably just run the conversion in my application.

If you have a relatively small list of characters you need to convert on a regular basis, you can use the char\_filter as described here: [https://www.elastic.co/guide/en/elasticsearch/guide/current/char-filters.html#\_tidying\_up\_punctuation](https://www.elastic.co/guide/en/elasticsearch/guide/current/char-filters.html#_tidying_up_punctuation)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:19pm UTC](https://discuss.elastic.co/t/indexing-special-characters-as-symbols-and-searching-with-unicode-values/60711/3 "2017-07-05T22:19:38Z")

</div>


