# Searching for Emoji Characters / Unicode

**URL:** <https://discuss.elastic.co/t/searching-for-emoji-characters-unicode/31625>\
**Category:** Elasticsearch\
**Created:** [October 5, 2015, 11:36am UTC](https://discuss.elastic.co/t/searching-for-emoji-characters-unicode/31625 "2015-10-05T11:36:35Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ben\_M](https://avatars.discourse-cdn.com/v4/letter/b/ee7513/32.png) [@Ben\_M](https://discuss.elastic.co/u/Ben_M)\
**Post date:** [October 5, 2015, 11:36am UTC](https://discuss.elastic.co/t/searching-for-emoji-characters-unicode/31625/1 "2015-10-05T11:36:35Z")

</div>

I have some articles in my ES index that contain emoji characters. I'd like to perform a search for articles that contain specific emojis. Via ElasticHQ I can see the emojis are in the data - they are rendered as icons in OS X, so I assume the unicode data is stored correctly. However, when I run the query below I get no results. If I run a plain text search I do get my results. From my very limited experience of ES (1 day) I'm guessing I need to add an Analyzer that can handle this? I don't know where to start or if this is a correct diagnosis.

Best,  
Ben

```
  client.search({
            index: 'articles',
            body: {
                fields: ["code", "title"],
                query: {
                    query_string: {
                        query:"😕"
                    }
                }
            }
        }, function(error, response) {
            // handling error / response here ...
        });
        return;
    }
    
    //results: {"took":6,"timed_out":false,"_shards":{"total":5,"successful":5,"failed":0},"hits":{"total":0,"max_score":null,"hits":[]}}
```

---

<div class="post-metadata">

**Author:** ![rmuir](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rmuir/32/44949_2.png) [@rmuir](https://discuss.elastic.co/u/rmuir)\
**Post date:** [October 5, 2015, 11:49am UTC](https://discuss.elastic.co/t/searching-for-emoji-characters-unicode/31625/2 "2015-10-05T11:49:42Z")

</div>

The default analyzer (StandardAnalyzer) uses the unicode word break algorithm ([http://unicode.org/reports/tr29/](http://unicode.org/reports/tr29/)).

The properties assigned to emoji are just "other", so they are treated no differently than "trash" like ^, , etc. Sorry this would be my personal opinion of them, too.

Anyway, yes you will need a custom analyzer if you want to make sense of emoji. You will have to decide how to make sense of them, e.g. if each should be its own word, or if someone writes 87 smileys in a row, what should happen then, and so on.

---

<div class="post-metadata">

**Author:** ![Ben\_M](https://avatars.discourse-cdn.com/v4/letter/b/ee7513/32.png) [@Ben\_M](https://discuss.elastic.co/u/Ben_M)\
**Post date:** [October 5, 2015, 12:02pm UTC](https://discuss.elastic.co/t/searching-for-emoji-characters-unicode/31625/3 "2015-10-05T12:02:30Z")

</div>

We work in the messenger app space, so emoji is a character set we need to support. For now I'd be happy to search for individual emoji characters. And maybe later on have something more complex that understands the context of a series of emoji. Which analyzer would I use in the basic scenario and where would I learn to integrate it? Many thanks.

---

<div class="post-metadata">

**Author:** ![rmuir](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rmuir/32/44949_2.png) [@rmuir](https://discuss.elastic.co/u/rmuir)\
**Post date:** [October 5, 2015, 1:51pm UTC](https://discuss.elastic.co/t/searching-for-emoji-characters-unicode/31625/4 "2015-10-05T13:51:32Z")

</div>

There is not such an analyzer. This stuff is hardly stable and not really standardized yet, so there is not yet well-accepted best practices and so on. You may have to write custom code, unless you want to do something very simple like use mappingcharfilter with mappings like 📀 -\> DVD

There are a lot of possibilities depending on use case, but its nowhere near the maturity level where we have incorporated anything in lucene, ready to support backwards compatibility for it, etc.

You can get some ideas from [http://unicode.org/reports/tr51/#Searching](http://unicode.org/reports/tr51/#Searching) and think about how you want it to work for your app.

---

<div class="post-metadata">

**Author:** ![Ben\_M](https://avatars.discourse-cdn.com/v4/letter/b/ee7513/32.png) [@Ben\_M](https://discuss.elastic.co/u/Ben_M)\
**Post date:** [October 5, 2015, 2:27pm UTC](https://discuss.elastic.co/t/searching-for-emoji-characters-unicode/31625/5 "2015-10-05T14:27:47Z")

</div>

Thanks so much for your help. Ben

---

<div class="post-metadata">

**Author:** ![Damien\_Alexandre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/damien_alexandre/32/21566_2.png) [@Damien\_Alexandre](https://discuss.elastic.co/u/Damien_Alexandre)\
**Post date:** [March 16, 2016, 8:40am UTC](https://discuss.elastic.co/t/searching-for-emoji-characters-unicode/31625/6 "2016-03-16T08:40:23Z")

</div>

> [@Ben\_M](#):
>
> Which analyzer would I use in the basic scenario and where would I learn to integrate it? Many thanks.

For my needs, I used a [whitespace tokenizer with quite a lot of customizations](http://jolicode.com/blog/search-for-emoji-with-elasticsearch). Did you came up with a good solution?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:07pm UTC](https://discuss.elastic.co/t/searching-for-emoji-characters-unicode/31625/7 "2017-07-05T23:07:59Z")

</div>


