# Tokenize email address

**URL:** <https://discuss.elastic.co/t/tokenize-email-address/6666>\
**Category:** Elasticsearch\
**Created:** [February 10, 2012, 1:06am UTC](https://discuss.elastic.co/t/tokenize-email-address/6666 "2012-02-10T01:06:27Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ask\_Bjorn\_Hansen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ask_bjorn_hansen/32/2995_2.png) [@Ask\_Bjorn\_Hansen](https://discuss.elastic.co/u/Ask_Bjorn_Hansen)\
**Post date:** [February 10, 2012, 1:06am UTC](https://discuss.elastic.co/t/tokenize-email-address/6666/1 "2012-02-10T01:06:27Z")

</div>

Hi everyone,

We're indexing a user database for an admin interface and would like to  
search on email addresses. It works fine except for email addresses with  
'.'s in them where our users expect to be able to search on either name  
("some.user@domain" should match "some" or "user" or "domain"). Anyway, we  
can't get any of the built-in tokenizers to split on the "." or on "\_" for  
that matter. The standard tokenizer works as expected on "-", but most  
email addresses have dots. 🙂

Any tips? What am I doing wrong?

Ask

$ curl  
'[http://indexdev1.la.sol:9200/us-devel-rms-v1/\_analyze?pretty=1&tokenizer=uax\_url\_email](http://indexdev1.la.sol:9200/us-devel-rms-v1/_analyze?pretty=1&tokenizer=uax_url_email)'  
-d 'some.user@domain'  
{  
"tokens" : [ {  
"token" : "some.user",  
"start\_offset" : 0,  
"end\_offset" : 9,  
"type" : "",  
"position" : 1  
}, {  
"token" : "domain",  
"start\_offset" : 10,  
"end\_offset" : 16,  
"type" : "",  
"position" : 2  
} ]  
}

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [February 10, 2012, 11:59am UTC](https://discuss.elastic.co/t/tokenize-email-address/6666/2 "2012-02-10T11:59:27Z")

</div>

Hi Ask

> We're indexing a user database for an admin interface and would like  
> to search on email addresses. It works fine except for email  
> addresses with '.'s in them where our users expect to be able to  
> search on either name ("some.user@domain" should match "some" or  
> "user" or "domain"). Anyway, we can't get any of the built-in  
> tokenizers to split on the "." or on "\_" for that matter. The  
> standard tokenizer works as expected on "-", but most email addresses  
> have dots. 🙂

If your field contains just email addresses, then use the 'simple'  
analyzer.

If your email addresses are embedded in other text, then you may want to  
do something smarter with the (undocumented) pattern replace filter:

> <https://github.com/elastic/elasticsearch/pull/1108>
>
> This adds the Pattern Replace Filter to Elasticsearch.
> 
> It allows to easily hand…le string replacement during the analysis process using the power of regular expression. 
> 
> The regular expression pattern has to be specified with the setting "pattern".
> 
> The replacement expression has to be specified with the setting "replacement".
> 
> \<pre\>
> 
> index:
> analysis:
> analyzer:
> default:
> tokenizer: standard
> filter: \[pattern\_replace\]
> filter:
> pattern\_replace:
> type: pattern\_replace
> pattern: (?&lt;=\[\\d\])(,)(?=\[\\d\])
> 
> \</pre\>
> 
> 
> By default the replacement expression is an empty string.

clint

---

<div class="post-metadata">

**Author:** ![Ask\_Bjorn\_Hansen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ask_bjorn_hansen/32/2995_2.png) [@Ask\_Bjorn\_Hansen](https://discuss.elastic.co/u/Ask_Bjorn_Hansen)\
**Post date:** [February 14, 2012, 5:57am UTC](https://discuss.elastic.co/t/tokenize-email-address/6666/3 "2012-02-14T05:57:44Z")

</div>

On Feb 10, 3:59 am, Clinton Gormley [cl...@traveljury.com](mailto:cl...@traveljury.com) wrote:

> If your field contains justemailaddresses, then use the 'simple'  
> analyzer.

Brilliant, thank you! This works perfectly.

I have a variation of this now - we have another query where we need  
to make sure that the email address matches exactly, except for case  
variations. I'm indexing it with the email mapping below now, but  
that requires me to lower case the input (and thus it gets messed up  
for display in \_source).

I have another field too (an ID type field) where I need to be able to  
filter for it case-insensitively, but keep it in the \_source as the  
original. There I have two fields in my mapping, like:

```
                provider_uid => {type => "string", index =>

```

"not\_analyzed"},  
provider\_uid\_lc =\> {type =\> "string", index =\>  
"not\_analyzed"},

The \_lc version is excluded from \_source and I lowercase it when  
indexing, and then use provider\_uid for display etc in my  
application. Is there a better way? It seems like the multi\_field is  
what I want, but I couldn't figure out how to lowercase, but not  
tokenize.

Here's the email mapping:

```
                email => {
                    type => "multi_field",
                    fields => {
                        "email" => {
                            type => "string",
                            index => "analyzed",
                            analyzer => "simple",
                        },
                        "raw" => {
                            type => "string",
                            index => "not_analyzed",
                        },
                    }
                },

```

Thanks again! The excellent software wouldn't be as good as it is  
without the great community helping the lost and clueless such as  
myself. 🙂

Ask

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [February 14, 2012, 9:40am UTC](https://discuss.elastic.co/t/tokenize-email-address/6666/4 "2012-02-14T09:40:01Z")

</div>

> I have a variation of this now - we have another query where we need  
> to make sure that the email address matches exactly, except for case  
> variations. I'm indexing it with the email mapping below now, but  
> that requires me to lower case the input (and thus it gets messed up  
> for display in \_source).
> 
> The \_lc version is excluded from \_source and I lowercase it when  
> indexing, and then use provider\_uid for display etc in my  
> application. Is there a better way? It seems like the multi\_field is  
> what I want, but I couldn't figure out how to lowercase, but not  
> tokenize.

Yes, a multi-field is the way to go. You need a custom analyzer which  
uses the 'keyword' tokenizer and the 'lowercase' filter:

> <https://gist.github.com/clintongormley/1825357>

clint

---

<div class="post-metadata">

**Author:** ![Ask\_Bjorn\_Hansen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ask_bjorn_hansen/32/2995_2.png) [@Ask\_Bjorn\_Hansen](https://discuss.elastic.co/u/Ask_Bjorn_Hansen)\
**Post date:** [February 15, 2012, 12:10am UTC](https://discuss.elastic.co/t/tokenize-email-address/6666/5 "2012-02-15T00:10:40Z")

</div>

On Feb 14, 1:40 am, Clinton Gormley [cl...@traveljury.com](mailto:cl...@traveljury.com) wrote:

> Yes, a multi-field is the way to go. You need a custom analyzer which  
> uses the 'keyword'tokenizerand the 'lowercase' filter:
> 
> [gist:1825357 · GitHub](https://gist.github.com/1825357)

Brilliant, thank you! I completely missed that the 'keyword' analyzer  
is the 'null tokenizer' I couldn't find.

Ask

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:39am UTC](https://discuss.elastic.co/t/tokenize-email-address/6666/6 "2017-07-06T03:39:19Z")

</div>


