# Special Characters not indexed and hence not searchable

**URL:** <https://discuss.elastic.co/t/special-characters-not-indexed-and-hence-not-searchable/8530>\
**Category:** Elasticsearch\
**Created:** [July 26, 2012, 6:28pm UTC](https://discuss.elastic.co/t/special-characters-not-indexed-and-hence-not-searchable/8530 "2012-07-26T18:28:46Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Praveen\_Kariyanahall](https://avatars.discourse-cdn.com/v4/letter/p/4491bb/32.png) [@Praveen\_Kariyanahall](https://discuss.elastic.co/u/Praveen_Kariyanahall)\
**Post date:** [July 26, 2012, 6:28pm UTC](https://discuss.elastic.co/t/special-characters-not-indexed-and-hence-not-searchable/8530/1 "2012-07-26T18:28:46Z")

</div>

I saw many threads discuss it. I have crossed few hurdles, last one is  
still bothering me. I had email address, ":", "/" in my data (which need to  
be indexed and searched). Now I am able to search the email address  
[test@myco.com](mailto:test@myco.com), _but I still cannot have the following characters indexed:  
":" "/" and "-"_. Any help is greatly appreciated. Is it issue with my  
index\_analyzer or search\_analyzer?

Thanks in Advance  
-Praveen

Here is my mapping:

ESINDEX = {  
"number\_of\_shards": 1,  
"analysis": {  
"filter": {  
"mynGram" : {  
"type" : "nGram",  
"min\_gram": 1,  
"max\_gram": 50  
}  
},  
"analyzer": {  
"a1" : {  
"type" :"custom",  
"tokenizer":"uax\_url\_email",  
"filter" : ["mynGram"]  
}  
}  
}  
}

ESMAPPINGS = {  
"index\_analyzer" : "a1",  
"search\_analyzer" : "whitespace",  
"properties" : {  
u'test\_field1' : {  
'index' : 'not\_analyzed',  
'type' : u'string',  
'store' : 'yes'  
},  
u'testfield2' : {  
'index' : 'not\_analyzed',  
'type' : u'string',  
'store' : 'yes'  
},  
u'email' : {  
'index': 'not\_analyzed',  
'type' : u'string',  
'store': 'yes'  
},  
:::::::::::  
}

---

<div class="post-metadata">

**Author:** ![Joe\_Wong](https://avatars.discourse-cdn.com/v4/letter/j/e495f1/32.png) [@Joe\_Wong](https://discuss.elastic.co/u/Joe_Wong)\
**Post date:** [July 26, 2012, 8:53pm UTC](https://discuss.elastic.co/t/special-characters-not-indexed-and-hence-not-searchable/8530/2 "2012-07-26T20:53:57Z")

</div>

Do you mean you want to have ":", "/" searchable?  
if so you may have to use a custom tokenizer and specify a pattern to  
tokenize.

This will include all unicode characters plus the special characters such  
as "/" and ":"

'tokenizer' : {

```
        'email_tokenizer' : {

            'type' : { 'pattern',

            'pattern' => "[@*\/*:*\.*\\w\\p{L}]+"

             }

        }

    }

```

On Thursday, July 26, 2012 2:28:46 PM UTC-4, Praveen Kariyanahalli wrote:

> I saw many threads discuss it. I have crossed few hurdles, last one is  
> still bothering me. I had email address, ":", "/" in my data (which need to  
> be indexed and searched). Now I am able to search the email address  
> [test@myco.com](mailto:test@myco.com), _but I still cannot have the following characters indexed:  
> ":" "/" and "-"_. Any help is greatly appreciated. Is it issue with my  
> index\_analyzer or search\_analyzer?
> 
> Thanks in Advance  
> -Praveen
> 
> Here is my mapping:
> 
> ESINDEX = {  
> "number\_of\_shards": 1,  
> "analysis": {  
> "filter": {  
> "mynGram" : {  
> "type" : "nGram",  
> "min\_gram": 1,  
> "max\_gram": 50  
> }  
> },  
> "analyzer": {  
> "a1" : {  
> "type" :"custom",  
> "tokenizer":"uax\_url\_email",  
> "filter" : ["mynGram"]  
> }  
> }  
> }  
> }
> 
> ESMAPPINGS = {  
> "index\_analyzer" : "a1",  
> "search\_analyzer" : "whitespace",  
> "properties" : {  
> u'test\_field1' : {  
> 'index' : 'not\_analyzed',  
> 'type' : u'string',  
> 'store' : 'yes'  
> },  
> u'testfield2' : {  
> 'index' : 'not\_analyzed',  
> 'type' : u'string',  
> 'store' : 'yes'  
> },  
> u'email' : {  
> 'index': 'not\_analyzed',  
> 'type' : u'string',  
> 'store': 'yes'  
> },  
> :::::::::::  
> }

---

<div class="post-metadata">

**Author:** ![Praveen\_Kariyanahall](https://avatars.discourse-cdn.com/v4/letter/p/4491bb/32.png) [@Praveen\_Kariyanahall](https://discuss.elastic.co/u/Praveen_Kariyanahall)\
**Post date:** [July 26, 2012, 9:57pm UTC](https://discuss.elastic.co/t/special-characters-not-indexed-and-hence-not-searchable/8530/3 "2012-07-26T21:57:36Z")

</div>

Hi Joe

As per you suggestion, I changed my tokenizer to the following. But it  
doesnt help. I have lost my initial ngram indexing too? Am I missing  
something?

ESINDEX = {  
"number\_of\_shards": 1,  
"analysis": {  
"filter": {  
"mynGram" : {  
"type" : "nGram",  
"min\_gram": 1,  
"max\_gram": 50  
}  
},  
"analyzer": {  
"a1" : {  
"type" : "pattern",  
"pattern" : "[@/:.\w\p{L}]+",  
"filter" : ["mynGram"]  
}  
}  
}  
}

I did not understand the significance of the '\*' in your pattern. Also you  
are saying 'Letters' and 'word' and mention any number of occurrences of  
that (the end +?).

Can you please clarify?

Thanks  
-Praveen

On Thursday, July 26, 2012 1:53:57 PM UTC-7, Joe Wong wrote:

> Do you mean you want to have ":", "/" searchable?  
> if so you may have to use a custom tokenizer and specify a pattern to  
> tokenize.
> 
> This will include all unicode characters plus the special characters such  
> as "/" and ":"
> 
> 'tokenizer' : {
> 
> ```
> 'email_tokenizer' : {
> 
> 'type' : { 'pattern',
> 
> 'pattern' => "[@*\/*:*\.*\\w\\p{L}]+"
> 
> }
> 
> }
> 
> }
> 
> ```
> 
> On Thursday, July 26, 2012 2:28:46 PM UTC-4, Praveen Kariyanahalli wrote:
> 
> > I saw many threads discuss it. I have crossed few hurdles, last one is  
> > still bothering me. I had email address, ":", "/" in my data (which need to  
> > be indexed and searched). Now I am able to search the email address  
> > [test@myco.com](mailto:test@myco.com), _but I still cannot have the following characters  
> > indexed: ":" "/" and "-"_. Any help is greatly appreciated. Is it issue  
> > with my index\_analyzer or search\_analyzer?
> > 
> > Thanks in Advance  
> > -Praveen
> > 
> > Here is my mapping:
> > 
> > ESINDEX = {  
> > "number\_of\_shards": 1,  
> > "analysis": {  
> > "filter": {  
> > "mynGram" : {  
> > "type" : "nGram",  
> > "min\_gram": 1,  
> > "max\_gram": 50  
> > }  
> > },  
> > "analyzer": {  
> > "a1" : {  
> > "type" :"custom",  
> > "tokenizer":"uax\_url\_email",  
> > "filter" : ["mynGram"]  
> > }  
> > }  
> > }  
> > }
> > 
> > ESMAPPINGS = {  
> > "index\_analyzer" : "a1",  
> > "search\_analyzer" : "whitespace",  
> > "properties" : {  
> > u'test\_field1' : {  
> > 'index' : 'not\_analyzed',  
> > 'type' : u'string',  
> > 'store' : 'yes'  
> > },  
> > u'testfield2' : {  
> > 'index' : 'not\_analyzed',  
> > 'type' : u'string',  
> > 'store' : 'yes'  
> > },  
> > u'email' : {  
> > 'index': 'not\_analyzed',  
> > 'type' : u'string',  
> > 'store': 'yes'  
> > },  
> > :::::::::::  
> > }

---

<div class="post-metadata">

**Author:** ![Praveen\_Kariyanahall](https://avatars.discourse-cdn.com/v4/letter/p/4491bb/32.png) [@Praveen\_Kariyanahall](https://discuss.elastic.co/u/Praveen_Kariyanahall)\
**Post date:** [July 26, 2012, 10:16pm UTC](https://discuss.elastic.co/t/special-characters-not-indexed-and-hence-not-searchable/8530/4 "2012-07-26T22:16:50Z")

</div>

Just in case I was not clear,

I want my tokens to be anything that is made of letters, digits, @, -, .,  
:, /

Here is the pattern I am trying: "pattern" :  
"[@-.:/\p{L}\d]+"

On this token I need ngram filter.

Any help is greatly appreciated.

Thanks in Advnace  
-pk

ESINDEX = {  
"number\_of\_shards": 1,  
"analysis": {  
"filter": {  
"mynGram" : {  
"type" : "nGram",  
"min\_gram": 1,  
"max\_gram": 50  
}  
},  
"analyzer": {  
"a1" : {  
"type" : "pattern",  
"filter" : ["mynGram"],  
"pattern" : "[@-.:/\p{L}\d]+"  
}  
}  
}  
}

ESMAPPINGS = {  
"index\_analyzer" : "a1",  
"search\_analyzer" : "whitespace",  
"date\_formats" : ["yyyy-MM-dd", "MM-dd-yyyy"],  
"properties" : {  
u'my\_field' : {  
'index' : 'not\_analyzed',  
'type' : u'string',  
'store' : 'yes'  
},  
::::::::::::::  
}

---

<div class="post-metadata">

**Author:** ![Joe\_Wong](https://avatars.discourse-cdn.com/v4/letter/j/e495f1/32.png) [@Joe\_Wong](https://discuss.elastic.co/u/Joe_Wong)\
**Post date:** [July 26, 2012, 10:19pm UTC](https://discuss.elastic.co/t/special-characters-not-indexed-and-hence-not-searchable/8530/5 "2012-07-26T22:19:38Z")

</div>

ah yes '\*' aren't required

you need to define the custom tokenizer for your analyzer

i.e.

ESINDEX = {  
"number\_of\_shards": 1,  
"analysis": {  
"filter": {  
"mynGram" : {  
"type" : "nGram",  
"min\_gram": 1,  
"max\_gram": 50  
}  
},  
"analyzer": {  
"a1" : {  
"type" :"custom",  
"tokenizer":"email\_tokenizer",  
"filter" : ["mynGram"]  
}  
},  
"tokenizer" : {  
"email\_tokenizer" : {  
"type" : 'pattern',  
"pattern" =\> "[@/:.\w\p{L}]+"  
}  
}  
}  
}

On Thursday, July 26, 2012 5:57:36 PM UTC-4, Praveen Kariyanahalli wrote:

> Hi Joe
> 
> As per you suggestion, I changed my tokenizer to the following. But it  
> doesnt help. I have lost my initial ngram indexing too? Am I missing  
> something?
> 
> ESINDEX = {  
> "number\_of\_shards": 1,  
> "analysis": {  
> "filter": {  
> "mynGram" : {  
> "type" : "nGram",  
> "min\_gram": 1,  
> "max\_gram": 50  
> }  
> },  
> "analyzer": {  
> "a1" : {  
> "type" : "pattern",  
> "pattern" : "[@/:.\w\p{L}]+",  
> "filter" : ["mynGram"]  
> }  
> },

> ```
> }
> }
> 
> ```
> 
> I did not understand the significance of the '\*' in your pattern. Also you  
> are saying 'Letters' and 'word' and mention any number of occurrences of  
> that (the end +?).
> 
> Can you please clarify?
> 
> Thanks  
> -Praveen
> 
> On Thursday, July 26, 2012 1:53:57 PM UTC-7, Joe Wong wrote:
> 
> > Do you mean you want to have ":", "/" searchable?  
> > if so you may have to use a custom tokenizer and specify a pattern to  
> > tokenize.
> > 
> > This will include all unicode characters plus the special characters such  
> > as "/" and ":"
> > 
> > 'tokenizer' : {
> > 
> > ```
> > 'email_tokenizer' : {
> > 
> > 'type' : { 'pattern',
> > 
> > 'pattern' => "[@*\/*:*\.*\\w\\p{L}]+"
> > 
> > }
> > 
> > }
> > 
> > }
> > 
> > ```
> > 
> > On Thursday, July 26, 2012 2:28:46 PM UTC-4, Praveen Kariyanahalli wrote:
> > 
> > > I saw many threads discuss it. I have crossed few hurdles, last one is  
> > > still bothering me. I had email address, ":", "/" in my data (which need to  
> > > be indexed and searched). Now I am able to search the email address  
> > > [test@myco.com](mailto:test@myco.com), _but I still cannot have the following characters  
> > > indexed: ":" "/" and "-"_. Any help is greatly appreciated. Is it issue  
> > > with my index\_analyzer or search\_analyzer?
> > > 
> > > Thanks in Advance  
> > > -Praveen
> > > 
> > > Here is my mapping:
> > > 
> > > ESINDEX = {  
> > > "number\_of\_shards": 1,  
> > > "analysis": {  
> > > "filter": {  
> > > "mynGram" : {  
> > > "type" : "nGram",  
> > > "min\_gram": 1,  
> > > "max\_gram": 50  
> > > }  
> > > },  
> > > "analyzer": {  
> > > "a1" : {  
> > > "type" :"custom",  
> > > "tokenizer":"uax\_url\_email",  
> > > "filter" : ["mynGram"]  
> > > }  
> > > }  
> > > }  
> > > }
> > > 
> > > ESMAPPINGS = {  
> > > "index\_analyzer" : "a1",  
> > > "search\_analyzer" : "whitespace",  
> > > "properties" : {  
> > > u'test\_field1' : {  
> > > 'index' : 'not\_analyzed',  
> > > 'type' : u'string',  
> > > 'store' : 'yes'  
> > > },  
> > > u'testfield2' : {  
> > > 'index' : 'not\_analyzed',  
> > > 'type' : u'string',  
> > > 'store' : 'yes'  
> > > },  
> > > u'email' : {  
> > > 'index': 'not\_analyzed',  
> > > 'type' : u'string',  
> > > 'store': 'yes'  
> > > },  
> > > :::::::::::  
> > > }

---

<div class="post-metadata">

**Author:** ![Praveen\_Kariyanahall](https://avatars.discourse-cdn.com/v4/letter/p/4491bb/32.png) [@Praveen\_Kariyanahall](https://discuss.elastic.co/u/Praveen_Kariyanahall)\
**Post date:** [July 26, 2012, 11:39pm UTC](https://discuss.elastic.co/t/special-characters-not-indexed-and-hence-not-searchable/8530/6 "2012-07-26T23:39:07Z")

</div>

This finally worked (see below). I had to negate the pattern that I want  
to be tokenized. It looks like pattern is defining what my delimiter is. So  
in my case keyword pattern is saying, tokenize until you see character  
other than @, :, /, ., !, =, -, letter or digit, then go on to apply filter  
on them. I reread the documentation, then I got  
it: [http://www.elasticsearch.org/guide/reference/index-modules/analysis/pattern-analyzer.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/pattern-analyzer.html)

```
                "tokenizer" : {
                    "email_tokenizer" : {
                        "type" : "pattern",
                       * "pattern" : "[^@:\/\.\!\=\-\\w\\p{L}\\d]+"*
                    }
                },
                "analyzer": {
                    "a1" : {
                        "type" : "custom",
                        "tokenizer":"email_tokenizer",
                        "filter" : ["mynGram"]
                    }
                }
```

---

<div class="post-metadata">

**Author:** ![Anusha](https://avatars.discourse-cdn.com/v4/letter/a/c5a1d2/32.png) [@Anusha](https://discuss.elastic.co/u/Anusha)\
**Post date:** [March 20, 2015, 9:56am UTC](https://discuss.elastic.co/t/special-characters-not-indexed-and-hence-not-searchable/8530/7 "2015-03-20T09:56:01Z")

</div>

Am using the pattern "[^@:/.!=-\w\p{L}\d]+" in sense for settings. It is showing showing Bad string sytax error. Do I need to add anything for this pattern inorder to accept the string in sense.

---

<div class="post-metadata">

**Author:** ![Anusha](https://avatars.discourse-cdn.com/v4/letter/a/c5a1d2/32.png) [@Anusha](https://discuss.elastic.co/u/Anusha)\
**Post date:** [March 20, 2015, 10:00am UTC](https://discuss.elastic.co/t/special-characters-not-indexed-and-hence-not-searchable/8530/8 "2015-03-20T10:00:56Z")

</div>

Where I would like to add special characters like'-','/','(',')' to my pattern..

---

<div class="post-metadata">

**Author:** ![Anusha](https://avatars.discourse-cdn.com/v4/letter/a/c5a1d2/32.png) [@Anusha](https://discuss.elastic.co/u/Anusha)\
**Post date:** [March 20, 2015, 10:21am UTC](https://discuss.elastic.co/t/special-characters-not-indexed-and-hence-not-searchable/8530/9 "2015-03-20T10:21:45Z")

</div>

Am using the pattern "[^@:/.!=-\w\p{L}\d]+" in sense for settings.  
It is showing showing Bad string sytax error. Do I need to add anything for  
this pattern inorder to accept the string in sense.  
Where I would like to add special characters like'-','/','(',')' to my  
pattern..

Here is my settings:  
"analysis": {  
"analyzer": {  
"my\_analyzer":  
{  
"type":"custom",  
"tokenizer":"special\_tokenizer",  
"filter" : ["mynGram"]  
}  
},  
"tokenizer": {  
"special\_tokenizer":  
{  
"type" : "pattern",  
"pattern" :  
"[^-/\w\p{L}\d]+" Here am  
getting Bad String syntax error in sense , any other way of giving the  
string  
}  
},  
"filter": {  
"mynGram" : {  
"type" : "nGram",  
"min\_gram": 1,  
"max\_gram": 50  
}  
}

```
    }

```

On Thursday, July 26, 2012 at 11:58:46 PM UTC+5:30, Praveen Kariyanahalli  
wrote:

> I saw many threads discuss it. I have crossed few hurdles, last one is  
> still bothering me. I had email address, ":", "/" in my data (which need to  
> be indexed and searched). Now I am able to search the email address  
> [te...@myco.com](mailto:te...@myco.com) \<javascript:\>, _but I still cannot have the following  
> characters indexed: ":" "/" and "-"_. Any help is greatly appreciated. Is  
> it issue with my index\_analyzer or search\_analyzer?
> 
> Thanks in Advance  
> -Praveen
> 
> Here is my mapping:
> 
> ESINDEX = {  
> "number\_of\_shards": 1,  
> "analysis": {  
> "filter": {  
> "mynGram" : {  
> "type" : "nGram",  
> "min\_gram": 1,  
> "max\_gram": 50  
> }  
> },  
> "analyzer": {  
> "a1" : {  
> "type" :"custom",  
> "tokenizer":"uax\_url\_email",  
> "filter" : ["mynGram"]  
> }  
> }  
> }  
> }
> 
> ESMAPPINGS = {  
> "index\_analyzer" : "a1",  
> "search\_analyzer" : "whitespace",  
> "properties" : {  
> u'test\_field1' : {  
> 'index' : 'not\_analyzed',  
> 'type' : u'string',  
> 'store' : 'yes'  
> },  
> u'testfield2' : {  
> 'index' : 'not\_analyzed',  
> 'type' : u'string',  
> 'store' : 'yes'  
> },  
> u'email' : {  
> 'index': 'not\_analyzed',  
> 'type' : u'string',  
> 'store': 'yes'  
> },  
> :::::::::::  
> }

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/c431b584-a4a2-4b86-89ce-b9cce43d3e79%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/c431b584-a4a2-4b86-89ce-b9cce43d3e79%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:25am UTC](https://discuss.elastic.co/t/special-characters-not-indexed-and-hence-not-searchable/8530/10 "2017-07-06T00:25:23Z")

</div>


