# Indexing and searching on special characters?

**URL:** <https://discuss.elastic.co/t/indexing-and-searching-on-special-characters/9312>\
**Category:** Elasticsearch\
**Created:** [October 11, 2012, 7:36am UTC](https://discuss.elastic.co/t/indexing-and-searching-on-special-characters/9312 "2012-10-11T07:36:46Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![tsandstr](https://avatars.discourse-cdn.com/v4/letter/t/94ad74/32.png) [@tsandstr](https://discuss.elastic.co/u/tsandstr)\
**Post date:** [October 11, 2012, 7:36am UTC](https://discuss.elastic.co/t/indexing-and-searching-on-special-characters/9312/1 "2012-10-11T07:36:46Z")

</div>

I am having some troubles with how the indexing works on strings. I use normal mapping for strings and query\_string for my queries.

I index the following:

curl -XPUT [http://localhost:9200/twitter/tweet/1](http://localhost:9200/twitter/tweet/1) -d '{  
"message": "A&B: This is just a test"  
}'  
curl -XPUT [http://localhost:9200/twitter/tweet/2](http://localhost:9200/twitter/tweet/2) -d '{  
"message": "A+B: This is just a test"  
}'  
curl -XPUT [http://localhost:9200/twitter/tweet/3](http://localhost:9200/twitter/tweet/3) -d '{  
"message": "A-B: This is just a test"  
}'  
curl -XPUT [http://localhost:9200/twitter/tweet/4](http://localhost:9200/twitter/tweet/4) -d '{  
"message": "A_B: This is just a test"  
}'  
curl -XPUT [http://localhost:9200/twitter/tweet/5](http://localhost:9200/twitter/tweet/5) -d '{  
"message": "A&C: This is just a test"  
}'  
curl -XPUT [http://localhost:9200/twitter/tweet/6](http://localhost:9200/twitter/tweet/6) -d '{  
"message": "A+C: This is just a test"  
}'  
curl -XPUT [http://localhost:9200/twitter/tweet/7](http://localhost:9200/twitter/tweet/7) -d '{  
"message": "A-C: This is just a test"  
}'  
curl -XPUT [http://localhost:9200/twitter/tweet/8](http://localhost:9200/twitter/tweet/8) -d '{  
"message": "A_C: This is just a test"  
}'

and I get the following results with my queries:

{'query': {'query\_string': {'query': 'A+B', 'default\_operator': 'AND', 'default\_field': 'message'}}} {u'message': u'A&B: This is just a test'} {u'message': u'A+B: This is just a test'} {u'message': u'A-B: This is just a test'} {u'message': u'A\*B: This is just a test'} --\> I Expected to get only: A+B: This is just a test

{'query': {'query\_string': {'query': 'A&B', 'default\_operator': 'AND', 'default\_field': 'message'}}}  
{u'message': u'A&B: This is just a test'}  
{u'message': u'A+B: This is just a test'}  
{u'message': u'A-B: This is just a test'}  
{u'message': u'A\*B: This is just a test'}  
--\> I Expected to get only: A&B: This is just a test

{'query': {'query\_string': {'query': 'A\*B', 'default\_operator': 'AND', 'default\_field': 'message'}}}  
--\> Got no results at all, this is VERY confusing

{'query': {'query\_string': {'query': 'A-B', 'default\_operator': 'AND', 'default\_field': 'message'}}}  
{u'message': u'A&B: This is just a test'}  
{u'message': u'A+B: This is just a test'}  
{u'message': u'A-B: This is just a test'}  
{u'message': u'A\*B: This is just a test'}  
--\> I Expected to get only: A&B: This is just a test

{'query': {'query\_string': {'query': 'A B', 'default\_operator': 'AND', 'default\_field': 'message'}}}  
{u'message': u'A&B: This is just a test'}  
{u'message': u'A+B: This is just a test'}  
{u'message': u'A-B: This is just a test'}  
{u'message': u'A\*B: This is just a test'}

Why does the query A\*B not return anything at all?  
Is there a simple way to know how special characters are indexed by default?  
It seems to me that A&B, A+B and A-B are all indexed the same way as A B, am I correct?

What about all other characters? How do they behave and how should I construct my query in order to get correct results? Escaping them does not seem to help.

Any help would be appreciated.

Thank you!

- Tommy

---

<div class="post-metadata">

**Author:** ![tsandstr](https://avatars.discourse-cdn.com/v4/letter/t/94ad74/32.png) [@tsandstr](https://discuss.elastic.co/u/tsandstr)\
**Post date:** [October 11, 2012, 11:31am UTC](https://discuss.elastic.co/t/indexing-and-searching-on-special-characters/9312/2 "2012-10-11T11:31:43Z")

</div>

Nevermind, it seems that I should escape all extra characters. And since the standard analyzer either replaces the special character with spaces or does something else. I will end up with several results no mather which special character is used.

This is okay for me, since I do not really care about the extra characters and do not need a search on them.

In some cases though it could be good to do the indexing of the extra characters too. But I guess the only way to do that is to create an own analyzer and use it?

One more question. How are endings such as 's handled? For example: Twitter's

I would rather remove the 's endings, instead of having them, during index. Can this behaviour be easily achieved?

- Tommy

---

<div class="post-metadata">

**Author:** ![Chris\_Male](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chris_male/32/2607_2.png) [@Chris\_Male](https://discuss.elastic.co/u/Chris_Male)\
**Post date:** [October 13, 2012, 7:14am UTC](https://discuss.elastic.co/t/indexing-and-searching-on-special-characters/9312/3 "2012-10-13T07:14:01Z")

</div>

Tommy,

On Friday, October 12, 2012 12:31:43 AM UTC+13, tommy wrote:

> Nevermind, it seems that I should escape all extra characters. And since  
> the  
> standard analyzer either replaces the special character with spaces or  
> does  
> something else. I will end up with several results no mather which special  
> character is used.
> 
> This is okay for me, since I do not really care about the extra characters  
> and do not need a search on them.
> 
> In some cases though it could be good to do the indexing of the extra  
> characters too. But I guess the only way to do that is to create an own  
> analyzer and use it?

StandardAnalyzer is really just built of a Tokenizer and a number of  
TokenFilters. To have different behaviour (such as indexing extra  
characters) you can just define a new Analyzer your ES mapping which uses a  
different Tokenizer or TokenFilters.

> One more question. How are endings such as 's handled? For example:  
> Twitter's
> 
> I would rather remove the 's endings, instead of having them, during  
> index.  
> Can this behaviour be easily achieved?
> 
> - Tommy
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Indexing-and-searching-on-special-characters-tp4023761p4023781.html](http://elasticsearch-users.115913.n3.nabble.com/Indexing-and-searching-on-special-characters-tp4023761p4023781.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:08am UTC](https://discuss.elastic.co/t/indexing-and-searching-on-special-characters/9312/4 "2017-07-06T03:08:57Z")

</div>


