# How to Index special characters and Search those special characters in Elasticsearch

**URL:** https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506
**Category:** Elasticsearch
**Created:** [February 23, 2016, 3:01pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506 "2016-02-23T15:01:50Z")
**Posts on this page:** 14
**Page:** 1

<div class="post-metadata">

### Author: ![Aravinthan\_Asokan](https://avatars.discourse-cdn.com/v4/letter/a/dfb087/32.png) [@Aravinthan\_Asokan](https://discuss.elastic.co/u/Aravinthan_Asokan)
#### Post date: [February 23, 2016, 3:01pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/1 "2016-02-23T15:01:50Z")

</div>

Hi

I have been trying to fix this issue for more than 20 days , but couldn't make it working.  
Also I am new to Elasticsearch as this is our first project to implement.

Step 1 :  
I have Installed Elasticsearch 2.0 in Ubuntu 14.04. I able to create new Index using below code

$hosts = array('our ip address:9200');  
$client = \Elasticsearch\ClientBuilder::create()-\>setHosts($hosts)-\>build();  
$index = "IndexName";  
$params['index'] = $index;  
$params['type'] = 'xyz';  
$params['body']["id"] = "1";  
$params['body']["title"] = "C++ Developer - C# Developer";  
$client-\>index($params);

once the above code runs Index successfully created.

Step 2 :  
Able to look into the created Index using below link

`http://our ip address:9200/IndexName/\_search?q=C%23&pretty

{  
"took" : 30,  
"timed\_out" : false,  
"\_shards" : {  
"total" : 5,  
"successful" : 5,  
"failed" : 0  
},  
"hits" : {  
"total" : 9788,  
"max\_score" : 0.8968174,  
"hits" : [ {  
"\_index" : "IndexName",  
"\_type" : "xyz",  
"\_id" : "1545680",  
"\_score" : 0.8968174,  
"\_source":{"id":"1545680","title":"C\+\+ and C\# \- Software Engineer"}  
}, {  
"\_index" : "IndexName",  
"\_type" : "xyz",  
"\_id" : "1539778",  
"\_score" : 0.853807,  
"\_source":{"id":"1539778","title":"Rebaca Technologies Hiring in C\+\+"}  
}  
....

[http://our](http://our) ip address:9200/IndexName/\_search?q=C%23&pretty

{  
"took" : 30,  
"timed\_out" : false,  
"\_shards" : {  
"total" : 5,  
"successful" : 5,  
"failed" : 0  
},  
"hits" : {  
"total" : 9788,  
"max\_score" : 0.8968174,  
"hits" : [ {  
"\_index" : "IndexName",  
"\_type" : "xyz",  
"\_id" : "1545680",  
"\_score" : 0.8968174,  
"\_source":{"id":"1545680","title":"C\+\+ and C\# \- Software Engineer"}  
}, {  
"\_index" : "IndexName",  
"\_type" : "xyz",  
"\_id" : "1539778",  
"\_score" : 0.853807,  
"\_source":{"id":"1539778","title":"Rebaca Technologies Hiring in C\+\+"}  
}  
....

If you note the above search result i am getting 2nd result which is not having c#. Even i am getting the same result for search "C" only

I am not getting relavant search result according to the keywords which contains special characters like +, #, or .

I am preserving the special characters as per the below guide

Escaping Special Characters

Lucene supports escaping special characters that are part of the query syntax. The current list special characters are

`+ - && || ! ( ) { } [] ^ " ~ * ? : \`  
`  
To escape these character use the \ before the character. For example to search for (1+1):2 use the query:

`(1+1):2

`I added # in the group of escape charaters.

Step 3:

In php while passing the special characters into Elasticsearch search function i am escaping like below

$keyword = str\_replace(""",'"',$keyword);  
$keyword = str\_replace("+","+",$keyword);  
$keyword = str\_replace(".",".",$keyword);  
$keyword = str\_replace("#","#",$keyword);  
$keyword = str\_replace("/","/",$keyword);  
$keyword = trim($keyword);

$params['body']['query']['query\_string'] = array("query" =\> $keyword,"default\_operator" =\> "AND" ,"fields" =\> array("title"));  
$client-\>search($params); `

Please help me how to make the special character work

Thanks

---

<div class="post-metadata">

### Author: ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)
#### Post date: [March 1, 2016, 2:14pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/2 "2016-03-01T14:14:47Z")

</div>

Hi,

you're looking in the wrong spot. Your problem is related to something called analysis. I suggest you read more about [analysis in the definitive guide](https://www.elastic.co/guide/en/elasticsearch/guide/current/analysis-intro.html).

By default, Elasticsearch uses the "standard" analyzer to analyze text. You can try this by yourself in Sense:

```auto
GET /_analyze?analyzer=standard
{
    "text": "C# developer"
}

```

This produces:

```auto
{
   "tokens": [
      {
         "token": "c",
         "start_offset": 6,
         "end_offset": 7,
         "type": "<ALPHANUM>",
         "position": 1
      },
      {
         "token": "developer",
         "start_offset": 9,
         "end_offset": 18,
         "type": "<ALPHANUM>",
         "position": 2
      }
   ]
}

```

You can see that Elasticsearch's standard analyzer just strips the "#" character (and similarly "++"). The analyzer is applied at index time so your text never makes it into the index as you want it.

Hence, one solution to this problem is to define your own analyzer. Here is a minimal example that should get you going:

First, we create a custom analyzer. We use the whitespace tokenizer here but you should check the [documentation on custom analyzers](https://www.elastic.co/guide/en/elasticsearch/guide/current/custom-analyzers.html) and decide whether this really fits your use case.

```auto
PUT /my_index
{
   "settings": {
      "analysis": {
         "analyzer": {
            "my_analyzer": {
               "type": "custom",
               "filter": [
                  "lowercase"
               ],
               "tokenizer": "whitespace"
            }
         }
      }
   }
}

```

We can already try that now the special characters are preserved:

```auto
GET /my_index/_analyze?analyzer=my_analyzer
{
    "text": "C# developer"
}

```

This produces:

```auto
{
   "tokens": [
      {
         "token": "c#",
         "start_offset": 0,
         "end_offset": 2,
         "type": "word",
         "position": 0
      },
      {
         "token": "developer",
         "start_offset": 3,
         "end_offset": 12,
         "type": "word",
         "position": 1
      }
   ]
}

```

Note that "c#" is still present as a token. This is key to understand the rest.

Now we have to use our custom analyzer. For that we define a new type called "jobs":

```auto
PUT /my_index/_mapping/jobs
{
   "properties": {
      "content": {
         "type": "string",
         "analyzer": "my_analyzer"
      }
   }
}

```

We can now index some documents:

```auto
POST /_bulk
{"index":{"_index":"my_index","_type":"jobs"}}
{"content":"We are looking for C++ and C# developers"}
{"index":{"_index":"my_index","_type":"jobs"}}
{"content":"We are looking for C developers"}
{"index":{"_index":"my_index","_type":"jobs"}}
{"content":"We are looking for project managers"}

```

And if we search now for "C#":

```auto
GET /my_index/jobs/_search
{
   "query": {
      "match": {
         "content": {
            "query": "C#"
         }
      }
   }
}

```

we get the expected result:

```auto
{
   "took": 3,
   "timed_out": false,
   "_shards": {
      "total": 5,
      "successful": 5,
      "failed": 0
   },
   "hits": {
      "total": 1,
      "max_score": 0.095891505,
      "hits": [
         {
            "_index": "my_index",
            "_type": "jobs",
            "_id": "AVMyfdxBfIbbKEiejUJ3",
            "_score": 0.095891505,
            "_source": {
               "content": "We are looking for C++ and C# developers"
            }
         }
      ]
   }
}

```

I can heartily recommend the [Definitive Guide](https://www.elastic.co/guide/en/elasticsearch/guide/current/index.html) to get a deeper understanding of Elasticsearch.

Daniel

---

<div class="post-metadata">

### Author: ![Rudolf\_Reddy\_Macejka](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rudolf_reddy_macejka/32/12875_2.png) [@Rudolf\_Reddy\_Macejka](https://discuss.elastic.co/u/Rudolf_Reddy_Macejka)
#### Post date: [November 1, 2016, 2:50pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/3 "2016-11-01T14:50:09Z")

</div>

Hi Daniel,

I like your solution.  
However how can i apply the customized analyzer and the job for new indexes dynamically created every day?

I am using logstash for that.

Thanks, Reddy

---

<div class="post-metadata">

### Author: ![cbuescher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cbuescher/32/60402_2.png) [@cbuescher](https://discuss.elastic.co/u/cbuescher)
#### Post date: [November 1, 2016, 2:58pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/4 "2016-11-01T14:58:32Z")

</div>

Hi @Rudolf_Reddy_Macejka,

you can use [Index Templates](https://www.elastic.co/guide/en/elasticsearch/reference/5.0/indices-templates.html) for doing just that.

---

<div class="post-metadata">

### Author: ![Rudolf\_Reddy\_Macejka](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rudolf_reddy_macejka/32/12875_2.png) [@Rudolf\_Reddy\_Macejka](https://discuss.elastic.co/u/Rudolf_Reddy_Macejka)
#### Post date: [November 1, 2016, 3:21pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/6 "2016-11-01T15:21:43Z")

</div>

Thank you @cbuescher,

created

> PUT \_template/all\_whitespace\_only  
> {  
> "template": "\*",  
> "version": 1,  
> "settings": {  
> "number\_of\_shards": 1,  
> "analysis": {  
> "analyzer": {  
> "anal\_whitespace\_only": {  
> "type": "custom",  
> "filter": [  
> "lowercase"  
> ],  
> "tokenizer": "whitespace"  
> }  
> }  
> }  
> },  
> "mappings": {  
> "jobs": {  
> "properties": {  
> "content": {  
> "type": "string",  
> "analyzer": "anal\_whitespace\_only"  
> }  
> }  
> }  
> }  
> }

Will provide the result then,

Rudo

---

<div class="post-metadata">

### Author: ![Rudolf\_Reddy\_Macejka](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rudolf_reddy_macejka/32/12875_2.png) [@Rudolf\_Reddy\_Macejka](https://discuss.elastic.co/u/Rudolf_Reddy_Macejka)
#### Post date: [November 21, 2016, 12:58pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/7 "2016-11-21T12:58:53Z")

</div>

Hello,

finally I can continue. the above template is applicable to the field content.

But how I can create it for all string fields generaly?

I have found a solution and after amending it seems by following:

```
PUT _template/all_whitespace_only
{
  "template": "tracking*",
  "version": 1,
  "settings": {
    "number_of_shards": 1,
    "analysis": {
      "analyzer": {
        "default": {
           "type": "custom",
             "filter": [
                "lowercase"
             ],
              "tokenizer": "whitespace"
            }
         }
      }
  },
  "mappings": {
    "_default_": {
      "dynamic_templates": {
        "string_fields": {
          "match": "*",
          "match_mapping_type": "string",
          "mapping": {
            "type": "string",
            "index": "analyzed",
            "analyzer": "anal_whitespace_only",
            "fielddata": {
                "format": "disabled"
            }
          }
        }
      }
    }
  }
}

```

However I am getting error

`Failed to parse mapping [_default_]: java.util.LinkedHashMap cannot be cast to java.util.List", "caused_by"=>{"type"=>"class_cast_exception", "reason"=>"java.util.LinkedHashMap cannot be cast to java.util.List"}}}}, :level=>:warn}←[0m`

Be honest I do not understand what is wrong.

Can you please help?

Thank you!  
Reddy

---

<div class="post-metadata">

### Author: ![cbuescher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cbuescher/32/60402_2.png) [@cbuescher](https://discuss.elastic.co/u/cbuescher)
#### Post date: [November 21, 2016, 1:04pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/8 "2016-11-21T13:04:50Z")

</div>

Which client are you using? The error looks like you're not putting the template using curl or the sense plugin, but I might be mistaken. I suspect there's an error using the client.

---

<div class="post-metadata">

### Author: ![Rudolf\_Reddy\_Macejka](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rudolf_reddy_macejka/32/12875_2.png) [@Rudolf\_Reddy\_Macejka](https://discuss.elastic.co/u/Rudolf_Reddy_Macejka)
#### Post date: [November 21, 2016, 1:09pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/9 "2016-11-21T13:09:06Z")

</div>

It is error from logstash. Here is my configuration:

```
input { 
  jdbc { jdbc_driver_library => "e:\Utility\elasticsearch-2.1.0\addons\logstash-2.1.0\lib\ojdbc6.jar"
                     jdbc_driver_class => "Java::oracle.jdbc.driver.OracleDriver"
           jdbc_connection_string => "jdbc:oracle:thin:@ *******:1521/****"
           jdbc_user => " *****"
           jdbc_password => " *******"
                     parameters => { }
                     schedule => "08 * * * *"
                     statement => "select eg.identifier id, al.DATA_CORRELATIONID IdSourceMessage, eg.SNRF IdMasterMessage, 'EMEA DCG MB UAT' sourceSystem, 'EMEA' as region, 'DCG MB UAT' as platform, to_char(eg.datetime, 'YYYY-MM-DD\"T\"HH24:MI') timestamp, eg.sender Sender, eg.receiver Receiver, eg.aprf APRF, eg.snrf SNRF, eg.doctrackid trackingID, decode(substr(lower(eg.Status),1,11),'translation',trim(REGEXP_REPLACE(eg.Status,'[[:alpha:]]'))) SessionID, eg.status Status, eg.text Action, eg.details Details from gxsmailbox_uat.tgeg_log eg left join dcgplatform_uat.gl_utils_audit_log al on eg.doctrackid = substr(al.data_message,instr(lower(al.data_message),'indentifier')+13,18) where eg.class = 'DATA' and eg.datetime > sysdate-1.1/24 order by eg.identifier"
      } 
}

filter { 
    mutate { 
        gsub => [
            "timestamp","[\\]", ""
         ] 
    }
}

output {
  elasticsearch { "index" => "tracking-%{+YYYY.MM.dd}"
                                      "document_type" => "eg"
                                    "document_id" => "eg-uat-%{id}" 
    }

}
```

---

<div class="post-metadata">

### Author: ![cbuescher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cbuescher/32/60402_2.png) [@cbuescher](https://discuss.elastic.co/u/cbuescher)
#### Post date: [November 21, 2016, 2:28pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/10 "2016-11-21T14:28:11Z")

</div>

According to the [documentation](https://www.elastic.co/guide/en/elasticsearch/reference/5.0/dynamic-templates.html), `dynamic_templates` needs to be an array. Does this solve your problem?

---

<div class="post-metadata">

### Author: ![Rudolf\_Reddy\_Macejka](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rudolf_reddy_macejka/32/12875_2.png) [@Rudolf\_Reddy\_Macejka](https://discuss.elastic.co/u/Rudolf_Reddy_Macejka)
#### Post date: [November 21, 2016, 2:39pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/11 "2016-11-21T14:39:20Z")

</div>

Will check .,.. and let you know

---

<div class="post-metadata">

### Author: ![Rudolf\_Reddy\_Macejka](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rudolf_reddy_macejka/32/12875_2.png) [@Rudolf\_Reddy\_Macejka](https://discuss.elastic.co/u/Rudolf_Reddy_Macejka)
#### Post date: [November 22, 2016, 8:28am UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/12 "2016-11-22T08:28:44Z")

</div>

Hi,

i have removed all mappings from the template and remain only the analyzer part and it works as I wanted ... I tried to too much combine things and thinks 🙂

thank you.

---

<div class="post-metadata">

### Author: ![mateuspadua](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mateuspadua/32/16946_2.png) [@mateuspadua](https://discuss.elastic.co/u/mateuspadua)
#### Post date: [March 31, 2017, 11:47am UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/13 "2017-03-31T11:47:23Z")

</div>

Tip: to scape all characters in python, do:

```
>>> q = "your_search+here-now&"
>>> q = re.sub(pattern=r'([+\-=&|><(){}\[\]\^"~*?:\/])', repl=r'\\\1', string=q)
>>> print q 
your_search\+here\-now\&
```

---

<div class="post-metadata">

### Author: ![Shellbye\_Bai](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shellbye_bai/32/29850_2.png) [@Shellbye\_Bai](https://discuss.elastic.co/u/Shellbye_Bai)
#### Post date: [April 5, 2017, 12:20pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/14 "2017-04-05T12:20:50Z")

</div>

That's really helpful.Thanks

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 5, 2017, 10:01pm UTC](https://discuss.elastic.co/t/how-to-index-special-characters-and-search-those-special-characters-in-elasticsearch/42506/15 "2017-07-05T22:01:29Z")

</div>


