# How can I specify the UTF-8 codec when searching ES from python?

**URL:** <https://discuss.elastic.co/t/how-can-i-specify-the-utf-8-codec-when-searching-es-from-python/23088>\
**Category:** Elasticsearch\
**Created:** [April 3, 2015, 6:13pm UTC](https://discuss.elastic.co/t/how-can-i-specify-the-utf-8-codec-when-searching-es-from-python/23088 "2015-04-03T18:13:01Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Peter\_Trei](https://avatars.discourse-cdn.com/v4/letter/p/9f8e36/32.png) [@Peter\_Trei](https://discuss.elastic.co/u/Peter_Trei)\
**Post date:** [April 3, 2015, 6:13pm UTC](https://discuss.elastic.co/t/how-can-i-specify-the-utf-8-codec-when-searching-es-from-python/23088/1 "2015-04-03T18:13:01Z")

</div>

This issue is probably due to my noobishness to ELK, Python, and Unicode.

I have an index containing logstash-digested logs, including a field  
'host\_req', which contains a host name.  
Using Elasticsearch-py, I'm pulling that host name out of the record, and  
using it to search in another index.  
However, if the hostname contains multibyte characters, it fails with a  
UnicodeDecodeError  
Exactly the same query works fine when I enter it from the command line  
with 'curl -XGET'  
The unicode character is a lowercase 'a' with a diaeresis (two dots). The  
UTF-8 value is C3 A4,  
and the unicode code point seems to be 00E4 (the language is Swedish).

These curl commands work just fine from the command line:

curl -XGET  
'[http://localhost:9200/logstash-2015.01.30/logs/\_search?pretty=1](http://localhost:9200/logstash-2015.01.30/logs/_search?pretty=1)' -d ' {  
"query" : {"match" :{"req\_host" : "www.utkl\u00E4dningskl\u00E4derna.se"  
}}}'  
curl -XGET  
'[http://localhost:9200/logstash-2015.01.30/logs/\_search?pretty=1](http://localhost:9200/logstash-2015.01.30/logs/_search?pretty=1)' -d ' {  
"query" : {"match" :{"req\_host" : "www.utklädningskläderna.se" }}}'

They find and return the record

(the second line shows how the hostname appears in the log I pull it from,  
showing the lowercase 'a' with a diaersis, in two places)

I've written a very short Python script to show the problem: It uses  
hardwired queries, printing them and their type, then trying to use them  
in a search.  
---- start code ----

#!/usr/bin/python

# -_- coding: utf-8 -_-

import json  
import elasticsearch

es = elasticsearch.Elasticsearch()

if **name** ==" **main**":  
#uq = u'{ "query": { "match": { "req\_host": "www.utklädningskläderna.se"  
}}}' # raw utf-8 characters. does not work  
#uq = u'{ "query": { "match": { "req\_host":  
"www.utkl\u00E4dningskl\u00E4derna.se" }}}' # quoted unicode characters.  
does not work  
#uq = u'{ "query": { "match": { "req\_host":  
"www.utkl\uC3A4dningskl\uC3A4derna.se" }}}' # quoted uft-8 characters. does  
not work  
uq = u'{ "query": { "match": { "req\_host": "[www.facebook.com](http://www.facebook.com)"  
}}}' # non-unicode. works fine  
print "uq", type(uq), uq  
result =  
es.search(index="logstash-2015.01.30",doc\_type="logs",timeout=1000,body=uq);  
if result["hits"]["total"] == 0:  
print "nothing found"  
else:  
print "found some"

--- end code ----

If I run it as shown, with the 'facebook' query, it's fine - the output is:

$python testutf8b.py  
uq \<type 'unicode'\> { "query": { "match": { "req\_host": "[www.facebook.com](http://www.facebook.com)"  
}}}  
found some  
$

Note that the query string 'uq' is unicode.

But if I use the other three strings, which include the Unicode  
characters, I get

python testutf8b.py  
uq \<type 'unicode'\> { "query": { "match": { "req\_host":  
"www.utklädningskläderna.se" }}}  
Traceback (most recent call last):  
File "testutf8b.py", line 15, in   
result =  
es.search(index="logstash-2015.01.30",doc\_type="logs",timeout=1000,body=uq);  
File "build/bdist.linux-x86\_64/egg/elasticsearch/client/utils.py", line  
68, in \_wrapped  
File "build/bdist.linux-x86\_64/egg/elasticsearch/client/ **init**.py",  
line 497, in search  
File "build/bdist.linux-x86\_64/egg/elasticsearch/transport.py", line 307,  
in perform\_request  
File  
"build/bdist.linux-x86\_64/egg/elasticsearch/connection/http\_urllib3.py",  
line 82, in perform\_request  
elasticsearch.exceptions.ConnectionError: ConnectionError('ascii' codec  
can't decode byte 0xc3 in position 45: ordinal not in range(128)) caused  
by: UnicodeDecodeError('ascii' codec can't decode byte 0xc3 in position 45:  
ordinal not in range(128))

This is under Centos 7, using ES 1.5.0. The logs were digested into ES  
under a slightly older version, using logstasth-1.4.2

Any ideas? ES documentation contains sections about codecs, but that's for  
analysis. This looks to me like an  
elasticsearch-py library issue, (or I'm doing something stupid).

thanks!

PT

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/505b8b8f-faeb-4954-8fcb-e3107c2c80b6%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/505b8b8f-faeb-4954-8fcb-e3107c2c80b6%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![honzakral](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/honzakral/32/44958_2.png) [@honzakral](https://discuss.elastic.co/u/honzakral)\
**Post date:** [April 3, 2015, 9:24pm UTC](https://discuss.elastic.co/t/how-can-i-specify-the-utf-8-codec-when-searching-es-from-python/23088/2 "2015-04-03T21:24:29Z")

</div>

Hi Peter,

the easiest way is to just use python dictionaries - not strings containing  
json. Then everything will be correctly encoded and decoded into utf-8 and  
you can just work with unicode strings in your app.

Hope this helps

On Fri, Apr 3, 2015 at 8:13 PM, Peter Trei [petertrei@gmail.com](mailto:petertrei@gmail.com) wrote:

> This issue is probably due to my noobishness to ELK, Python, and Unicode.
> 
> I have an index containing logstash-digested logs, including a field  
> 'host\_req', which contains a host name.  
> Using Elasticsearch-py, I'm pulling that host name out of the record, and  
> using it to search in another index.  
> However, if the hostname contains multibyte characters, it fails with a  
> UnicodeDecodeError  
> Exactly the same query works fine when I enter it from the command line  
> with 'curl -XGET'  
> The unicode character is a lowercase 'a' with a diaeresis (two dots). The  
> UTF-8 value is C3 A4,  
> and the unicode code point seems to be 00E4 (the language is Swedish).
> 
> These curl commands work just fine from the command line:
> 
> curl -XGET '  
> [http://localhost:9200/logstash-2015.01.30/logs/\_search?pretty=1](http://localhost:9200/logstash-2015.01.30/logs/_search?pretty=1)' -d ' {  
> "query" : {"match" :{"req\_host" : "www.utkl\u00E4dningskl\u00E4derna.se"  
> }}}'  
> curl -XGET '  
> [http://localhost:9200/logstash-2015.01.30/logs/\_search?pretty=1](http://localhost:9200/logstash-2015.01.30/logs/_search?pretty=1)' -d ' {  
> "query" : {"match" :{"req\_host" : "www.utklädningskläderna.se  
> [http://www.utklädningskläderna.se](http://www.xn--utkldningsklderna-tqbi.se)" }}}'
> 
> They find and return the record
> 
> (the second line shows how the hostname appears in the log I pull it from,  
> showing the lowercase 'a' with a diaersis, in two places)
> 
> I've written a very short Python script to show the problem: It uses  
> hardwired queries, printing them and their type, then trying to use them  
> in a search.  
> ---- start code ----
> 
> #!/usr/bin/python
> 
> # -_- coding: utf-8 -_-
> 
> import json  
> import elasticsearch
> 
> es = elasticsearch.Elasticsearch()
> 
> if **name** ==" **main**":  
> #uq = u'{ "query": { "match": { "req\_host": "www.utklädningskläderna.se  
> [http://www.utklädningskläderna.se](http://www.xn--utkldningsklderna-tqbi.se)" }}}' # raw  
> utf-8 characters. does not work  
> #uq = u'{ "query": { "match": { "req\_host":  
> "www.utkl\u00E4dningskl\u00E4derna.se" }}}' # quoted unicode characters.  
> does not work  
> #uq = u'{ "query": { "match": { "req\_host":  
> "www.utkl\uC3A4dningskl\uC3A4derna.se" }}}' # quoted uft-8 characters. does  
> not work  
> uq = u'{ "query": { "match": { "req\_host": "[www.facebook.com](http://www.facebook.com)"  
> }}}' # non-unicode. works fine  
> print "uq", type(uq), uq  
> result =  
> es.search(index="logstash-2015.01.30",doc\_type="logs",timeout=1000,body=uq);  
> if result["hits"]["total"] == 0:  
> print "nothing found"  
> else:  
> print "found some"
> 
> --- end code ----
> 
> If I run it as shown, with the 'facebook' query, it's fine - the output is:
> 
> $python testutf8b.py  
> uq \<type 'unicode'\> { "query": { "match": { "req\_host": "[www.facebook.com](http://www.facebook.com)"  
> }}}  
> found some  
> $
> 
> Note that the query string 'uq' is unicode.
> 
> But if I use the other three strings, which include the Unicode  
> characters, I get
> 
> python testutf8b.py  
> uq \<type 'unicode'\> { "query": { "match": { "req\_host": "  
> www.utklädningskläderna.se [http://www.utklädningskläderna.se](http://www.xn--utkldningsklderna-tqbi.se)" }}}  
> Traceback (most recent call last):  
> File "testutf8b.py", line 15, in   
> result =  
> es.search(index="logstash-2015.01.30",doc\_type="logs",timeout=1000,body=uq);  
> File "build/bdist.linux-x86\_64/egg/elasticsearch/client/utils.py", line  
> 68, in \_wrapped  
> File "build/bdist.linux-x86\_64/egg/elasticsearch/client/ **init**.py",  
> line 497, in search  
> File "build/bdist.linux-x86\_64/egg/elasticsearch/transport.py", line  
> 307, in perform\_request  
> File  
> "build/bdist.linux-x86\_64/egg/elasticsearch/connection/http\_urllib3.py",  
> line 82, in perform\_request  
> elasticsearch.exceptions.ConnectionError: ConnectionError('ascii' codec  
> can't decode byte 0xc3 in position 45: ordinal not in range(128)) caused  
> by: UnicodeDecodeError('ascii' codec can't decode byte 0xc3 in position 45:  
> ordinal not in range(128))
> 
> This is under Centos 7, using ES 1.5.0. The logs were digested into ES  
> under a slightly older version, using logstasth-1.4.2
> 
> Any ideas? ES documentation contains sections about codecs, but that's for  
> analysis. This looks to me like an  
> elasticsearch-py library issue, (or I'm doing something stupid).
> 
> thanks!
> 
> PT
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/505b8b8f-faeb-4954-8fcb-e3107c2c80b6%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/505b8b8f-faeb-4954-8fcb-e3107c2c80b6%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/505b8b8f-faeb-4954-8fcb-e3107c2c80b6%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/505b8b8f-faeb-4954-8fcb-e3107c2c80b6%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
Honza Král  
Python Engineer  
[honza.kral@elastic.co](mailto:honza.kral@elastic.co)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAC4Vrtx7CVJ%3DsRwqyj572W7vy9\_Wvtk37wYa1WDyY%2BegUF%2Bssg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAC4Vrtx7CVJ%3DsRwqyj572W7vy9_Wvtk37wYa1WDyY%2BegUF%2Bssg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:21am UTC](https://discuss.elastic.co/t/how-can-i-specify-the-utf-8-codec-when-searching-es-from-python/23088/3 "2017-07-06T00:21:44Z")

</div>


