# UTF8 encoding is longer than the max length 32766

**URL:** https://discuss.elastic.co/t/utf8-encoding-is-longer-than-the-max-length-32766/816
**Category:** Elasticsearch
**Created:** [May 18, 2015, 2:20am UTC](https://discuss.elastic.co/t/utf8-encoding-is-longer-than-the-max-length-32766/816 "2015-05-18T02:20:00Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![avin](https://avatars.discourse-cdn.com/v4/letter/a/ee7513/32.png) [@avin](https://discuss.elastic.co/u/avin)
#### Post date: [May 18, 2015, 2:20am UTC](https://discuss.elastic.co/t/utf8-encoding-is-longer-than-the-max-length-32766/816/1 "2015-05-18T02:20:00Z")

</div>

I have a requirement to store a text larger than 64K. I dont want to index it, but still while inserting into the index, I see the following exception.

IllegalArgumentException[Document contains at least one immense term in field="kvdatav1" (whose UTF8 encoding is longer than the max length 32766), all of which were skipped. Please correct the analyzer to not produce such terms. The prefix of the first immense term is: '[86, 78, 49, 66, 55, 81, 86, 115, 74, 100, 85, 120, 98, 112, 73, 122, 67, 78, 77, 71, 116, 81, 107, 53, 65, 79, 90, 120, 119, 120]...', original message: bytes can be at most 32766 in length; got 45000]; nested: MaxBytesLengthExceededException[bytes can be at most 32766 in length; got 45000];

How do I work around the problem. I am using ES version 1.4.2

curl -XDELETE '[http://localhost:9200/test](http://localhost:9200/test)'  
curl -XPUT '[http://localhost:9200/test/](http://localhost:9200/test/)' -d '  
{  
"settings": {  
"number\_of\_shards": 5, //Default for number\_of\_shards is 5  
"number\_of\_replicas": 1, //Default for number\_of\_replicas is 1  
"analysis" : {  
"analyzer" : {  
"default" : {  
"type" : "keyword"  
}  
}  
}  
},  
"mappings": {  
"data" : {  
"properties" : {  
"kvdatav1" : {"type" : "string", "index" : "no"},"kReq" : {"type" : "string", "index" : "no"},"kResp" : {"type" : "string", "index" : "no"}  
}  
}  
}  
}'

---

<div class="post-metadata">

### Author: ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)
#### Post date: [May 18, 2015, 4:51am UTC](https://discuss.elastic.co/t/utf8-encoding-is-longer-than-the-max-length-32766/816/2 "2015-05-18T04:51:05Z")

</div>

You can try `ignore_above: 256`

From [the docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/mapping-core-types.html#mapping-core-types);

> ignore\_above  
> The analyzer will ignore strings larger than this size. Useful for generic not\_analyzed fields that should ignore long text.

---

<div class="post-metadata">

### Author: ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)
#### Post date: [May 18, 2015, 1:21pm UTC](https://discuss.elastic.co/t/utf8-encoding-is-longer-than-the-max-length-32766/816/3 "2015-05-18T13:21:03Z")

</div>

It looks like you just want the field in the \_source and not searchable. If  
that's true then you can just set "index": "no" and it won't be searchable  
but will still be return-able in the \_source.

---

<div class="post-metadata">

### Author: ![avin](https://avatars.discourse-cdn.com/v4/letter/a/ee7513/32.png) [@avin](https://discuss.elastic.co/u/avin)
#### Post date: [May 18, 2015, 10:05pm UTC](https://discuss.elastic.co/t/utf8-encoding-is-longer-than-the-max-length-32766/816/4 "2015-05-18T22:05:07Z")

</div>

Tried that as well. However it doesn't seem to be working fine.

```
curl -XDELETE 'http://localhost:9200/test'

curl -XPUT 'http://localhost:9200/test/' -d '
{
   "settings": {
       "number_of_shards": 5, //Default for number_of_shards is 5
       "number_of_replicas": 1, //Default for number_of_replicas is 1
       "analysis" : {
           "analyzer" : {
               "default" : {
                   "type" : "keyword",
                   "ignore_above" : 256
               }
           }
       }
   },

   "mappings": {
       "data" : {
           "properties" : {
               "kvdatav1" : {"type" : "string", "ignore_above" : 256, "index" : "no"},"kReq" : {"type" : "string", "ignore_above" : 256, "index" : "no"},"kResp" : {"type" : "string", "ignore_above" : 256, "index" : "no"}
           }
       }
   }
}'
```

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 12:13am UTC](https://discuss.elastic.co/t/utf8-encoding-is-longer-than-the-max-length-32766/816/5 "2017-07-06T00:13:24Z")

</div>


