# Encoding is longer than the max length 32766

**URL:** https://discuss.elastic.co/t/encoding-is-longer-than-the-max-length-32766/17819
**Category:** Elasticsearch
**Created:** [May 29, 2014, 8:47pm UTC](https://discuss.elastic.co/t/encoding-is-longer-than-the-max-length-32766/17819 "2014-05-29T20:47:37Z")
**Posts on this page:** 7
**Page:** 1

<div class="post-metadata">

### Author: ![Jeff\_Dupont](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jeff_dupont/32/1355_2.png) [@Jeff\_Dupont](https://discuss.elastic.co/u/Jeff_Dupont)
#### Post date: [May 29, 2014, 8:47pm UTC](https://discuss.elastic.co/t/encoding-is-longer-than-the-max-length-32766/17819/1 "2014-05-29T20:47:37Z")

</div>

We’re running into a peculiar issue when updating indexes with content for  
the document.

"document contains at least one immense term in (whose utf8 encoding is  
longer than the max length 32766), all of which were skipped. please  
correct the analyzer to not produce such terms”

I’m hoping that there’s a simple fix or setting that can resolve this.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/01a22ff3-056d-4b54-8b28-a17f95d91f4b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/01a22ff3-056d-4b54-8b28-a17f95d91f4b%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![karmi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/karmi/32/44951_2.png) [@karmi](https://discuss.elastic.co/u/karmi)
#### Post date: [June 3, 2014, 4:18pm UTC](https://discuss.elastic.co/t/encoding-is-longer-than-the-max-length-32766/17819/2 "2014-06-03T16:18:37Z")

</div>

This is actually a change in Lucene -- previously, the long term was  
silently dropped, now it raises an exception, see Lucene  
ticket [[LUCENE-5710] DefaultIndexingChain swallows useful information from MaxBytesLengthExceededException - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-5710)

You might want to add a `length` filter to your analyzer  
([Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/analysis-length-tokenfilter.html#analysis-length-tokenfilter)).

All in all, it hints at some strange data, because such "immense" term  
shouldn't probably be in the index in the first place.

Karel

On Thursday, May 29, 2014 10:47:37 PM UTC+2, Jeff Dupont wrote:

> We’re running into a peculiar issue when updating indexes with content for  
> the document.
> 
> "document contains at least one immense term in (whose utf8 encoding is  
> longer than the max length 32766), all of which were skipped. please  
> correct the analyzer to not produce such terms”
> 
> I’m hoping that there’s a simple fix or setting that can resolve this.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/a91895cb-437a-4642-8734-4445bb420125%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a91895cb-437a-4642-8734-4445bb420125%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![Andrew\_Mehler](https://avatars.discourse-cdn.com/v4/letter/a/41988e/32.png) [@Andrew\_Mehler](https://discuss.elastic.co/u/Andrew_Mehler)
#### Post date: [July 1, 2014, 7:22pm UTC](https://discuss.elastic.co/t/encoding-is-longer-than-the-max-length-32766/17819/3 "2014-07-01T19:22:54Z")

</div>

For not analyzed fields, Is there a way of capturing the old behavior?  
From what I can tell, you need to specify a tokenizer to have a token  
filter.

On Tuesday, June 3, 2014 12:18:37 PM UTC-4, Karel Minařík wrote:

> This is actually a change in Lucene -- previously, the long term was  
> silently dropped, now it raises an exception, see Lucene ticket  
> [[LUCENE-5710] DefaultIndexingChain swallows useful information from MaxBytesLengthExceededException - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-5710)
> 
> You might want to add a `length` filter to your analyzer (  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/analysis-length-tokenfilter.html#analysis-length-tokenfilter)  
> ).
> 
> All in all, it hints at some strange data, because such "immense" term  
> shouldn't probably be in the index in the first place.
> 
> Karel
> 
> On Thursday, May 29, 2014 10:47:37 PM UTC+2, Jeff Dupont wrote:
> 
> > We’re running into a peculiar issue when updating indexes with content  
> > for the document.
> > 
> > "document contains at least one immense term in (whose utf8 encoding is  
> > longer than the max length 32766), all of which were skipped. please  
> > correct the analyzer to not produce such terms”
> > 
> > I’m hoping that there’s a simple fix or setting that can resolve this.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/26e3ad78-65a3-4853-ad26-8836c7bc2c7c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/26e3ad78-65a3-4853-ad26-8836c7bc2c7c%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![rore](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rore/32/399_2.png) [@rore](https://discuss.elastic.co/u/rore)
#### Post date: [October 30, 2014, 10:43am UTC](https://discuss.elastic.co/t/encoding-is-longer-than-the-max-length-32766/17819/4 "2014-10-30T10:43:26Z")

</div>

+1 on this question.

If the error is generated because of a not\_analyzed field, how is it  
possible to instruct ES to drop these values instead of failing the request?

On Tuesday, July 1, 2014 10:22:54 PM UTC+3, Andrew Mehler wrote:

> For not analyzed fields, Is there a way of capturing the old behavior?  
> From what I can tell, you need to specify a tokenizer to have a token  
> filter.
> 
> On Tuesday, June 3, 2014 12:18:37 PM UTC-4, Karel Minařík wrote:
> 
> > This is actually a change in Lucene -- previously, the long term was  
> > silently dropped, now it raises an exception, see Lucene ticket  
> > [[LUCENE-5710] DefaultIndexingChain swallows useful information from MaxBytesLengthExceededException - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-5710)
> > 
> > You might want to add a `length` filter to your analyzer (  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/analysis-length-tokenfilter.html#analysis-length-tokenfilter)  
> > ).
> > 
> > All in all, it hints at some strange data, because such "immense" term  
> > shouldn't probably be in the index in the first place.
> > 
> > Karel
> > 
> > On Thursday, May 29, 2014 10:47:37 PM UTC+2, Jeff Dupont wrote:
> > 
> > > We’re running into a peculiar issue when updating indexes with content  
> > > for the document.
> > > 
> > > "document contains at least one immense term in (whose utf8 encoding is  
> > > longer than the max length 32766), all of which were skipped. please  
> > > correct the analyzer to not produce such terms”
> > > 
> > > I’m hoping that there’s a simple fix or setting that can resolve this.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/8e6acaf8-7101-4d04-9566-43ea8845013c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/8e6acaf8-7101-4d04-9566-43ea8845013c%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![asthanaamish](https://avatars.discourse-cdn.com/v4/letter/a/6bbea6/32.png) [@asthanaamish](https://discuss.elastic.co/u/asthanaamish)
#### Post date: [January 2, 2015, 7:56pm UTC](https://discuss.elastic.co/t/encoding-is-longer-than-the-max-length-32766/17819/5 "2015-01-02T19:56:05Z")

</div>

How does this MAX\_LENGTH restriction impact on a custom\_all field where we  
may be copying data from different fields using some analyzer.  
Is the MAX\_LENGTH restriction also applicable on such custom\_all field  
which in turn implies that in such a case cumulative length is what matters.  
amish

On Thursday, October 30, 2014 3:43:26 AM UTC-7, Rotem wrote:

> +1 on this question.
> 
> If the error is generated because of a not\_analyzed field, how is it  
> possible to instruct ES to drop these values instead of failing the request?
> 
> On Tuesday, July 1, 2014 10:22:54 PM UTC+3, Andrew Mehler wrote:
> 
> > For not analyzed fields, Is there a way of capturing the old behavior?  
> > From what I can tell, you need to specify a tokenizer to have a token  
> > filter.
> > 
> > On Tuesday, June 3, 2014 12:18:37 PM UTC-4, Karel Minařík wrote:
> > 
> > > This is actually a change in Lucene -- previously, the long term was  
> > > silently dropped, now it raises an exception, see Lucene ticket  
> > > [[LUCENE-5710] DefaultIndexingChain swallows useful information from MaxBytesLengthExceededException - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-5710)
> > > 
> > > You might want to add a `length` filter to your analyzer (  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/analysis-length-tokenfilter.html#analysis-length-tokenfilter)  
> > > ).
> > > 
> > > All in all, it hints at some strange data, because such "immense" term  
> > > shouldn't probably be in the index in the first place.
> > > 
> > > Karel
> > > 
> > > On Thursday, May 29, 2014 10:47:37 PM UTC+2, Jeff Dupont wrote:
> > > 
> > > > We’re running into a peculiar issue when updating indexes with content  
> > > > for the document.
> > > > 
> > > > "document contains at least one immense term in (whose utf8 encoding is  
> > > > longer than the max length 32766), all of which were skipped. please  
> > > > correct the analyzer to not produce such terms”
> > > > 
> > > > I’m hoping that there’s a simple fix or setting that can resolve this.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/f19fafb9-a9e0-42a9-b290-a9b37d1da51d%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f19fafb9-a9e0-42a9-b290-a9b37d1da51d%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)
#### Post date: [January 3, 2015, 5:00am UTC](https://discuss.elastic.co/t/encoding-is-longer-than-the-max-length-32766/17819/6 "2015-01-03T05:00:45Z")

</div>

The max length restriction is per token so its unlikely you'll see it  
unless use not\_analyzed fields. You can work around it by setting the  
ignore\_above option on the string type. That'll just throw away the token.

Nik  
How does this MAX\_LENGTH restriction impact on a custom\_all field where we  
may be copying data from different fields using some analyzer.  
Is the MAX\_LENGTH restriction also applicable on such custom\_all field  
which in turn implies that in such a case cumulative length is what matters.  
amish

On Thursday, October 30, 2014 3:43:26 AM UTC-7, Rotem wrote:

> +1 on this question.
> 
> If the error is generated because of a not\_analyzed field, how is it  
> possible to instruct ES to drop these values instead of failing the request?
> 
> On Tuesday, July 1, 2014 10:22:54 PM UTC+3, Andrew Mehler wrote:
> 
> > For not analyzed fields, Is there a way of capturing the old behavior?  
> > From what I can tell, you need to specify a tokenizer to have a token  
> > filter.
> > 
> > On Tuesday, June 3, 2014 12:18:37 PM UTC-4, Karel Minařík wrote:
> > 
> > > This is actually a change in Lucene -- previously, the long term was  
> > > silently dropped, now it raises an exception, see Lucene ticket  
> > > [[LUCENE-5710] DefaultIndexingChain swallows useful information from MaxBytesLengthExceededException - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-5710)
> > > 
> > > You might want to add a `length` filter to your analyzer (  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/)  
> > > reference/current/analysis-length-tokenfilter.html#  
> > > analysis-length-tokenfilter).
> > > 
> > > All in all, it hints at some strange data, because such "immense" term  
> > > shouldn't probably be in the index in the first place.
> > > 
> > > Karel
> > > 
> > > On Thursday, May 29, 2014 10:47:37 PM UTC+2, Jeff Dupont wrote:
> > > 
> > > > We’re running into a peculiar issue when updating indexes with content  
> > > > for the document.
> > > > 
> > > > "document contains at least one immense term in (whose utf8 encoding is  
> > > > longer than the max length 32766), all of which were skipped. please  
> > > > correct the analyzer to not produce such terms”
> > > > 
> > > > I’m hoping that there’s a simple fix or setting that can resolve this.
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google Groups  
> > > "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send an  
> > > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > > To view this discussion on the web visit  
> > > [https://groups.google.com/d/msgid/elasticsearch/f19fafb9-a9e0-42a9-b290-a9b37d1da51d%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f19fafb9-a9e0-42a9-b290-a9b37d1da51d%40googlegroups.com)  
> > > [https://groups.google.com/d/msgid/elasticsearch/f19fafb9-a9e0-42a9-b290-a9b37d1da51d%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/f19fafb9-a9e0-42a9-b290-a9b37d1da51d%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > .  
> > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAPmjWd2mOPdgd0p\_DMek-a4hoiY7\_2eT96KUv\_biO%3D3ZgPLb4A%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAPmjWd2mOPdgd0p_DMek-a4hoiY7_2eT96KUv_biO%3D3ZgPLb4A%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 12:41am UTC](https://discuss.elastic.co/t/encoding-is-longer-than-the-max-length-32766/17819/7 "2017-07-06T00:41:04Z")

</div>


