# Exceptions during Highlighting: InvalidTokenOffsetsException

**URL:** <https://discuss.elastic.co/t/exceptions-during-highlighting-invalidtokenoffsetsexception/6145>\
**Category:** Elasticsearch\
**Created:** [December 13, 2011, 5:06pm UTC](https://discuss.elastic.co/t/exceptions-during-highlighting-invalidtokenoffsetsexception/6145 "2011-12-13T17:06:44Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![BowlingX](https://avatars.discourse-cdn.com/v4/letter/b/edb3f5/32.png) [@BowlingX](https://discuss.elastic.co/u/BowlingX)\
**Post date:** [December 13, 2011, 5:06pm UTC](https://discuss.elastic.co/t/exceptions-during-highlighting-invalidtokenoffsetsexception/6145/1 "2011-12-13T17:06:44Z")

</div>

Hello,

im currently working on an autocompletion for a large dataset  
(geonames).  
My Schema: [https://gist.github.com/424ce0205a9a16e7afe1](https://gist.github.com/424ce0205a9a16e7afe1)

I've imported a few Countries to run some tests. Most Queries are  
successfull, but sometimes the following exception rises:

[...]  
Fetch Failed [Failed to highlight field [name.partial]]]; nested:  
InvalidTokenOffsetsException[Token dussvitz exceeds length of provided  
text sized 7];  
[...]

The Query looks like (see my comment in gist):

> <https://gist.github.com/BowlingX/424ce0205a9a16e7afe1>

Exception ONLY rises when highlighting on Fields like "name.partial or  
name.partial\_non\_ascii or alternateNames.partial etc."

I've no idea anymore and hope that someone can help me out.

---

<div class="post-metadata">

**Author:** ![BowlingX](https://avatars.discourse-cdn.com/v4/letter/b/edb3f5/32.png) [@BowlingX](https://discuss.elastic.co/u/BowlingX)\
**Post date:** [December 14, 2011, 10:12am UTC](https://discuss.elastic.co/t/exceptions-during-highlighting-invalidtokenoffsetsexception/6145/2 "2011-12-14T10:12:11Z")

</div>

I found it :), it seems like a bug in "edgeNGram" Filter, i switched  
the position of my filters from:

index.analysis.analyzer.partial.filter.3: name\_ngrams  
index.analysis.analyzer.partial.filter.2: asciifolding  
index.analysis.analyzer.partial.filter.1: lowercase  
index.analysis.analyzer.partial.filter.0: standard

to

index.analysis.analyzer.partial.filter.3: asciifolding  
index.analysis.analyzer.partial.filter.2: name\_ngrams  
index.analysis.analyzer.partial.filter.1: lowercase  
index.analysis.analyzer.partial.filter.0: standard

an it works, could be that problem:

[http://mail-archives.apache.org/mod\_mbox/lucene-solr-user/201112.mbox/\<CAOdYfZU\_pe3P-xspsACOhuJwNYTj+=K47uE6a4LYa7=jabB+2A@mail.gmail.com\>](http://mail-archives.apache.org/mod_mbox/lucene-solr-user/201112.mbox/%3CCAOdYfZU_pe3P-xspsACOhuJwNYTj+=K47uE6a4LYa7=jabB+2A@mail.gmail.com%3E)

[https://issues.apache.org/jira/browse/LUCENE-1500](https://issues.apache.org/jira/browse/LUCENE-1500)

Maybe i should file a bug report

On 13 Dez., 18:06, BowlingX [heidrich.da...@googlemail.com](mailto:heidrich.da...@googlemail.com) wrote:

> Hello,
> 
> im currently working on an autocompletion for a large dataset  
> (geonames).  
> My Schema:[Geonames.org Meta Settings and Mapping · GitHub](https://gist.github.com/424ce0205a9a16e7afe1)
> 
> I've imported a few Countries to run some tests. Most Queries are  
> successfull, but sometimes the following exception rises:
> 
> [...]  
> Fetch Failed [Failed to highlight field [name.partial]]]; nested:  
> InvalidTokenOffsetsException[Token dussvitz exceeds length of provided  
> text sized 7];  
> [...]
> 
> The Query looks like (see my comment in gist):[Geonames.org Meta Settings and Mapping · GitHub](https://gist.github.com/424ce0205a9a16e7afe1#comments)
> 
> Exception ONLY rises when highlighting on Fields like "name.partial or  
> name.partial\_non\_ascii or alternateNames.partial etc."
> 
> I've no idea anymore and hope that someone can help me out.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [December 16, 2011, 3:35pm UTC](https://discuss.elastic.co/t/exceptions-during-highlighting-invalidtokenoffsetsexception/6145/3 "2011-12-16T15:35:19Z")

</div>

Nice catch!

On Wed, Dec 14, 2011 at 12:12 PM, BowlingX [heidrich.david@googlemail.com](mailto:heidrich.david@googlemail.com)wrote:

> I found it :), it seems like a bug in "edgeNGram" Filter, i switched  
> the position of my filters from:
> 
> index.analysis.analyzer.partial.filter.3: name\_ngrams  
> index.analysis.analyzer.partial.filter.2: asciifolding  
> index.analysis.analyzer.partial.filter.1: lowercase  
> index.analysis.analyzer.partial.filter.0: standard
> 
> to
> 
> index.analysis.analyzer.partial.filter.3: asciifolding  
> index.analysis.analyzer.partial.filter.2: name\_ngrams  
> index.analysis.analyzer.partial.filter.1: lowercase  
> index.analysis.analyzer.partial.filter.0: standard
> 
> an it works, could be that problem:
> 
> [http://mail-archives.apache.org/mod\_mbox/lucene-solr-user/201112.mbox/\<CAOdYfZU\_pe3P-xspsACOhuJwNYTj+=K47uE6a4LYa7=jabB+2A@mail.gmail.com\>](http://mail-archives.apache.org/mod_mbox/lucene-solr-user/201112.mbox/%3CCAOdYfZU_pe3P-xspsACOhuJwNYTj+=K47uE6a4LYa7=jabB+2A@mail.gmail.com%3E)
> 
> [[LUCENE-1500] Highlighter throws StringIndexOutOfBoundsException - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-1500)
> 
> Maybe i should file a bug report
> 
> On 13 Dez., 18:06, BowlingX [heidrich.da...@googlemail.com](mailto:heidrich.da...@googlemail.com) wrote:
> 
> > Hello,
> > 
> > im currently working on an autocompletion for a large dataset  
> > (geonames).  
> > My Schema:[Geonames.org Meta Settings and Mapping · GitHub](https://gist.github.com/424ce0205a9a16e7afe1)
> > 
> > I've imported a few Countries to run some tests. Most Queries are  
> > successfull, but sometimes the following exception rises:
> > 
> > [...]  
> > Fetch Failed [Failed to highlight field [name.partial]]]; nested:  
> > InvalidTokenOffsetsException[Token dussvitz exceeds length of provided  
> > text sized 7];  
> > [...]
> > 
> > The Query looks like (see my comment in gist):  
> > [Geonames.org Meta Settings and Mapping · GitHub](https://gist.github.com/424ce0205a9a16e7afe1#comments)
> > 
> > Exception ONLY rises when highlighting on Fields like "name.partial or  
> > name.partial\_non\_ascii or alternateNames.partial etc."
> > 
> > I've no idea anymore and hope that someone can help me out.

---

<div class="post-metadata">

**Author:** ![BowlingX](https://avatars.discourse-cdn.com/v4/letter/b/edb3f5/32.png) [@BowlingX](https://discuss.elastic.co/u/BowlingX)\
**Post date:** [December 16, 2011, 11:23pm UTC](https://discuss.elastic.co/t/exceptions-during-highlighting-invalidtokenoffsetsexception/6145/4 "2011-12-16T23:23:53Z")

</div>

So, is it a bug? Or am I doing anythin wrong?

2011/12/16 Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com)

> Nice catch!
> 
> On Wed, Dec 14, 2011 at 12:12 PM, BowlingX [heidrich.david@googlemail.com](mailto:heidrich.david@googlemail.com)wrote:
> 
> > I found it :), it seems like a bug in "edgeNGram" Filter, i switched  
> > the position of my filters from:
> > 
> > index.analysis.analyzer.partial.filter.3: name\_ngrams  
> > index.analysis.analyzer.partial.filter.2: asciifolding  
> > index.analysis.analyzer.partial.filter.1: lowercase  
> > index.analysis.analyzer.partial.filter.0: standard
> > 
> > to
> > 
> > index.analysis.analyzer.partial.filter.3: asciifolding  
> > index.analysis.analyzer.partial.filter.2: name\_ngrams  
> > index.analysis.analyzer.partial.filter.1: lowercase  
> > index.analysis.analyzer.partial.filter.0: standard
> > 
> > an it works, could be that problem:
> > 
> > [http://mail-archives.apache.org/mod\_mbox/lucene-solr-user/201112.mbox/\<CAOdYfZU\_pe3P-xspsACOhuJwNYTj+=K47uE6a4LYa7=jabB+2A@mail.gmail.com\>](http://mail-archives.apache.org/mod_mbox/lucene-solr-user/201112.mbox/%3CCAOdYfZU_pe3P-xspsACOhuJwNYTj+=K47uE6a4LYa7=jabB+2A@mail.gmail.com%3E)
> > 
> > [[LUCENE-1500] Highlighter throws StringIndexOutOfBoundsException - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-1500)
> > 
> > Maybe i should file a bug report
> > 
> > On 13 Dez., 18:06, BowlingX [heidrich.da...@googlemail.com](mailto:heidrich.da...@googlemail.com) wrote:
> > 
> > > Hello,
> > > 
> > > im currently working on an autocompletion for a large dataset  
> > > (geonames).  
> > > My Schema:[Geonames.org Meta Settings and Mapping · GitHub](https://gist.github.com/424ce0205a9a16e7afe1)
> > > 
> > > I've imported a few Countries to run some tests. Most Queries are  
> > > successfull, but sometimes the following exception rises:
> > > 
> > > [...]  
> > > Fetch Failed [Failed to highlight field [name.partial]]]; nested:  
> > > InvalidTokenOffsetsException[Token dussvitz exceeds length of provided  
> > > text sized 7];  
> > > [...]
> > > 
> > > The Query looks like (see my comment in gist):  
> > > [Geonames.org Meta Settings and Mapping · GitHub](https://gist.github.com/424ce0205a9a16e7afe1#comments)
> > > 
> > > Exception ONLY rises when highlighting on Fields like "name.partial or  
> > > name.partial\_non\_ascii or alternateNames.partial etc."
> > > 
> > > I've no idea anymore and hope that someone can help me out.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [December 20, 2011, 3:12pm UTC](https://discuss.elastic.co/t/exceptions-during-highlighting-invalidtokenoffsetsexception/6145/5 "2011-12-20T15:12:13Z")

</div>

I don't know, need to check in Lucene.

On Sat, Dec 17, 2011 at 1:23 AM, David Heidrich \<  
[heidrich.david@googlemail.com](mailto:heidrich.david@googlemail.com)\> wrote:

> So, is it a bug? Or am I doing anythin wrong?
> 
> 2011/12/16 Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com)
> 
> > Nice catch!
> > 
> > On Wed, Dec 14, 2011 at 12:12 PM, BowlingX \<[heidrich.david@googlemail.com](mailto:heidrich.david@googlemail.com)
> > 
> > > wrote:
> > 
> > > I found it :), it seems like a bug in "edgeNGram" Filter, i switched  
> > > the position of my filters from:
> > > 
> > > index.analysis.analyzer.partial.filter.3: name\_ngrams  
> > > index.analysis.analyzer.partial.filter.2: asciifolding  
> > > index.analysis.analyzer.partial.filter.1: lowercase  
> > > index.analysis.analyzer.partial.filter.0: standard
> > > 
> > > to
> > > 
> > > index.analysis.analyzer.partial.filter.3: asciifolding  
> > > index.analysis.analyzer.partial.filter.2: name\_ngrams  
> > > index.analysis.analyzer.partial.filter.1: lowercase  
> > > index.analysis.analyzer.partial.filter.0: standard
> > > 
> > > an it works, could be that problem:
> > > 
> > > [http://mail-archives.apache.org/mod\_mbox/lucene-solr-user/201112.mbox/\<CAOdYfZU\_pe3P-xspsACOhuJwNYTj+=K47uE6a4LYa7=jabB+2A@mail.gmail.com\>](http://mail-archives.apache.org/mod_mbox/lucene-solr-user/201112.mbox/%3CCAOdYfZU_pe3P-xspsACOhuJwNYTj+=K47uE6a4LYa7=jabB+2A@mail.gmail.com%3E)
> > > 
> > > [[LUCENE-1500] Highlighter throws StringIndexOutOfBoundsException - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-1500)
> > > 
> > > Maybe i should file a bug report
> > > 
> > > On 13 Dez., 18:06, BowlingX [heidrich.da...@googlemail.com](mailto:heidrich.da...@googlemail.com) wrote:
> > > 
> > > > Hello,
> > > > 
> > > > im currently working on an autocompletion for a large dataset  
> > > > (geonames).  
> > > > My Schema:[Geonames.org Meta Settings and Mapping · GitHub](https://gist.github.com/424ce0205a9a16e7afe1)
> > > > 
> > > > I've imported a few Countries to run some tests. Most Queries are  
> > > > successfull, but sometimes the following exception rises:
> > > > 
> > > > [...]  
> > > > Fetch Failed [Failed to highlight field [name.partial]]]; nested:  
> > > > InvalidTokenOffsetsException[Token dussvitz exceeds length of provided  
> > > > text sized 7];  
> > > > [...]
> > > > 
> > > > The Query looks like (see my comment in gist):  
> > > > [Geonames.org Meta Settings and Mapping · GitHub](https://gist.github.com/424ce0205a9a16e7afe1#comments)
> > > > 
> > > > Exception ONLY rises when highlighting on Fields like "name.partial or  
> > > > name.partial\_non\_ascii or alternateNames.partial etc."
> > > > 
> > > > I've no idea anymore and hope that someone can help me out.

---

<div class="post-metadata">

**Author:** ![Kairos](https://avatars.discourse-cdn.com/v4/letter/k/ecc23a/32.png) [@Kairos](https://discuss.elastic.co/u/Kairos)\
**Post date:** [March 6, 2015, 8:52pm UTC](https://discuss.elastic.co/t/exceptions-during-highlighting-invalidtokenoffsetsexception/6145/6 "2015-03-06T20:52:32Z")

</div>

The bug is caused from a wrong calculation of the tokens' offsets. There are some filters that generate additional tokens with a text.length longer than the original token (the one before the (re-)analyzing).  
This provokes the wrong calculation of the start and end offsets, ending in a literal hell when trying to highlight.

In fact it will try to add some highlighting tags before the start offset and after the end offset of the token.  
But if the token has wrong offset calculated and it is in the end of the field value, then with a high probability the highlighter tries to write the tags over the field length....... raising the exception.

Am having the same problem with an additional plug-in for decompounding German words.

And actually I don't have idea of how to solve it (but I wonder if others filters, like the stemmers, generate longer tokens, and how they manage the offsets correctly).

It's an old post, but maybe can be useful to someone...

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:27am UTC](https://discuss.elastic.co/t/exceptions-during-highlighting-invalidtokenoffsetsexception/6145/7 "2017-07-06T00:27:58Z")

</div>


