# PatternCaptureGroupTokenFilter throwing error - startOffset must be non-negative, and endOffset must be \>= startOffset, and offsets must not go backwards

**URL:** <https://discuss.elastic.co/t/patterncapturegrouptokenfilter-throwing-error-startoffset-must-be-non-negative-and-endoffset-must-be-startoffset-and-offsets-must-not-go-backwards/366199>\
**Category:** Elasticsearch\
**Created:** [September 8, 2024, 10:28am UTC](https://discuss.elastic.co/t/patterncapturegrouptokenfilter-throwing-error-startoffset-must-be-non-negative-and-endoffset-must-be-startoffset-and-offsets-must-not-go-backwards/366199 "2024-09-08T10:28:53Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![shikha65786](https://avatars.discourse-cdn.com/v4/letter/s/d78d45/32.png) [@shikha65786](https://discuss.elastic.co/u/shikha65786)\
**Post date:** [September 8, 2024, 10:28am UTC](https://discuss.elastic.co/t/patterncapturegrouptokenfilter-throwing-error-startoffset-must-be-non-negative-and-endoffset-must-be-startoffset-and-offsets-must-not-go-backwards/366199/1 "2024-09-08T10:28:53Z")

</div>

I'm using the `PatternCaptureGroupTokenFilter` in my code to generate tokens based on multiple regular expressions and **highlight** matches in the string. I'm working with Lucene 9, but it's returning the following error.

`{"error":{"root_cause":[{"type":"illegal_argument_exception","reason":"startOffset must be non-negative, and endOffset must be >= startOffset, and offsets must not go backwards startOffset=0,endOffset=5,lastStartOffset=4 for field 'title_special.en'"}],"type":"illegal_argument_exception","reason":"startOffset must be non-negative, and endOffset must be >= startOffset, and offsets must not go backwards startOffset=0,endOffset=5,lastStartOffset=4 for field 'title_special.en'"},"status":400}`

```auto
{
“tokens” : [
{
“token” : “test:data”,
“start_offset” : 0,
“end_offset” : 9,
“type” : “word”,
“position” : 0
},
{
“token” : “test”,
“start_offset” : 0,
“end_offset” : 4,
“type” : “word”,
“position” : 0
},
{
“token” : “:data”,
“start_offset” : 4,
“end_offset” : 9,
“type” : “word”,
“position” : 0
},
{
“token” : “test:”,
“start_offset” : 0,
“end_offset” : 5,
“type” : “word”,
“position” : 0
},
{
“token” : “test”,
“start_offset” : 10,
“end_offset” : 14,
“type” : “word”,
“position” : 1
}
]
}

```

I am using the following Java code

```auto
public final class PatternCaptureGroupTokenFilter extends TokenFilter {

    private final CharTermAttribute charTermAttr = addAttribute(CharTermAttribute.class);
    private final PositionIncrementAttribute posAttr = addAttribute(PositionIncrementAttribute.class);
    private final OffsetAttribute offsetAtt = addAttribute(OffsetAttribute.class);

    private final TypeAttribute typeAttribute = addAttribute(TypeAttribute.class);
    private State state;
    private final Matcher[] matchers;
    private final CharsRefBuilder spare = new CharsRefBuilder();
    private final int[] groupCounts;
    private final boolean preserveOriginal;
    private int[] currentGroup;
    private int currentMatcher;
    private int main_token_start;
    private int main_token_end;

    public PatternCaptureGroupTokenFilter(TokenStream input,
        boolean preserveOriginal, Pattern... patterns) {
        super(input);

        this.preserveOriginal = preserveOriginal;
        this.matchers = new Matcher[patterns.length];
        this.groupCounts = new int[patterns.length];
        this.currentGroup = new int[patterns.length];
        for (int i = 0; i < patterns.length; i++) {
            this.matchers[i] = patterns[i].matcher("");
            this.groupCounts[i] = this.matchers[i].groupCount();
            this.currentGroup[i] = -1;
        }
    }

    private boolean nextCapture() {

        int min_offset = Integer.MAX_VALUE;
        currentMatcher = -1;
        Matcher matcher;

        for (int i = 0; i < matchers.length; i++) {
            matcher = matchers[i];
            if (currentGroup[i] == -1) {
                currentGroup[i] = matcher.find() ? 1 : 0;
            }
            if (currentGroup[i] != 0) {
                while (currentGroup[i] < groupCounts[i] + 1) {
                    final int start = matcher.start(currentGroup[i]);
                    final int end = matcher.end(currentGroup[i]);

                    if (start == end || preserveOriginal && start == 0
                            && spare.length() == end) {
                        currentGroup[i]++;
                        continue;
                    }
                    if (start < min_offset) {
                        min_offset = start;
                        currentMatcher = i;
                    }
                    break;
                }
                if (currentGroup[i] == groupCounts[i] + 1) {
                    currentGroup[i] = -1;
                    i--;
                }
            }
        }
        return currentMatcher != -1;
    }

    @Override
    public boolean incrementToken() throws IOException {
        if (currentMatcher != -1 && nextCapture()) {
            assert state != null;
            clearAttributes();
            restoreState(state);
            final int start = matchers[currentMatcher]
                    .start(currentGroup[currentMatcher]);
            final int end = matchers[currentMatcher]
                    .end(currentGroup[currentMatcher]);

            // modified code starts
            main_token_start = offsetAtt.startOffset();
            main_token_end = offsetAtt.endOffset();

            final int newStart = start + main_token_start;
            final int newEnd = end + main_token_start;

            offsetAtt.setOffset(newStart, newEnd);
            // modified code ends

            posAttr.setPositionIncrement(0);

            charTermAttr.copyBuffer(spare.chars(), start, end - start);
            currentGroup[currentMatcher]++;
            return true;
        }

        if (!input.incrementToken()) {
            return false;
        }

        char[] buffer = charTermAttr.buffer();
        int length = charTermAttr.length();
        spare.copyChars(buffer, 0, length);
        state = captureState();

        for (int i = 0; i < matchers.length; i++) {
            matchers[i].reset(spare.get());
            currentGroup[i] = -1;
        }

        if (preserveOriginal) {
            currentMatcher = 0;
        } else if (nextCapture()) {
            final int start = matchers[currentMatcher]
                    .start(currentGroup[currentMatcher]);
            final int end = matchers[currentMatcher]
                    .end(currentGroup[currentMatcher]);

            // if we start at 0 we can simply set the length and save the copy
            if (start == 0) {
                charTermAttr.setLength(end);
            } else {
                charTermAttr.copyBuffer(spare.chars(), start, end - start);
            }
            currentGroup[currentMatcher]++;
        }

        return true;

    }

    @Override
    public void reset() throws IOException {
        super.reset();
        state = null;
        currentMatcher = -1;
    }
}

```

This error is occurring in Lucene's latest versions. Please refer to: [Start offset going backwards has a legitimate purpose [LUCENE-8776] · Issue #9820 · apache/lucene · GitHub](https://github.com/apache/lucene/issues/9820)

Can anyone suggest how I can achieve this?

---

<div class="post-metadata">

**Author:** ![carly.richmond](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/carly.richmond/32/104935_2.png) [@carly.richmond](https://discuss.elastic.co/u/carly.richmond)\
**Post date:** [September 9, 2024, 1:39pm UTC](https://discuss.elastic.co/t/patterncapturegrouptokenfilter-throwing-error-startoffset-must-be-non-negative-and-endoffset-must-be-startoffset-and-offsets-must-not-go-backwards/366199/2 "2024-09-09T13:39:37Z")

</div>

From #Elastic Search to #Elasticsearch

---

<div class="post-metadata">

**Author:** ![carly.richmond](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/carly.richmond/32/104935_2.png) [@carly.richmond](https://discuss.elastic.co/u/carly.richmond)\
**Post date:** [September 9, 2024, 1:42pm UTC](https://discuss.elastic.co/t/patterncapturegrouptokenfilter-throwing-error-startoffset-must-be-non-negative-and-endoffset-must-be-startoffset-and-offsets-must-not-go-backwards/366199/3 "2024-09-09T13:42:22Z")

</div>

Hi @shikha65786,

Welcome! I see you've raised a [similar issue here](https://discuss.elastic.co/t/patterncapturegrouptokenfilter-is-creating-the-same-offset-positions-which-is-causing-highlighting-issue/366200) as well.

To confirm are you using Elasticsearch, or native Lucene? It looks like the latter so I wanted to confirm.

---

<div class="post-metadata">

**Author:** ![shikha65786](https://avatars.discourse-cdn.com/v4/letter/s/d78d45/32.png) [@shikha65786](https://discuss.elastic.co/u/shikha65786)\
**Post date:** [September 10, 2024, 5:11am UTC](https://discuss.elastic.co/t/patterncapturegrouptokenfilter-throwing-error-startoffset-must-be-non-negative-and-endoffset-must-be-startoffset-and-offsets-must-not-go-backwards/366199/4 "2024-09-10T05:11:16Z")

</div>

Hi @carly.richmond , Thanks for responding!

I am using the `PatternCaptureGroupTokenFilter` with multiple regex patterns to create tokens within a custom tokenizer.

I've tried two different approaches, but neither is delivering the expected results. One approach, as mentioned in this discussion, is causing an error during document indexing in Elasticsearch. In the other approach, I used Lucene's standard `PatternCaptureGroupTokenFilter`, but it is generating the same offset positions for each token, which makes it unsuitable for highlighting.

In summary, I'm using the Lucene `PatternCaptureGroupTokenFilter` within a custom tokenizer to index data in Elasticsearch.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 8, 2024, 5:11am UTC](https://discuss.elastic.co/t/patterncapturegrouptokenfilter-throwing-error-startoffset-must-be-non-negative-and-endoffset-must-be-startoffset-and-offsets-must-not-go-backwards/366199/5 "2024-10-08T05:11:31Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
