# Multiple tokens with same position

**URL:** <https://discuss.elastic.co/t/multiple-tokens-with-same-position/50006>\
**Category:** Elasticsearch\
**Created:** [May 13, 2016, 2:49pm UTC](https://discuss.elastic.co/t/multiple-tokens-with-same-position/50006 "2016-05-13T14:49:32Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![philv](https://avatars.discourse-cdn.com/v4/letter/p/ebca7d/32.png) [@philv](https://discuss.elastic.co/u/philv)\
**Post date:** [May 13, 2016, 2:49pm UTC](https://discuss.elastic.co/t/multiple-tokens-with-same-position/50006/1 "2016-05-13T14:49:32Z")

</div>

I have a situation where I have multiple tokens at the same position (post a PatternCaptureTokenFilter).  
If I had a sentence like this for example:  
"blue car\_automobile\_wheeled\_thing road"

And I split on underscore, I'd end up with a token list looking like this

Token Position Token String  
0 blue  
1 car  
1 automobile  
1 wheeled  
1 thing  
2 road

This is akin to the multi word synonym problem as described in this brilliant blog [post](http://blog.mikemccandless.com/2012/04/lucenes-tokenstreams-are-actually.html)

Note that 'wheeled thing' is the multi word synonym effectively.

Lucene , afaik, doesn't use the PositionLengthAttribute, so the bag of tokens at position 1 is unordered.  
If wanted to search for the phrase 'wheeled thing', I'd find it, but if I searched for 'blue thing', I'd find that too, erroneously, because 'thing' is one position away from 'blue'. Has anyone got a solution to this kind of multi-word synonym issue ?

Many thanks,  
Phil

---

<div class="post-metadata">

**Author:** ![mikemccand](https://avatars.discourse-cdn.com/v4/letter/m/f04885/32.png) [@mikemccand](https://discuss.elastic.co/u/mikemccand)\
**Post date:** [May 13, 2016, 3:18pm UTC](https://discuss.elastic.co/t/multiple-tokens-with-same-position/50006/2 "2016-05-13T15:18:23Z")

</div>

Assuming you had indexed `blue car road` and had synonyms `car -> automobile` and `car -> wheeled thing`, then in your example, `thing` should be at position 2 not 1 (i.e., it overlaps `road` not `car`), the way synonym filter works today.

And then "blue thing" phrase query should _not_ match (good), but e.g. "wheeled thing road" won't match but should (bad).

Essentially, the synonym filter cannot create new positions, so it takes multi-term synonyms and lays them on top of the existing tokens.

There is a working patch on [https://issues.apache.org/jira/browse/LUCENE-6664](https://issues.apache.org/jira/browse/LUCENE-6664) to let synonym filter create new positions, so that it produces a correct graph, but it was controversial and got shelved.

---

<div class="post-metadata">

**Author:** ![philv](https://avatars.discourse-cdn.com/v4/letter/p/ebca7d/32.png) [@philv](https://discuss.elastic.co/u/philv)\
**Post date:** [May 13, 2016, 3:58pm UTC](https://discuss.elastic.co/t/multiple-tokens-with-same-position/50006/3 "2016-05-13T15:58:49Z")

</div>

Many thanks Mike !  
You're right of course, my example was wrong. A better example to explain our situation would be:

"blue big\_wheeled\_thing\_CONCEPTCAR road" where we embed the concept CAR into the same token position as the words "big", "wheeled" and "thing". We can search with CONCEPTCAR fine. The issue we face is that we want to be able to allow search hits for 'big wheeled thing' , but not hits for 'big thing'. It's the same issue underneath as that your blog mentioned I think. Are there any other tricks we might use to get around this lucene limitation and thereby achieve no hits for 'big thing' ?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:51pm UTC](https://discuss.elastic.co/t/multiple-tokens-with-same-position/50006/4 "2017-07-05T22:51:44Z")

</div>


