# Match every token position in the field when using synonyms

**URL:** <https://discuss.elastic.co/t/match-every-token-position-in-the-field-when-using-synonyms/15272>\
**Category:** Elasticsearch\
**Created:** [January 16, 2014, 2:12pm UTC](https://discuss.elastic.co/t/match-every-token-position-in-the-field-when-using-synonyms/15272 "2014-01-16T14:12:23Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dany\_Gielow](https://avatars.discourse-cdn.com/v4/letter/d/3bc359/32.png) [@Dany\_Gielow](https://discuss.elastic.co/u/Dany_Gielow)\
**Post date:** [January 16, 2014, 2:12pm UTC](https://discuss.elastic.co/t/match-every-token-position-in-the-field-when-using-synonyms/15272/1 "2014-01-16T14:12:23Z")

</div>

In my Elasticsearch index I have documents that have multiple tokens at the  
same position.

I want to get a document back when I match at least one token at every  
position.  
The order of the tokens is not important. How can I accomplish that?  
I use Elasticsearch 0.90.5.

_Example:_

I index a document like this.

```
{
    "field":"red car"
}

```

I use a synonym token filter that adds synonyms at the same positions as  
the original token.  
So now in the field, there are 2 positions:

- Position 1: "red"
- Position 2: "car", "automobile"

_My solution for now:_

To be able to ensure that all positions match, I index the maximum position  
as well.

```
{
    "field":"red car",
    "max_position": 2
}

```

I have a custom similarity that extends from DefaultSimilarity and returns  
1 tf(), idf() and lengthNorm(). The resulting score is the number of  
matching terms in the field.

Query:

```
{
    "custom_score": {
        "query": {
             "match": {
                 "field": "a car is an automobile"
             }
        },
        "_script": "_score*100/doc[\"max_position\"]+_score"
    },
    "min_score":"100"
}

```

Enter code here...

_Problem with my solution:_  
The above search should not match the document, because there is no token  
"red" in the query string. But it matches, because Elasticsearch counts the  
matches for car and automobile as two matches and that gives a score of 2  
which leads to a script score of 102, which satisfies the "min\_score".

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/77d52c69-8862-4e10-8036-470bf4ca8189%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/77d52c69-8862-4e10-8036-470bf4ca8189%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Sebastian\_Briesemeis](https://avatars.discourse-cdn.com/v4/letter/s/c57346/32.png) [@Sebastian\_Briesemeis](https://discuss.elastic.co/u/Sebastian_Briesemeis)\
**Post date:** [January 22, 2014, 8:29am UTC](https://discuss.elastic.co/t/match-every-token-position-in-the-field-when-using-synonyms/15272/2 "2014-01-22T08:29:54Z")

</div>

I am also very keen on answer!! If you find a solution, let me know!

Sebastian

On Thursday, 16 January 2014 15:12:23 UTC+1, Dany Gielow wrote:

> In my Elasticsearch index I have documents that have multiple tokens at  
> the same position.
> 
> I want to get a document back when I match at least one token at every  
> position.  
> The order of the tokens is not important. How can I accomplish that?  
> I use Elasticsearch 0.90.5.
> 
> _Example:_
> 
> I index a document like this.
> 
> ```
> {
> "field":"red car"
> }
> 
> ```
> 
> I use a synonym token filter that adds synonyms at the same positions as  
> the original token.  
> So now in the field, there are 2 positions:
> 
> - Position 1: "red"
> - Position 2: "car", "automobile"
> 
> _My solution for now:_
> 
> To be able to ensure that all positions match, I index the maximum  
> position as well.
> 
> ```
> {
> "field":"red car",
> "max_position": 2
> }
> 
> ```
> 
> I have a custom similarity that extends from DefaultSimilarity and returns  
> 1 tf(), idf() and lengthNorm(). The resulting score is the number of  
> matching terms in the field.
> 
> Query:
> 
> ```
> {
> "custom_score": {
> "query": {
> "match": {
> "field": "a car is an automobile"
> }
> },
> "_script": "_score*100/doc[\"max_position\"]+_score"
> },
> "min_score":"100"
> }
> 
> ```
> 
> Enter code here...
> 
> _Problem with my solution:_  
> The above search should not match the document, because there is no token  
> "red" in the query string. But it matches, because Elasticsearch counts the  
> matches for car and automobile as two matches and that gives a score of 2  
> which leads to a script score of 102, which satisfies the "min\_score".

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ca3aaaaa-dffc-4714-8940-0278cf70a7cf%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ca3aaaaa-dffc-4714-8940-0278cf70a7cf%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:55am UTC](https://discuss.elastic.co/t/match-every-token-position-in-the-field-when-using-synonyms/15272/3 "2017-07-06T01:55:27Z")

</div>


