# Annotated Text Plugin: full-text queries for annotated\_text

**URL:** https://discuss.elastic.co/t/annotated-text-plugin-full-text-queries-for-annotated-text/232308
**Category:** Elasticsearch
**Created:** [May 12, 2020, 7:40pm UTC](https://discuss.elastic.co/t/annotated-text-plugin-full-text-queries-for-annotated-text/232308 "2020-05-12T19:40:44Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![JohnathanBostrom](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/johnathanbostrom/32/68189_2.png) [@JohnathanBostrom](https://discuss.elastic.co/u/JohnathanBostrom)
#### Post date: [May 12, 2020, 7:40pm UTC](https://discuss.elastic.co/t/annotated-text-plugin-full-text-queries-for-annotated-text/232308/1 "2020-05-12T19:40:44Z")

</div>

I would like to use a full-text query on annotated text. [A similar question](https://discuss.elastic.co/t/can-you-change-the-type-of-a-token-emitted-by-an-analyzer/191100) was asked last year, but didn't receive any responses.  
Just as that poster, I want to be able to do the following:  
if I have the annotated text

```auto
Text1: "[Thomas Jefferson](_president_) was born in [Virginia](_place_)"
Text2: "[Thomas Jefferson](_writer_) was born in [Virginia](_place_)."  

```

I'd like to be able to execute the query

```auto
{
    "match_phrase": {
        "annotatedField": "_president_ was born in Virginia"
    }
}

```

and have it match Text1 and but not Text2.  
Is there any way to do this with annotated text? If not, is there a way to do it with payloads?  
Thanks!

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [May 12, 2020, 8:11pm UTC](https://discuss.elastic.co/t/annotated-text-plugin-full-text-queries-for-annotated-text/232308/2 "2020-05-12T20:11:23Z")

</div>

The problem you are facing is that the match\_phrase query clause will feed your input through an analyzer which will probably strip the underscores off your ‘_president_’ and fail to match. Your search string has to be presented in a way where the terms are not tokenized which will involve more JSON - see the Span or Interval query clauses.

---

<div class="post-metadata">

### Author: ![JohnathanBostrom](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/johnathanbostrom/32/68189_2.png) [@JohnathanBostrom](https://discuss.elastic.co/u/JohnathanBostrom)
#### Post date: [May 12, 2020, 8:25pm UTC](https://discuss.elastic.co/t/annotated-text-plugin-full-text-queries-for-annotated-text/232308/3 "2020-05-12T20:25:17Z")

</div>

Thanks for the reply! I just tried adding the annotations without the underscores, and it worked. Thanks for pointing me in the right direction!

---

<div class="post-metadata">

### Author: ![JohnathanBostrom](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/johnathanbostrom/32/68189_2.png) [@JohnathanBostrom](https://discuss.elastic.co/u/JohnathanBostrom)
#### Post date: [May 12, 2020, 9:08pm UTC](https://discuss.elastic.co/t/annotated-text-plugin-full-text-queries-for-annotated-text/232308/4 "2020-05-12T21:08:10Z")

</div>

It looks like match\_phrase doesn't work well when the annotated text contains spaces.  
If I have the following:

```auto
Text1: "when [Thomas Jefferson](ispresident) was born in [Virginia](isplace)"
Text2: "when [Jefferson](ispresident) was born in [Virginia](isplace)." 

```

then this matches Text2, but not Text1:

```auto
{
    "match_phrase": {
        "annotatedField": "when ispresident was born in Virginia"
    }
}

```

I would expect it to match both. Any insights?

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [May 13, 2020, 7:21am UTC](https://discuss.elastic.co/t/annotated-text-plugin-full-text-queries-for-annotated-text/232308/5 "2020-05-13T07:21:08Z")

</div>

In text1 the “ispresident” token is anchored to the same position as the first token it annotates- so in this case “Thomas” and not “Jefferson”. Adding a small slop factor to the query will help allow for this gap

---

<div class="post-metadata">

### Author: ![JohnathanBostrom](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/johnathanbostrom/32/68189_2.png) [@JohnathanBostrom](https://discuss.elastic.co/u/JohnathanBostrom)
#### Post date: [May 13, 2020, 4:37pm UTC](https://discuss.elastic.co/t/annotated-text-plugin-full-text-queries-for-annotated-text/232308/6 "2020-05-13T16:37:41Z")

</div>

Unfortunately a slop wouldn't work well for my purposes. My annotated text is up to 5 words long. A slop that large would cause many unwanted matches.  
I can think of two possible approaches around this issue. I would appreciate some feedback on the feasibility of these.

1. Tokenize the annotated text as a single token. if my understanding is right, if `"Thomas Jefferson"` were a keyword token, then it would occupy the same position as `ispresident` and my query would match. Is there any way to mark certain phrases as keywords during the tokenization process? In my case, all instances of `"Thomas Jefferson"` would be keywords and I know all possible keywords at the outset. If I could turn them into single tokens at index/query time it seems like it should solve my issue?

2. Create `isPresident` annotations of differing lengths. This idea is pretty hacky, but since I have a finite number of phrases that could be annotated with `isPresident`, I could create `"isPresident0"`, `"ispresident0 ispresident1"`, and `"ispresident0 ispresident1 ispresident2"` annotations to match texts of different lengths. I could then expand my query on that field to turn ispresident into all of those alternatives. How are annotations tokenized? Will multi word annotations take up multiple token positions?  
Thanks!

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [May 13, 2020, 4:59pm UTC](https://discuss.elastic.co/t/annotated-text-plugin-full-text-queries-for-annotated-text/232308/7 "2020-05-13T16:59:25Z")

</div>

> Tokenize the annotated text as a single token

That approach would preclude users searching for "Thomas Jefferson" or just "Jefferson" and matching that text.  
Unlike synonyms, annotations are not an index-wide policy definition attached to an analyzer. They are overlays on selected pieces of text so that not all Thomas Jeffersons have to be presidents.

> [@JohnathanBostrom](#):
>
> How are annotations tokenized?

You can use the `_analyze` api to see that. There's an example in [this blog](https://www.elastic.co/blog/search-for-things-not-strings-with-the-annotated-text-plugin)

> [@JohnathanBostrom](#):
>
> Will multi word annotations take up multiple token positions?

Nope. They're always anchored to the same position as the first token they annotate. Practically speaking we had to pick the annotation's position as either the first or last token and we chose the former. We could have positioned an instance of the annotation over both or maybe _every_ token in the covered text but that would have messed with term frequency (TF) scoring and any searches to find presidents near mentions of other presidents.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [June 10, 2020, 4:59pm UTC](https://discuss.elastic.co/t/annotated-text-plugin-full-text-queries-for-annotated-text/232308/8 "2020-06-10T16:59:39Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
