# Data Type to specify exact tokens to index

**URL:** https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637
**Category:** Elasticsearch
**Created:** [August 18, 2019, 8:44am UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637 "2019-08-18T08:44:00Z")
**Posts on this page:** 15
**Page:** 1

<div class="post-metadata">

### Author: ![robmartin11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robmartin11/32/99177_2.png) [@robmartin11](https://discuss.elastic.co/u/robmartin11)
#### Post date: [August 18, 2019, 8:44am UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/1 "2019-08-18T08:44:00Z")

</div>

Is there any way to index content by specifying the exact tokens to index?  
For example:

```auto
PUT /index/type/1
{
	“content”: “Some content to index”,
	“tokens”: [
		{
			“value”: “data”
			“start”: 5,
			“end”: 11
		}
	]
}

```

I know the annotated text plugin does something similar, but its unsuitable as we already know the tokens we want to index and some tokens may overlap.  
Would I need to create a custom plugin?

Thanks in advance……

---

<div class="post-metadata">

### Author: ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)
#### Post date: [August 24, 2019, 11:59pm UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/2 "2019-08-24T23:59:12Z")

</div>

Currently there is not a special field datatype for injecting your own tokens. You would need to develop your own plugin.  
I have filed an [issue](https://github.com/elastic/elasticsearch/issues/45942) to discuss a possibility of implementing it.

---

<div class="post-metadata">

### Author: ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)
#### Post date: [August 26, 2019, 8:58pm UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/3 "2019-08-26T20:58:07Z")

</div>

@robmartin11 We would like to know more about your use case.

- First, why you don't you do analysis on the server side and take advantage of the existing analyzers?
- Second, provided that we implement a new field type that allows to inject tokens, how would you query it? Would you use `Keyword` analyzer?

---

<div class="post-metadata">

### Author: ![robmartin11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robmartin11/32/99177_2.png) [@robmartin11](https://discuss.elastic.co/u/robmartin11)
#### Post date: [August 30, 2019, 7:33am UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/4 "2019-08-30T07:33:59Z")

</div>

So the use case is that we have our own Named Entity Recognition engine which has already found the tokens in the text. This includes tokens that overlap to represent entities with overlapping words. I expect it would work very much like the annotated text plugin, but rather than markup the content with the entities (which doesn't allow for overlapping entities, only multiple entities that have exactly the same start/end index), it would return the set of tokens in a separate property as a list. I believe with the annotated text plugin you can still provide your own analyzer

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [August 30, 2019, 8:54am UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/5 "2019-08-30T08:54:14Z")

</div>

> [@robmartin11](#):
>
> I believe with the annotated text plugin you can still provide your own analyzer

Yes, because injected annotations have to occupy a _position_ in the indexed information dictated by a choice of tokenizer - this is effectively measured as word-number rather than character offset. So the word "foo" is the fourth word in this sentence. The concept of proximity when running positional queries like phrase queries and span queries in Lucene is measured using the token's `position` info - not by comparing character offsets for proximity. Character offsets are only used for things like highlighting. It can be useful to allow positional queries that mix structured annotations and free-text in proximity queries - [demo](https://youtu.be/_ThbDlhp1AQ?t=519)

Annotations using the annotated text plugin can include more than one token to inject - these can be encoded and separated by & characters. Maybe you could put the annotation braces `[]` around the longest span of any overlapping entities and use the ampersand to separate the two tokens eg `(shortEntity&longEntity)`?

Can you share any examples of your overlapping annotations?

The alternative to using annotated text to inject tokens would be to write your own Analyzer plugin in Java and take full control of converting the text into indexed elements.

---

<div class="post-metadata">

### Author: ![robmartin11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robmartin11/32/99177_2.png) [@robmartin11](https://discuss.elastic.co/u/robmartin11)
#### Post date: [August 30, 2019, 6:57pm UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/6 "2019-08-30T18:57:00Z")

</div>

Thanks for the clarification.  
I suppose what we are after is something like the annotated text plugin but with the highlighting correct for overlapping entities.  
Does the annotated text plugin suffer from lucenes sausagization issues described in [http://blog.mikemccandless.com/2012/04/lucenes-tokenstreams-are-actually.html](http://blog.mikemccandless.com/2012/04/lucenes-tokenstreams-are-actually.html)

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [August 30, 2019, 9:13pm UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/7 "2019-08-30T21:13:33Z")

</div>

The annotations themselves are not tokenised so even if they contain whitespace they are considered as a single token.  
The position and length of each of these is dependent on the host Analyzer’s policy for tokenising the text they decorate. [Here](https://github.com/elastic/elasticsearch/blob/master/plugins/mapper-annotated-text/src/main/java/org/elasticsearch/index/mapper/annotatedtext/AnnotatedTextFieldMapper.java#L464) is the code for injecting the annotations with the position information from the wrapped Analyzer

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [September 2, 2019, 4:49pm UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/8 "2019-09-02T16:49:06Z")

</div>

I put together a multi-value annotation example [here](https://gist.github.com/markharwood/881a91f3b25a352fe3578c7f44ccde7a)

---

<div class="post-metadata">

### Author: ![robmartin11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robmartin11/32/99177_2.png) [@robmartin11](https://discuss.elastic.co/u/robmartin11)
#### Post date: [September 4, 2019, 11:43am UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/9 "2019-09-04T11:43:02Z")

</div>

Thanks for the example.  
We have also noticed that nesting seems to work correctly, for example: `[President of the [United States](USA)](President)`. This wasn't mentioned in the documentation so just wanted to check if this was done by design and if so if multiple levels of nesting are supported.

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [September 4, 2019, 7:18pm UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/10 "2019-09-04T19:18:03Z")

</div>

> [@robmartin11](#):
>
> nesting seems to work correctly,

No, I don't think it does and wasn't designed to either. Note only the inner `USA` annotation is parsed. You can see the outputs using the "analyze" api:

```
DELETE test
PUT test
{
  "settings": {
	"number_of_shards": 1
  },
  "mappings": {
	  "properties":{
		"text":{
		  "type":"annotated_text"
		}
	  }
  }
}
GET test/_analyze
{
  "field": "text",
  "text":"[President of the [United States](USA)](President)"
}

```

Note the outer `President` annotation is lower cased because it is considered to be text.

---

<div class="post-metadata">

### Author: ![robmartin11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robmartin11/32/99177_2.png) [@robmartin11](https://discuss.elastic.co/u/robmartin11)
#### Post date: [September 11, 2019, 1:51pm UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/11 "2019-09-11T13:51:38Z")

</div>

One solution we are looking at is marking up each word separately so we can support nested entities, for example if we have the entity 'leg' (ANAT:456) within 'short leg syndrome' (IND:123): `[short](IND:123) [leg](IND:123&ANAT:456) [syndrome](IND:123)` I understand this may have some downsides, for example skewing relevancy but is there any other major issues which this may cause with the plugin?

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [September 11, 2019, 2:23pm UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/12 "2019-09-11T14:23:01Z")

</div>

I don't envisage any issues with the plugin.  
I tried your example and the highlighting seems to work OK

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [September 23, 2019, 3:28pm UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/13 "2019-09-23T15:28:06Z")

</div>

Is this approach working for you?

---

<div class="post-metadata">

### Author: ![robmartin11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robmartin11/32/99177_2.png) [@robmartin11](https://discuss.elastic.co/u/robmartin11)
#### Post date: [September 25, 2019, 8:57am UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/14 "2019-09-25T08:57:36Z")

</div>

We have started prototyping the next generation of our document search UI which relies heavily on NER. We have built is on top of ES + annotated text plugin using the word per token strategy. Things are looking good so far and it seems to deal with overlapping and nested word wells. Still need to see how it effects relevancy but I think in our case this is unlikely to be an issue

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [October 23, 2019, 8:57am UTC](https://discuss.elastic.co/t/data-type-to-specify-exact-tokens-to-index/195637/15 "2019-10-23T08:57:36Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
