# Populating TermVector when having tokenizer outside ES

**URL:** https://discuss.elastic.co/t/populating-termvector-when-having-tokenizer-outside-es/17258
**Category:** Elasticsearch
**Created:** [April 29, 2014, 12:05pm UTC](https://discuss.elastic.co/t/populating-termvector-when-having-tokenizer-outside-es/17258 "2014-04-29T12:05:10Z")
**Posts on this page:** 2
**Page:** 1

<div class="post-metadata">

### Author: ![Neeraj\_Makam](https://avatars.discourse-cdn.com/v4/letter/n/94ad74/32.png) [@Neeraj\_Makam](https://discuss.elastic.co/u/Neeraj_Makam)
#### Post date: [April 29, 2014, 12:05pm UTC](https://discuss.elastic.co/t/populating-termvector-when-having-tokenizer-outside-es/17258/1 "2014-04-29T12:05:10Z")

</div>

Hi,

I have a mapping in which there is a nested list of words (which is  
generated by a tokenizer residing outside ES). Each word has fields  
'token\_offset' and 'character\_offset' which is populated by my tokenizer.  
This is the mapping i am using (say):

{  
"contract": {  
"\_id" : {  
"path" : "objectId"  
},  
"properties": {  
"filepath": {  
"type": "string",  
"index": "not\_analyzed"  
},  
"objectId": {  
"type": "string",  
"index": "no"  
},  
"_words_": {  
"type": "_nested_",  
"properties": {  
"_characterOffset_": {  
"type": "long",  
"index": "no"  
},  
"_wordType_": {  
"type": "long"  
},  
"_tokenOffset_": {  
"type": "long",  
"index": "no"  
},  
"_value_": {  
"type": "string",  
"index": "not\_analyzed"  
}  
}  
}  
}  
}  
}

I want to be able to do a query which says:  
"value" == "foo" AND "wordType" == 5.  
This made me map the list "words" as nested. [1]

For eg:  
if the text is "_this is foo and bar_", my tokenizer separates out each  
word and associates wordType for each word, and also generates  
characterOffset and tokenOffset.  
i.e  
word[0].value = "this"  
word[0].wordType = 5  
word[0].characterOffset = 0  
word[0].tokenOffset = 0

Now how do i populate the termvector of ES so as to leverage its phrase  
search and other features such as "AND/OR/NEAR" etc?? [2]

_[1] - Is there a way i can implement this without using the concept of  
nested (because this will separate out each word into a separate document)_  
_[2] - Can a custom analyzer be used to populate the term vector of ES  
while having the tokenizer outside ES (assuming due to business necessity,  
moving the tokenizer inside ES is not feasible)._  
//The feature i need to implement is _phrase search, supporting AND/OR/NEAR  
and highlighting._

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/7c363b2d-3dc0-4096-8d47-ab70ee20d181%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/7c363b2d-3dc0-4096-8d47-ab70ee20d181%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 1:32am UTC](https://discuss.elastic.co/t/populating-termvector-when-having-tokenizer-outside-es/17258/2 "2017-07-06T01:32:47Z")

</div>


