# Stored term vectors still slow when retrieving their scores (terms filtering)

**URL:** https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218
**Category:** Elasticsearch
**Created:** [January 14, 2016, 9:56am UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218 "2016-01-14T09:56:26Z")
**Posts on this page:** 14
**Page:** 1

<div class="post-metadata">

### Author: ![sandros](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sandros/32/6348_2.png) [@sandros](https://discuss.elastic.co/u/sandros)
#### Post date: [January 14, 2016, 9:56am UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/1 "2016-01-14T09:56:26Z")

</div>

Hi all,

I want to get the most characteristic words for each document I stored.  
The most straightforward way I tried this was (calling from python with es my ES-instance) looping over all my ids:  
`es.termvectors(index = INDEX_NAME,doc_type=TYPE_NAME,id=ii,field_statistics=True,fields =fields4TermVec,term_statistics=True,dfs= False,positions=False,offsets=False,body=bod_4TermVecs)`

where  
`fields4TermVec` contains 14 fields  
`bod_4TermVecs={ "filter" : { "max_num_terms" : 52 "min_term_freq" : 2, "min_doc_freq" : 1 }}`

this turns out to be not feasible, since way too slow.

So I thought that if I reindex into a new index with the mapping including for the 14 fields in `field4TermVec` having set `"term_vector": "yes"` I would get a substantial increase in computing speed. But this was not the case, anyone knows why?

The only reason I come up with is that it still needs to compute the doc\_freq and term\_freq and so storing the term vectors does not make a difference.

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [January 14, 2016, 10:42am UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/2 "2016-01-14T10:42:47Z")

</div>

> [@sandros](#):
>
> The only reason I come up with is that it still needs to compute the doc\_freq and term\_freq

Right - there's information that can be stored with a document that is static (e.g. the frequency of the term in the document) and there's stuff that is dynamic as the index changes such as the number of docs that contain a term. The latter requires a lot of look-ups to gather frequencies. Lookups=random disk seeks=slow.

---

<div class="post-metadata">

### Author: ![sandros](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sandros/32/6348_2.png) [@sandros](https://discuss.elastic.co/u/sandros)
#### Post date: [January 14, 2016, 10:44am UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/3 "2016-01-14T10:44:41Z")

</div>

Ok, so it is like I thought. Thanks a lot!

---

<div class="post-metadata">

### Author: ![Shai\_Erera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shai_erera/32/13492_2.png) [@Shai\_Erera](https://discuss.elastic.co/u/Shai_Erera)
#### Post date: [May 11, 2017, 11:58am UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/4 "2017-05-11T11:58:40Z")

</div>

@Mark_Harwood, you're right that there are lookups, but I think the implementation could be improved by storing per-request `Map<String,TermStatistics>` so that if you ask for TVs with filtering of many docs (MultiTVRequest), such local (transient) cache could be used to avoid looking up same terms over and over, as well potentially reducing the seeks that are performed per document and term. I ran into the same issue today ... 🙂

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [May 11, 2017, 1:55pm UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/5 "2017-05-11T13:55:24Z")

</div>

> [@Shai\_Erera](#):
>
> cache could be used to avoid looking up same terms over and over

We do this sort of caching when running `significant_terms` aggregations. Could you achieve the results you need using significant\_terms? I know that currently relies on memory-hungry fielddata but that's something [I'm working on](https://github.com/elastic/elasticsearch/issues/23674)

---

<div class="post-metadata">

### Author: ![Shai\_Erera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shai_erera/32/13492_2.png) [@Shai\_Erera](https://discuss.elastic.co/u/Shai_Erera)
#### Post date: [May 16, 2017, 1:30pm UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/6 "2017-05-16T13:30:19Z")

</div>

I read about `significant_terms` and I don't think it can help me. I need to fetch the TermVectors of multiple documents (say 100), and in order to reduce the returned payload size, I would like to return only the top-K tf-idf scoring terms (if docs have few 1000s of unique terms, and K=20,50,100, I expect to get a much smaller payload).

Perhaps I'm missing something about `significant_terms`, but it doesn't look like it addresses this requirement.

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [May 16, 2017, 1:59pm UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/7 "2017-05-16T13:59:54Z")

</div>

> [@Shai\_Erera](#):
>
> Perhaps I'm missing something about significant\_terms, but it doesn't look like it addresses this requirement.

Here's a query for example docs, referring to them directly by a unique ID and asking for significant terms:

```
GET signalmedia/_search
{
  "query": {
	"terms": {
	  "my_id": [
		"AVwMkdbCMeByIGkpzTjN",
		"AVwMk4PIMeByIGkp0VTd",
		"AVwMlg6pMeByIGkp1xGD",
		"AVwMlAc1MeByIGkp0opg",
		"AVwMlCLlMeByIGkp0s5T",
		"AVwMktkhMeByIGkpz7Hp",
		"AVwMldmJMeByIGkp1p7h",
		"AVwMkg4dMeByIGkpzcQh",
		"AVwMlCLlMeByIGkp0s67"
	  ]
	}
  },
  "size": 0,
  "aggs": {
	"keywords": {
	  "significant_terms": {
		"field": "content",
		"size": 20
	  }
	}
  }
}

```

Here's the results (they happen to be docs about elasticsearch):

```
{
  "took": 84,
  ...
  "aggregations": {
	"keywords": {
	  "doc_count": 9,
	  "buckets": [
		{
		  "key": "logstash",
		  "doc_count": 3,
		  "score": 27777.44444444444,
		  "bg_count": 4
		},
		{
		  "key": "elasticsearch",
		  "doc_count": 8,
		  "score": 22574.067019400354,
		  "bg_count": 35
		},
		{
		  "key": "kibana",
		  "doc_count": 3,
		  "score": 22221.888888888887,
		  "bg_count": 5
		}
	  ]
	}
  }
}

```

This requires `fielddata:true` and can be costly which is why I'm busy working on a significant\_text agg that retokenizes top matches on the fly from the stored source.

---

<div class="post-metadata">

### Author: ![Shai\_Erera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shai_erera/32/13492_2.png) [@Shai\_Erera](https://discuss.elastic.co/u/Shai_Erera)
#### Post date: [May 16, 2017, 2:10pm UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/8 "2017-05-16T14:10:35Z")

</div>

Thanks for the example, however this returns the top terms for **all** queried documents, and not the top ones per document (as it's an aggregation), which is what I need... any way to do that with that agg?

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [May 16, 2017, 2:22pm UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/9 "2017-05-16T14:22:00Z")

</div>

> [@Shai\_Erera](#):
>
> any way to do that with that agg?

No. If I understand your question correctly you see each doc as independent and the keywords you want for doc 1 (which might be about fish) is not influenced in any way by your other choice of docs which might be about something entirely different like bicycles or chocolate?  
i.e. you might as well make separate requests for each doc, were it not for the added network costs. If so, maybe try the "more like this" query, and set `size:1` and `explain:true`setting to see what the MLT logic picked out as the interesting keywords in the example doc.

---

<div class="post-metadata">

### Author: ![Shai\_Erera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shai_erera/32/13492_2.png) [@Shai\_Erera](https://discuss.elastic.co/u/Shai_Erera)
#### Post date: [May 16, 2017, 3:19pm UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/10 "2017-05-16T15:19:23Z")

</div>

This issue is about the slowness of when using terms filtering when retrieving the TVs of multiple documents. I think we agree that the implementation can be improved, right?

About your proposal, my current use case is this: I execute a query `Q` and retrieve `N` results. For each I would like to fetch the term vectors and do some post-processing at the client side. I use a multi-TV request, so I only have one additional round-trip to the server.

Due to network latency, fetching those TVs (think top 100 docs, each has few hundreds to thousands of unique terms) is slow (big response payload + network latency). So I thought to retrieve only the top-K terms of each document, in order to reduce the size of the TV response payload. However, due to the current implementation, the response time of terms filtering is actually much higher (and that's something I measured on my local laptop, i.e. no network latency...).

I will consider the MLT approach you mentioned, but I don't have an example document and I don't want to issue a request per document.

Do you see any reason not to improve the implementation in ES, when serving a multi-TV request with terms filtering?

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [May 16, 2017, 3:34pm UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/11 "2017-05-16T15:34:22Z")

</div>

> [@Shai\_Erera](#):
>
> This issue is about the slowness of when using terms filtering when retrieving the TVs of multiple documents. I think we agree that the implementation can be improved, right?

Ah yes. OK so even unrelated docs will share terms [and, of, the, if, when, I, you, .....] and you want to avoid looking those words up multiple times. Got it.

> [@Shai\_Erera](#):
>
> I execute a query Q and retrieve N results

OK so there _is_ a common theme to the docs - they all match the same query. If you replace the ids in my previous significant\_terms example with your choice of query (e.g. "bird flu") then analyzing those docs should spot "h5n1".  
You would need to do a follow-up query though for those top docs and a `terms` agg with an `include` clause listing the significant terms in order to discover which docs had those keywords.

The advantage of looking across the docs as a set rather than as individual hits is you can figure out that H5N1 is highly significant when it might be mentioned only once in a handful of docs (low TF).

> [@Shai\_Erera](#):
>
> Do you see any reason not to improve the implementation in ES, when serving a multi-TV request with terms filtering?

Significant\_text is intended to tackle this sort of thing and adds sequence de-duplication which I see as necessary for use on typical real-world text.

---

<div class="post-metadata">

### Author: ![Shai\_Erera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shai_erera/32/13492_2.png) [@Shai\_Erera](https://discuss.elastic.co/u/Shai_Erera)
#### Post date: [May 17, 2017, 7:03pm UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/12 "2017-05-17T19:03:38Z")

</div>

Thanks @Mark_Harwood. I intend to experiment with `significant_terms` more, as it looks interesting (it's more than the simple 'terms' aggregation that I thought it is before). So far, and without diving too deep into it, it's not that fast (4-5 seconds on my laptop, against a local index and 500K docs, one-word query), but I still need to experiment with it, so I don't mind the times too much yet. And I know you're working on improving it.

Parallel to that though, the TV I fetch for each result document is taken as its _profile_, and by taking only the top-K terms I consider them to be a _truncated-profile_ and that's still required by my application. Therefore I do wish the implementation of multi-TV with filters will be improved.

If I took a stab at it, do you think it's something that you would consider having in the code? I may not get to it right away, but if you think positively about this improvement, I'll try to allocate some time to it, as I do rely on that API.

---

<div class="post-metadata">

### Author: ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)
#### Post date: [May 24, 2017, 1:53pm UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/13 "2017-05-24T13:53:47Z")

</div>

> [@Shai\_Erera](#):
>
> and without diving too deep into it, it's not that fast

So sampling with the `sampler` (or `diversified_sampler`) aggregation is not only beneficial to performance but also results quality.  
Another key issue with most real-world text content is that of the various forms of content duplication that throw off statistical analysis. Check out the approach used in this new `significant_text` aggregation coming in 6.0 that deals with on-the-fly de-duplication on real-world examples: [https://www.youtube.com/watch?v=zH7bizwjj20](https://www.youtube.com/watch?v=zH7bizwjj20)

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 5, 2017, 10:00pm UTC](https://discuss.elastic.co/t/stored-term-vectors-still-slow-when-retrieving-their-scores-terms-filtering/39218/14 "2017-07-05T22:00:07Z")

</div>


