# Term vector of all documents

**URL:** <https://discuss.elastic.co/t/term-vector-of-all-documents/60128>\
**Category:** Elasticsearch\
**Created:** [September 9, 2016, 2:25am UTC](https://discuss.elastic.co/t/term-vector-of-all-documents/60128 "2016-09-09T02:25:26Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![tyka](https://avatars.discourse-cdn.com/v4/letter/t/49beb7/32.png) [@tyka](https://discuss.elastic.co/u/tyka)\
**Post date:** [September 9, 2016, 2:25am UTC](https://discuss.elastic.co/t/term-vector-of-all-documents/60128/1 "2016-09-09T02:25:27Z")

</div>

I have a corpus of documents indexed . I also stored the term vectors when indexing. Now I want to retrieve term vectors of all documents satisfying some filtering options.  
I was able to get term vector for a single document or for a set of documents by providing the document IDs. But is there a way to get term vectors for all the documents without providing document IDs?  
Eventually what I want to do is to get the frequency counts of all the terms in a field, for all documents in an index (i.e., a bag of words matrix).

I am using elasticsearch-py as a client.

Appreciate any pointers. Thanks!

---

<div class="post-metadata">

**Author:** ![tyka](https://avatars.discourse-cdn.com/v4/letter/t/49beb7/32.png) [@tyka](https://discuss.elastic.co/u/tyka)\
**Post date:** [September 9, 2016, 6:24pm UTC](https://discuss.elastic.co/t/term-vector-of-all-documents/60128/2 "2016-09-09T18:24:21Z")

</div>

For this task, is there a way to aggregate on termvectors for a field?

---

<div class="post-metadata">

**Author:** ![javanna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javanna/32/4698_2.png) [@javanna](https://discuss.elastic.co/u/javanna)\
**Post date:** [September 12, 2016, 8:50am UTC](https://discuss.elastic.co/t/term-vector-of-all-documents/60128/3 "2016-09-12T08:50:50Z")

</div>

There is no way to aggregate on term\_vectors. The only way to retrieve term\_vectors that I'm aware of is per document id, the way to retrieve them for all documents matching a query would be to run a search scroll and retrieve term\_vectors for each document returned by id. Actually there's also the multi term vector api that allows to retrieve term\_vectors for multiple documents at the same time which is a better fit, so you could batch them.

---

<div class="post-metadata">

**Author:** ![tyka](https://avatars.discourse-cdn.com/v4/letter/t/49beb7/32.png) [@tyka](https://discuss.elastic.co/u/tyka)\
**Post date:** [September 12, 2016, 6:06pm UTC](https://discuss.elastic.co/t/term-vector-of-all-documents/60128/4 "2016-09-12T18:06:42Z")

</div>

Thanks!  
Yes, I tried the multi-termvector approach, but still I have to provide the list of document IDs which is huge, in the order of hundreds of millions.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:21pm UTC](https://discuss.elastic.co/t/term-vector-of-all-documents/60128/5 "2017-07-05T22:21:00Z")

</div>


