# Elasticsearch: total term frequency and doc count from given set of documents

**URL:** <https://discuss.elastic.co/t/elasticsearch-total-term-frequency-and-doc-count-from-given-set-of-documents/115223>\
**Category:** Elasticsearch\
**Created:** [January 12, 2018, 6:17am UTC](https://discuss.elastic.co/t/elasticsearch-total-term-frequency-and-doc-count-from-given-set-of-documents/115223 "2018-01-12T06:17:24Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![prasoonkirar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/prasoonkirar/32/26462_2.png) [@prasoonkirar](https://discuss.elastic.co/u/prasoonkirar)\
**Post date:** [January 12, 2018, 6:17am UTC](https://discuss.elastic.co/t/elasticsearch-total-term-frequency-and-doc-count-from-given-set-of-documents/115223/1 "2018-01-12T06:17:24Z")

</div>

I am trying to get total term frequency and document count from given set of documents, but \_termvectors in elasticsearch returns ttf and doc\_count from all documents within the index. Is there any way so that I can specify list of documents (document ids) so that result will based on those documents only.

Below are documents details and query to get total term frequency:  
**Index details:**

```
PUT /twitter
{ "mappings": {
    "tweets": {
      "properties": {
	"name": {
	  "type": "text",
	  "analyzer":"english"
	}
      }
    }
  },
  "settings" : {
    "index" : {
      "number_of_shards" : 1,
      "number_of_replicas" : 0
    }
  }
}

```

**Document Details:**

```
PUT /twitter/tweets/1
{
  "name":"Hello bar"
}

PUT /twitter/tweets/2
{
  "name":"Hello foo"
}

PUT /twitter/tweets/3
{
  "name":"Hello foo bar"
}

```

It will create three document with ids 1, 2 and 3. Now suppose tweets with ids 1 and 2 belongs to user1 and 3 belong to another user and I want to get the termvectors for user1.

**Query to get this result:**

```
GET /twitter/tweets/_mtermvectors
{
  "ids" : ["1", "2"],
  "parameters": {
      "fields": ["name"],
      "term_statistics": true,
      "offsets":false,
      "payloads":false,
      "positions":false
  }
}

```

**Response:**

```
{
  "docs": [
    {
      "_index": "twitter",
      "_type": "tweets",
      "_id": "1",
      "_version": 1,
      "found": true,
      "took": 1,
      "term_vectors": {
        "name": {
          "field_statistics": {
            "sum_doc_freq": 7,
            "doc_count": 3,
            "sum_ttf": 7
          },
          "terms": {
            "bar": {
              "doc_freq": 2,
              "ttf": 2,
              "term_freq": 1
            },
            "hello": {
              "doc_freq": 3,
              "ttf": 3,
              "term_freq": 1
            }
          }
        }
      }
    },
    {
      "_index": "twitter",
      "_type": "tweets",
      "_id": "2",
      "_version": 1,
      "found": true,
      "took": 1,
      "term_vectors": {
        "name": {
          "field_statistics": {
            "sum_doc_freq": 7,
            "doc_count": 3,
            "sum_ttf": 7
          },
          "terms": {
            "foo": {
              "doc_freq": 2,
              "ttf": 2,
              "term_freq": 1
            },
            "hello": {
              "doc_freq": 3,
              "ttf": 3,
              "term_freq": 1
            }
          }
        }
      }
    }
  ]
}

```

Here we can see `hello` is having doc\_count 3 and ttf 3. How can I make it to consider only documents with given ids.

One approach I am thinking is to create different index for different users. But I am not sure if this approach is correct. With this approach indices will increase with users. Or can there be another solution?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 12, 2018, 7:28am UTC](https://discuss.elastic.co/t/elasticsearch-total-term-frequency-and-doc-count-from-given-set-of-documents/115223/2 "2018-01-12T07:28:09Z")

</div>

Please format your code using `</>` icon as explained in [this guide](https://discuss.elastic.co/t/about-the-elasticsearch-category/21). It will make your post more readable.

Or use markdown style like:

````
```
CODE
```
````

---

<div class="post-metadata">

**Author:** ![prasoonkirar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/prasoonkirar/32/26462_2.png) [@prasoonkirar](https://discuss.elastic.co/u/prasoonkirar)\
**Post date:** [January 12, 2018, 7:39am UTC](https://discuss.elastic.co/t/elasticsearch-total-term-frequency-and-doc-count-from-given-set-of-documents/115223/3 "2018-01-12T07:39:35Z")

</div>

Thanks for suggestion, I have updated formatting accordingly.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 12, 2018, 8:00am UTC](https://discuss.elastic.co/t/elasticsearch-total-term-frequency-and-doc-count-from-given-set-of-documents/115223/4 "2018-01-12T08:00:22Z")

</div>

Awesome. Thanks!

I don't have the answer to your question. Indeed creating specific indices would work. If it's a one shot or small operation, you could imagine something like:

- calling reindex API using the query you wish and reindex few docs to `tmp_timestamp` index for example
- call the termvectors API
- drop the `tmp_timestamp` index

But may be @jimczi has a much better idea?

---

<div class="post-metadata">

**Author:** ![jimczi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jimczi/32/47985_2.png) [@jimczi](https://discuss.elastic.co/u/jimczi)\
**Post date:** [January 12, 2018, 8:36am UTC](https://discuss.elastic.co/t/elasticsearch-total-term-frequency-and-doc-count-from-given-set-of-documents/115223/5 "2018-01-12T08:36:48Z")

</div>

Since you have the `term_freq` per term per document in the response, it should be straightforward to derive the total term frequency for each term (just sum up the `term_freq` of each document/term) and the doc count is just the number of documents in the response that contain the term.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 9, 2018, 8:37am UTC](https://discuss.elastic.co/t/elasticsearch-total-term-frequency-and-doc-count-from-given-set-of-documents/115223/6 "2018-02-09T08:37:04Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
