# What does “docCount” and "docFreq" mean in the Explain API?

**URL:** <https://discuss.elastic.co/t/what-does-doccount-and-docfreq-mean-in-the-explain-api/163850>\
**Category:** Elasticsearch\
**Created:** [January 11, 2019, 6:42am UTC](https://discuss.elastic.co/t/what-does-doccount-and-docfreq-mean-in-the-explain-api/163850 "2019-01-11T06:42:44Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Masanori\_Ohnishi](https://avatars.discourse-cdn.com/v4/letter/m/278dde/32.png) [@Masanori\_Ohnishi](https://discuss.elastic.co/u/Masanori_Ohnishi)\
**Post date:** [January 11, 2019, 6:42am UTC](https://discuss.elastic.co/t/what-does-doccount-and-docfreq-mean-in-the-explain-api/163850/1 "2019-01-11T06:42:44Z")

</div>

Here's the sample of mapping, register, and search query.

## mapping

```auto
curl -X PUT "es:9200/english1" -H 'Content-Type: application/json' -d'
{
  "mappings": {
    "_doc": {
      "properties": {
        "header" : {
          "type" : "text"
        },
        "body" : {
          "type" : "text"
        }
      }
    }
  }
}
'

```

## register

```auto
curl -X PUT "es:9200/english1/_doc/1?refresh" -H 'Content-Type: application/json' -d'
{
  "header": "something special",
  "body": "I am John"
}
'

curl -X PUT "es:9200/english1/_doc/2?refresh" -H 'Content-Type: application/json' -d'
{
  "header": "something better",
  "body": "You are Chris"
}
'

curl -X PUT "es:9200/english1/_doc/3?refresh" -H 'Content-Type: application/json' -d'
{
  "header": "anything hot",
  "body": "This is a cup"
}
'

curl -X PUT "es:9200/english1/_doc/4?refresh" -H 'Content-Type: application/json' -d'
{
  "header": "anything cold",
  "body": "That is a glass"
}
'

```

## search

```auto
curl -XGET 'es:9200/english1/_search?pretty' -H 'Content-Type: application/json' -d'
{
   "query" : {
        "simple_query_string":{
        "query": "something",
        "fields": ["header","body"]
      }
    },
    "explain": true
}'

```

## result

```auto
{
  "took" : 5,
  "timed_out" : false,
  "_shards" : {
    "total" : 5,
    "successful" : 5,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : 2,
    "max_score" : 0.6931472,
    "hits" : [
      {
        "_shard" : "[english1][2]",
        "_node" : "sN3QHj7oRF-rgbBbs4U6lw",
        "_index" : "english1",
        "_type" : "_doc",
        "_id" : "2",
        "_score" : 0.6931472,
        "_source" : {
          "header" : "something better",
          "body" : "You are Chris"
        },
        "_explanation" : {
          "value" : 0.6931472,
          "description" : "sum of:",
          "details" : [
            {
              "value" : 0.6931472,
              "description" : "weight(header:something in 0) [PerFieldSimilarity], result of:",
              "details" : [
                {
                  "value" : 0.6931472,
                  "description" : "score(doc=0,freq=1.0 = termFreq=1.0\n), product of:",
                  "details" : [
                    {
                      "value" : 0.6931472,
                      "description" : "idf, computed as log(1 + (docCount - docFreq + 0.5) / (docFreq + 0.5)) from:",
                      "details" : [
                        {
                          "value" : 1.0,
                          "description" : "docFreq",
                          "details" : []
                        },
                        {
                          "value" : 2.0,
                          "description" : "docCount",
                          "details" : []
                        }
                      ]
                    },
                    {
                      "value" : 1.0,
                      "description" : "tfNorm, computed as (freq * (k1 + 1)) / (freq + k1 * (1 - b + b * fieldLength / avgFieldLength)) from:",
                      "details" : [
                        {
                          "value" : 1.0,
                          "description" : "termFreq=1.0",
                          "details" : []
                        },
                        {
                          "value" : 1.2,
                          "description" : "parameter k1",
                          "details" : []
                        },
                        {
                          "value" : 0.75,
                          "description" : "parameter b",
                          "details" : []
                        },
                        {
                          "value" : 2.0,
                          "description" : "avgFieldLength",
                          "details" : []
                        },
                        {
                          "value" : 2.0,
                          "description" : "fieldLength",
                          "details" : []
                        }
...
      },
      {
        "_shard" : "[english1][3]",
        "_node" : "sN3QHj7oRF-rgbBbs4U6lw",
        "_index" : "english1",
        "_type" : "_doc",
        "_id" : "1",
        "_score" : 0.2876821,
        "_source" : {
          "header" : "something special",
          "body" : "I am John"
        },
        "_explanation" : {
          "value" : 0.2876821,
          "description" : "sum of:",
          "details" : [
            {
              "value" : 0.2876821,
              "description" : "weight(header:something in 0) [PerFieldSimilarity], result of:",
              "details" : [
                {
                  "value" : 0.2876821,
                  "description" : "score(doc=0,freq=1.0 = termFreq=1.0\n), product of:",
                  "details" : [
                    {
                      "value" : 0.2876821,
                      "description" : "idf, computed as log(1 + (docCount - docFreq + 0.5) / (docFreq + 0.5)) from:",
                      "details" : [
                        {
                          "value" : 1.0,
                          "description" : "docFreq",
                          "details" : []
                        },
                        {
                          "value" : 1.0,
                          "description" : "docCount",
                          "details" : []
                        }
                      ]
                    },
                    {
                      "value" : 1.0,
                      "description" : "tfNorm, computed as (freq * (k1 + 1)) / (freq + k1 * (1 - b + b * fieldLength / avgFieldLength)) from:",
                      "details" : [
                        {
                          "value" : 1.0,
                          "description" : "termFreq=1.0",
                          "details" : []
                        },
                        {
                          "value" : 1.2,
                          "description" : "parameter k1",
                          "details" : []
                        },
                        {
                          "value" : 0.75,
                          "description" : "parameter b",
                          "details" : []
                        },
                        {
                          "value" : 2.0,
                          "description" : "avgFieldLength",
                          "details" : []
                        },
                        {
                          "value" : 2.0,
                          "description" : "fieldLength",
                          "details" : []
                        }
...
}

```

According to this Q&A[[Understanding doc and docCount values in explain response](https://discuss.elastic.co/t/understanding-doc-and-doccount-values-in-explain-response/158383/2)],  
if `docCount` the number of docs in my index, I assume `docCount` will be 4(but it was wrong in this response.).

Also I cant't understand `docFreq`.

Please help me...

---

<div class="post-metadata">

**Author:** ![s1monw](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/s1monw/32/3637_2.png) [@s1monw](https://discuss.elastic.co/u/s1monw)\
**Post date:** [January 11, 2019, 7:03am UTC](https://discuss.elastic.co/t/what-does-doccount-and-docfreq-mean-in-the-explain-api/163850/2 "2019-01-11T07:03:21Z")

</div>

All these statistics are per shard not per index. Use a single shard instead of 5 then your stats will be accurate.

---

<div class="post-metadata">

**Author:** ![Masanori\_Ohnishi](https://avatars.discourse-cdn.com/v4/letter/m/278dde/32.png) [@Masanori\_Ohnishi](https://discuss.elastic.co/u/Masanori_Ohnishi)\
**Post date:** [January 11, 2019, 10:52am UTC](https://discuss.elastic.co/t/what-does-doccount-and-docfreq-mean-in-the-explain-api/163850/3 "2019-01-11T10:52:46Z")

</div>

Thank you very much!  
The solution works!!

couple of follow up question,

- the limit of size of a single shard seems to be 50GB according to the blog.([https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster))  
If the data size exeeds 50GB, how I can increase shard num?
- And Initially, is it common to use a single shard in making products?

Forgive me for stealing you time again....😥

---

<div class="post-metadata">

**Author:** ![s1monw](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/s1monw/32/3637_2.png) [@s1monw](https://discuss.elastic.co/u/s1monw)\
**Post date:** [January 11, 2019, 12:11pm UTC](https://discuss.elastic.co/t/what-does-doccount-and-docfreq-mean-in-the-explain-api/163850/4 "2019-01-11T12:11:37Z")

</div>

> [@Masanori\_Ohnishi](#):
>
> Forgive me for stealing you time again....😥

you are not stealing anybodies time.

> [@Masanori\_Ohnishi](#):
>
> the limit of size of a single shard seems to be 50GB according to the blog.

There are some limits but size is not the limit. Yet, that said massive single shards will at some point not give you the performance you need. You can use more than one shard and you should if you have enough data. The number of shards is determined at index creation time but you can still use the `_split` [API](https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-split-index.html) if you wanna use more shards.

> [@Masanori\_Ohnishi](#):
>
> - And Initially, is it common to use a single shard in making products?

it depends on how much data you have, yet it's not uncommon and a good place to start.

I take from your original question that you need the statistics to be accurate? Why is this the case? Can you explain why you need a accurate `docCount` values?

---

<div class="post-metadata">

**Author:** ![s1monw](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/s1monw/32/3637_2.png) [@s1monw](https://discuss.elastic.co/u/s1monw)\
**Post date:** [January 11, 2019, 12:15pm UTC](https://discuss.elastic.co/t/what-does-doccount-and-docfreq-mean-in-the-explain-api/163850/5 "2019-01-11T12:15:11Z")

</div>

> [@Masanori\_Ohnishi](#):
>
> Forgive me for stealing you time again....😥

you are not stealing anybodies time.

> [@Masanori\_Ohnishi](#):
>
> the limit of size of a single shard seems to be 50GB according to the blog.

There are some limits but size is not the limit. Yet, that said massive single shards will at some point not give you the performance you need. You can use more than one shard and you should if you have enough data. The number of shards is determined at index creation time but you can still use the `_split` [API](https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-split-index.html) if you wanna use more shards.

> [@Masanori\_Ohnishi](#):
>
> - And Initially, is it common to use a single shard in making products?

it depends on how much data you have, yet it's not uncommon and a good place to start.

I take from your original question that you need the statistics to be accurate? Why is this the case? Can you explain why you need a accurate `docCount` values?

---

<div class="post-metadata">

**Author:** ![s1monw](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/s1monw/32/3637_2.png) [@s1monw](https://discuss.elastic.co/u/s1monw)\
**Post date:** [January 11, 2019, 12:19pm UTC](https://discuss.elastic.co/t/what-does-doccount-and-docfreq-mean-in-the-explain-api/163850/6 "2019-01-11T12:19:04Z")

</div>

> [@Masanori\_Ohnishi](#):
>
> Forgive me for stealing you time again....😥

you are not stealing anybodies time.

> [@Masanori\_Ohnishi](#):
>
> the limit of size of a single shard seems to be 50GB according to the blog.

There are some limits but size is not the limit. Yet, that said massive single shards will at some point not give you the performance you need. You can use more than one shard and you should if you have enough data. The number of shards is determined at index creation time but you can still use the `_split` [API](https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-split-index.html) if you wanna use more shards.

> [@Masanori\_Ohnishi](#):
>
> - And Initially, is it common to use a single shard in making products?

it depends on how much data you have, yet it's not uncommon and a good place to start.

I take from your original question that you need the statistics to be accurate? Why is this the case? Can you explain why you need accurate `docCount` values?

---

<div class="post-metadata">

**Author:** ![Masanori\_Ohnishi](https://avatars.discourse-cdn.com/v4/letter/m/278dde/32.png) [@Masanori\_Ohnishi](https://discuss.elastic.co/u/Masanori_Ohnishi)\
**Post date:** [January 15, 2019, 5:15am UTC](https://discuss.elastic.co/t/what-does-doccount-and-docfreq-mean-in-the-explain-api/163850/7 "2019-01-15T05:15:19Z")

</div>

Thanks a lot !!  
(Sorry for replying late...)

> I take from your original question that you need the statistics to be accurate? Why is this the case? Can you explain why you need accurate `docCount` values?

Currently I deal with a small amount of data, so I want to need the statistics to be accurate.  
However, since data will increase in the future, I wanted to know how to cope when the data increased.

In conclusion, I will use a single shard, and use split api if necessary.

---

<div class="post-metadata">

**Author:** ![s1monw](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/s1monw/32/3637_2.png) [@s1monw](https://discuss.elastic.co/u/s1monw)\
**Post date:** [January 15, 2019, 8:37am UTC](https://discuss.elastic.co/t/what-does-doccount-and-docfreq-mean-in-the-explain-api/163850/8 "2019-01-15T08:37:23Z")

</div>

> [@Masanori\_Ohnishi](#):
>
> In conclusion, I will use a single shard, and use split api if necessary.

I agree that is a good solution. Once you are beyond one shard you can still use the [DFS query then fetch](https://www.elastic.co/blog/understanding-query-then-fetch-vs-dfs-query-then-fetch) search type to get accurate stats. That requires an additional roundtrip and might be overkill. I recommend reading the linked article.

simon

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 12, 2019, 8:37am UTC](https://discuss.elastic.co/t/what-does-doccount-and-docfreq-mean-in-the-explain-api/163850/9 "2019-02-12T08:37:27Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
