# Terms Aggregation not returning keys

**URL:** <https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634>\
**Category:** Elasticsearch\
**Created:** [January 29, 2021, 11:26am UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634 "2021-01-29T11:26:03Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![gabe.ks11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gabe.ks11/32/83109_2.png) [@gabe.ks11](https://discuss.elastic.co/u/gabe.ks11)\
**Post date:** [January 29, 2021, 11:26am UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/1 "2021-01-29T11:26:03Z")

</div>

Hi everyone,

So recently we ran into a problem using elastic, we are attempting to detect records which have a duplicate value (hash) and to patch them with a flag so that we iterate all records as we go. (27 million records in total, out of which 6 million populate with said hash).

This workflow worked fine on an index sitting on a single node. Eventually the index grew and we moved it to multiple nodes so that we retain performance.

After the move the aggregation result was not accurate anymore as some of the records which we know for sure have keys with more than 1 doc count, do not get returned. I tried to run the query with a min\_doc\_count of 1 and again the aggregation does not return some of the values (hashes) which should be returned.

In the query below if we remove the must\_not and add a different condition which should return the missing hashes will not work either. If we add a condition defining an explicit equal: key = value (hash) which we know is missing then the aggregation will return the correct result and count, otherwise the key is missing all together.

I am hoping that there are a few of few who can explain what is going on and if there is something that we might be doing wrong or if there is a way to rectify the situation.

```
 {
  "query": {
    "bool": {
      "must_not": [
        {
          "bool": {
            "filter": [
              {
                "exists": {
                  "field": "$type",
                  "boost": 1
                }
              },
              {
                "term": {
                  "kmeta:Misc": {
                    "value": "KBXD-R-1611392400028",
                    "boost": 1
                  }
                }
              }
            ],
            "adjust_pure_negative": true,
            "boost": 1
          }
        }
      ],
      "adjust_pure_negative": true,
      "boost": 1
    }
  }
  "aggregations": {
    "kmeta:fileHash": {
      "terms": {
        "field": "kmeta:fileHash",
        "size": 10000,
        "shard_size": 10000,
        "min_doc_count": 2,
        "shard_min_doc_count": 0
      }
    }
  }
}
```

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [January 29, 2021, 4:55pm UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/2 "2021-01-29T16:55:59Z")

</div>

The problem is that randomly-sharded data is suited to finding the most popular things only when the frequency of those top things are \> number of shards.

Assuming there's millions of hashes that match your query and they are spread somewhat randomly across shards how would multiple remote shards independently decide on the same subset of 10k (shard\_size) terms that would guarantee finding the duplicates? Each of the millions of hashes on each shard occur only once so they're all equally promising candidates in this isolated view.  
When the required global frequency is low (2 in your example), and there are millions of values to choose from then you have to get each shard to focus analysis on the same subset of all the matching terms in a request e.g. just the hashes beginning with "a".  
A more effective way to do this is the [term partitioning](https://www.elastic.co/guide/en/elasticsearch/reference/7.10/search-aggregations-bucket-terms-aggregation.html#_filtering_values_with_partitions) feature in the terms agg. It means you have to make multiple requests, one for each partition, but the results can be made accurate.

Another alternative is to reindex using routing to send docs with the same hash to the same shard.

---

<div class="post-metadata">

**Author:** ![gabe.ks11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gabe.ks11/32/83109_2.png) [@gabe.ks11](https://discuss.elastic.co/u/gabe.ks11)\
**Post date:** [January 29, 2021, 5:55pm UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/3 "2021-01-29T17:55:54Z")

</div>

Hi Mark,

Thank you very much for your answer, the query constraint is what we patch as we iterate in a loop, so that records returned are excluded for the next page. But as I mentioned patching all records found in the terms aggregation with no min doc count constraint will still not go through some of the hashes (I do not even care about the accuracy of the count at this time) . I also tried what you suggested with splitting the aggregation in parts, the missing hashes are still not returned in the set. The only way I was able to get the hash to be part of the aggregation was to add an equal constraint on it.

Do you have any other suggestions without reindexing?

I will document myself about the reindexing suggestion but I think that will be my last option

Thank you yet again for your help in this matter,  
Gabe.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [January 29, 2021, 6:03pm UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/4 "2021-01-29T18:03:01Z")

</div>

> [@gabe.ks11](#):
>
> I also tried what you suggested with splitting the aggregation in parts, the missing hashes are still not returned in the set.

Put a cardinality agg on the hash field. If the count it returns is \> shard\_size then your partition\_size is too small.

---

<div class="post-metadata">

**Author:** ![gabe.ks11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gabe.ks11/32/83109_2.png) [@gabe.ks11](https://discuss.elastic.co/u/gabe.ks11)\
**Post date:** [January 29, 2021, 6:35pm UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/5 "2021-01-29T18:35:56Z")

</div>

Hey Mark,

The last set on which I am doing my tests should return less than 10000 which is less than the default maximum allowed buckets for an aggregation. A single aggregation will return the same result as the partitioned requests.

I have tried to iterate through them using 20x 500 partitions and another 10 x 1000 partitions.

None of the sets include the missing hashes 🤔. Just wondering if there is some hidden mechanic which excludes these missing hashes. If it was a matter of accuracy I should have been able to get a match on them even if their doc count was 1 but they are missing altogether unless I add a strict equal query to match a specific hash.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [January 29, 2021, 10:02pm UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/6 "2021-01-29T22:02:48Z")

</div>

Try partitioning but with shard\_min\_doc\_count = 1  
I think 0 may try return terms that don’t match the query too.

---

<div class="post-metadata">

**Author:** ![gabe.ks11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gabe.ks11/32/83109_2.png) [@gabe.ks11](https://discuss.elastic.co/u/gabe.ks11)\
**Post date:** [January 31, 2021, 2:28pm UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/7 "2021-01-31T14:28:21Z")

</div>

Hi Mark,

I believe I have tried it before without success but I tried it yet again just to make sure and the results are still not returning the missing hashes. Both shard\_min\_doc\_count 0 and 1 return the same result, without the missing hashes.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [February 1, 2021, 9:55am UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/8 "2021-02-01T09:55:54Z")

</div>

This should be working (otherwise there's a bug).  
I think we'll need to see some JSON to see what's up.

Can you share :

1. The partitioned agg request which you tried and doesn't work
2. The "hash == x" request you did to prove the duplicate hashes are there
3. An example of the docs (relevant fields only) which are duplicates.

Thanks

---

<div class="post-metadata">

**Author:** ![gabe.ks11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gabe.ks11/32/83109_2.png) [@gabe.ks11](https://discuss.elastic.co/u/gabe.ks11)\
**Post date:** [February 1, 2021, 10:30am UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/9 "2021-02-01T10:30:53Z")

</div>

Hi Mark,

I will try to prepare this for you, a few clarifications:

1. Do you wish min shard count to be 0 or 1 ?
2. Do you with to have the results of a query with min\_doc\_count 1 or 2 ? (with 1 there will be more results)

I will start preparing the data for you and thank you very much for helping out with this one!

Kind regards,  
Gabriel K.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [February 1, 2021, 10:52am UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/10 "2021-02-01T10:52:28Z")

</div>

final min\_doc\_count = 2 (to only find the duplicates)  
shard\_min\_doc\_count = 1 (to only find docs that match you query at least once)

Is the query important here? We'll only find duplicate hashes if they exist in docs that match the query

---

<div class="post-metadata">

**Author:** ![gabe.ks11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gabe.ks11/32/83109_2.png) [@gabe.ks11](https://discuss.elastic.co/u/gabe.ks11)\
**Post date:** [February 1, 2021, 11:08am UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/11 "2021-02-01T11:08:15Z")

</div>

I am using the query to restrict the amount of records returned so that it is easier to focus on one hash instance which is missing from the result set.

I am collecting and cleaning the data so it should be here soon 🙂

---

<div class="post-metadata">

**Author:** ![gabe.ks11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gabe.ks11/32/83109_2.png) [@gabe.ks11](https://discuss.elastic.co/u/gabe.ks11)\
**Post date:** [February 1, 2021, 12:08pm UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/12 "2021-02-01T12:08:57Z")

</div>

Relavant metadata stored for the given hash  
&  
Query 1 and results returning one of the missing hashes:

> **[Relavant metadata stored for the given hash: "hits": \[{ "\_ind -...](https://pastebin.com/jJER92Vn)**
>
> Pastebin.com is the number one paste tool since 2002. Pastebin is a website where you can store text online for a set period of time.

Query 2 and results returning the set not containing the result of Query 1:  
Part 1: [Query 2 returning 10 pages not containing the result of Query 1 : [PART 1/2]G - Pastebin.com](https://pastebin.com/qMP97iVJ)  
Part 2: [Query 2 Part 2 - Pastebin.com](https://pastebin.com/11hsWGe0)

Let me know if there's anything more I can do to help 🙂

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [February 1, 2021, 12:29pm UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/13 "2021-02-01T12:29:51Z")

</div>

Thanks for that. I can see from the results that `doc_count_error_upper_bound` \>0 which means that the not all of the relevant data is being returned for consideration to the coordinating node.

This means that the number of partitions is too low for the size of results being considered. By increasing the number of partitions you will be reducing the number of terms being considered in any one request to a manageable subset and you should see the `doc_count_error_upper_bound` value become zero (meaning nothing was left behind on shards).

---

<div class="post-metadata">

**Author:** ![gabe.ks11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gabe.ks11/32/83109_2.png) [@gabe.ks11](https://discuss.elastic.co/u/gabe.ks11)\
**Post date:** [February 1, 2021, 12:44pm UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/14 "2021-02-01T12:44:22Z")

</div>

I understand, I will tweak my test so that the error upper bound becomes 0 and get back to you.

---

<div class="post-metadata">

**Author:** ![gabe.ks11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gabe.ks11/32/83109_2.png) [@gabe.ks11](https://discuss.elastic.co/u/gabe.ks11)\
**Post date:** [February 1, 2021, 2:41pm UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/15 "2021-02-01T14:41:00Z")

</div>

Hi Mark,

Good news, I was finally able to retrieve the hash which was missing part of the partitioned result set, thank you very much for the breakthrough.

I do have one more question, how would I go about finding the correct amount of partitions for a certain page size so that the errors will be 0, is there some sort of formula or algorithm I could use ? so I do not have to do guess work to get the optimal values.

Gabe.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [February 1, 2021, 2:59pm UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/16 "2021-02-01T14:59:06Z")

</div>

> [@gabe.ks11](#):
>
> Good news, I was finally able to retrieve the hash which was missing part of the partitioned result set, thank you very much for the breakthrough.

Great stuff!

> [@gabe.ks11](#):
>
> how would I go about finding the correct amount of partitions for a certain page size so that the errors will be 0,

[The docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html) include some guidance. I notice they say tweak settings until partitions have sorted results that include things you don't want (e.g. terms that only occur once).  
They should also say pay attention to the `doc_count_error_upper_bound` to make sure the calculations are including all counts from all shards.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 1, 2021, 2:59pm UTC](https://discuss.elastic.co/t/terms-aggregation-not-returning-keys/262634/17 "2021-03-01T14:59:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
