# Terms aggregation doesn't return all hits

**URL:** <https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353>\
**Category:** Elasticsearch\
**Created:** [November 13, 2020, 2:16pm UTC](https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353 "2020-11-13T14:16:18Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Zining](https://avatars.discourse-cdn.com/v4/letter/z/8baadc/32.png) [@Zining](https://discuss.elastic.co/u/Zining)\
**Post date:** [November 13, 2020, 2:16pm UTC](https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353/1 "2020-11-13T14:16:19Z")

</div>

Hi,  
We are using Elasticsearch 5.6 to store track events. Recently we run Terms aggregation on one index to find out duplicated events which have same **event type, device id, and event time**. Then we remove the duplicated ones from the index. The index contains about 300k events and most of them are unique.

The following query is used to find out duplications. We **loop sending** this query and remove the duplicated events until nothing is found.

```auto
{
	"size": 0,
	"aggs": {
		"DuplicatedEvents": {
			"terms": {
				"script": "return doc['evt_time'].value+doc['device_id'].value+doc['type'].value;",
				"size": 1000,
				"min_doc_count": 2
			},
			"aggs": {
				"hits": {
					"top_hits": {
						"size": 1000
					}
				}
			}
		}
	}
}

```

It runs smoothly, most of the duplications are removed. However, we noticed that few duplications are still there when we search events by a device id which had plenty of duplicated events before. We run the above query again to verify and it's pretty sure that nothing is returned. Then we try to **increase the size in terms aggregation** in the same query and this time it returns these duplications.

It confuses us a lot, why these duplicated events doesn't return until we increase the number of size in terms aggregation? We already loop sending the query and remove duplicated events until nothing is found.

Is there any better solution to remove duplicated documents in Elasticsearch 5.6? (upgrading to other version is not an option for us right now)

If we have to increase the size to a very large number, will it crash the cluster?

Many thanks!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 13, 2020, 3:10pm UTC](https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353/2 "2020-11-13T15:10:26Z")

</div>

I think that you have a solution described in the documentation (note that it will require that you upgrade which you should do anyway):

> **[Terms aggregation | Elasticsearch Guide \[8.11\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html#search-aggregations-bucket-terms-aggregation-size)**

> If you want to retrieve **all** terms or all combinations of terms in a nested `terms` aggregation you should use the [Composite](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-composite-aggregation.html) aggregation which allows to paginate over all possible terms rather than setting a size greater than the cardinality of the field in the `terms` aggregation. The `terms` aggregation is meant to return the `top` terms and does not allow pagination.

---

<div class="post-metadata">

**Author:** ![Zining](https://avatars.discourse-cdn.com/v4/letter/z/8baadc/32.png) [@Zining](https://discuss.elastic.co/u/Zining)\
**Post date:** [November 13, 2020, 3:24pm UTC](https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353/3 "2020-11-13T15:24:41Z")

</div>

Thanks David for your help! Unfortunately we need to solve the problem before we can upgrade to latest ES version.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 13, 2020, 4:47pm UTC](https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353/4 "2020-11-13T16:47:31Z")

</div>

How large is your index? How many shards does it have?

---

<div class="post-metadata">

**Author:** ![Zining](https://avatars.discourse-cdn.com/v4/letter/z/8baadc/32.png) [@Zining](https://discuss.elastic.co/u/Zining)\
**Post date:** [November 16, 2020, 8:41am UTC](https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353/5 "2020-11-16T08:41:19Z")

</div>

Hi Christian,  
The index has around 300K documents with 3 shards.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 16, 2020, 9:24am UTC](https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353/6 "2020-11-16T09:24:44Z")

</div>

If you have duplicates spread across multiple shards with less than 2 on all shards it is difficult to find them as only the top terms are returned from each shard and these can easily be missed. Given the small number of documents the best way might be to change to a single primary shard.

---

<div class="post-metadata">

**Author:** ![Zining](https://avatars.discourse-cdn.com/v4/letter/z/8baadc/32.png) [@Zining](https://discuss.elastic.co/u/Zining)\
**Post date:** [November 17, 2020, 7:49am UTC](https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353/7 "2020-11-17T07:49:44Z")

</div>

Thanks!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 17, 2020, 7:51am UTC](https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353/8 "2020-11-17T07:51:41Z")

</div>

If you do not want to merge you may also be able to run multiple aggregations against subsets of your data as that may increase the chance that spread out duplicates are detected.

---

<div class="post-metadata">

**Author:** ![Zining](https://avatars.discourse-cdn.com/v4/letter/z/8baadc/32.png) [@Zining](https://discuss.elastic.co/u/Zining)\
**Post date:** [November 17, 2020, 7:58am UTC](https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353/9 "2020-11-17T07:58:33Z")

</div>

Hi Christian,  
What is multiple aggregations? Could you please share an example to detect duplicates? Many thanks!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 17, 2020, 8:04am UTC](https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353/10 "2020-11-17T08:04:39Z")

</div>

Send multiple requests where each aggregation only aggregates across a subset of the data in the index, e.g. a timestamp range.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 15, 2020, 8:05am UTC](https://discuss.elastic.co/t/terms-aggregation-doesnt-return-all-hits/255353/11 "2020-12-15T08:05:11Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
