# Duplicate Field Aggregate crashes system

**URL:** <https://discuss.elastic.co/t/duplicate-field-aggregate-crashes-system/84025>\
**Category:** Elasticsearch\
**Created:** [April 28, 2017, 2:58pm UTC](https://discuss.elastic.co/t/duplicate-field-aggregate-crashes-system/84025 "2017-04-28T14:58:40Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Erik\_Miller](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/erik_miller/32/14913_2.png) [@Erik\_Miller](https://discuss.elastic.co/u/Erik_Miller)\
**Post date:** [April 28, 2017, 2:58pm UTC](https://discuss.elastic.co/t/duplicate-field-aggregate-crashes-system/84025/1 "2017-04-28T14:58:40Z")

</div>

I am trying to run a simple aggregation that returns all documents that have a duplicate field. The use case is that the field is considered an ID field and there should never be more than one specific field value across all documents.

I have tried running this against the entire document set, but then realized if there are many duplicates this may just be too much data.

So I tried running it with a filter against a know subset size that includes duplicates and only has 50 documents and it did not return after minutes worth of processing and one of the nodes heap size ballooned causing me to restart the node.

There ~50 million docs, spread across 6 nodes each with 32G memory, search query is below.

Any help would be appreciated.

```
{
  "filter": {
    "bool": {
      "must": [
        {
          "term": {
            "progservId": 96891
          }
        },
        {
          "term": {
            "schedDate": "2017-05-02"
          }
        }
      ]
    }
  },
  "aggs": {
    "duplicateNames": {
      "terms": {
        "field": "tsId",
        "size": 0,
        "min_doc_count": 2
      }
    }
  }
}
```

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 30, 2017, 8:36am UTC](https://discuss.elastic.co/t/duplicate-field-aggregate-crashes-system/84025/2 "2017-04-30T08:36:00Z")

</div>

As your duplicate documents can be spread out across multiple shards and basically all values of the duplicate field need to be compared centrally, this is very expensive. If this is something you will need to do regularly, I would recommend indexing using the duplicate field [as a routing key](https://www.elastic.co/guide/en/elasticsearch/reference/5.3/docs-index_.html#index-routing). This will make all documents with the same duplicate field to end up in the same shard, allowing you to use the [shard\_min\_doc\_count parameter](https://www.elastic.co/guide/en/elasticsearch/reference/5.3/search-aggregations-bucket-terms-aggregation.html#_minimum_document_count_3) (set to 2) to identify duplicates already at the shard level, which should be much more efficient.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 28, 2017, 8:44am UTC](https://discuss.elastic.co/t/duplicate-field-aggregate-crashes-system/84025/3 "2017-05-28T08:44:37Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
