# Getting All records with doc\_count \>= 2

**URL:** <https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407>\
**Category:** Elasticsearch\
**Tags:** kql-kibana-query-language\
**Created:** [April 14, 2022, 9:01am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407 "2022-04-14T09:01:39Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![minhaj\_shakeel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/minhaj_shakeel/32/104387_2.png) [@minhaj\_shakeel](https://discuss.elastic.co/u/minhaj_shakeel)\
**Post date:** [April 14, 2022, 9:01am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/1 "2022-04-14T09:01:39Z")

</div>

I have 2 es indices. There are some records common in both the indices, I want to fetch them. My preliminary query is :-

```auto
GET index1,index2/_search
{
  "size": 0,
  "aggs": {
    "duplicate record": {
      "terms": {
        "field": "Id"
      },
      "aggs": {
        "get duplicate": {
          "bucket_selector": {
            "buckets_path": {
              "doc_count": "_count"
            },
            "script": "params.doc_count == 2"
          }
        }
      }
    }
  }
}

```

This query returns the following result:-

```auto
{
  "took" : 687,
  "timed_out" : false,
  "_shards" : {
    "total" : 2,
    "successful" : 2,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 10000,
      "relation" : "gte"
    },
    "max_score" : null,
    "hits" : []
  },
  "aggregations" : {
    "duplicate record" : {
      "doc_count_error_upper_bound" : 2,
      "sum_other_doc_count" : 3000689,
      "buckets" : [
        {
          "key" : "1989861726",
          "doc_count" : 2
        },
        {
          "key" : "1989861734",
          "doc_count" : 2
        },
        {
          "key" : "1989861860",
          "doc_count" : 2
        },
        {
          "key" : "1989862267",
          "doc_count" : 2
        },
        {
          "key" : "1989862280",
          "doc_count" : 2
        },
        {
          "key" : "1989862642",
          "doc_count" : 2
        },
        {
          "key" : "1989862730",
          "doc_count" : 2
        },
        {
          "key" : "2004088312",
          "doc_count" : 2
        },
        {
          "key" : "2004088315",
          "doc_count" : 2
        },
        {
          "key" : "2004088321",
          "doc_count" : 2
        }
      ]
    }
  }
}

```

I want full records associated with these id's. Can someone help me with that?

---

<div class="post-metadata">

**Author:** ![RabBit\_BR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rabbit_br/32/82261_2.png) [@RabBit\_BR](https://discuss.elastic.co/u/RabBit_BR)\
**Post date:** [April 14, 2022, 1:06pm UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/2 "2022-04-14T13:06:47Z")

</div>

> [@minhaj\_shakeel](#):
>
> ` "script": "params.doc_count == 2"`

you do not want: params.doc\_count \>= 2 ?

---

<div class="post-metadata">

**Author:** ![anime\_lover](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anime_lover/32/100265_2.png) [@anime\_lover](https://discuss.elastic.co/u/anime_lover)\
**Post date:** [April 15, 2022, 5:24am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/3 "2022-04-15T05:24:08Z")

</div>

hi there ,  
i am confused with your question can you elaborate a bit ...like what are the fields containing in those indices ? are both indices consist of same fields ? or only id field is common?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [April 15, 2022, 12:39pm UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/4 "2022-04-15T12:39:44Z")

</div>

> [@minhaj\_shakeel](#):
>
> I want full records associated with these id's. Can someone help me with that?

The terms aggregation has a min\_doc\_count setting which does this but can be inaccurate if there are high numbers of ids that occur infrequently. It’s what I call the [Elizabeth Taylor problem](https://discuss.elastic.co/t/aggregation-based-on-terms-field-is-not-working/240423/3)  
If you have many unique rare terms you may need to look at a different strategy which we can discuss.

---

<div class="post-metadata">

**Author:** ![minhaj\_shakeel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/minhaj_shakeel/32/104387_2.png) [@minhaj\_shakeel](https://discuss.elastic.co/u/minhaj_shakeel)\
**Post date:** [April 18, 2022, 6:18am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/5 "2022-04-18T06:18:18Z")

</div>

both the indices contain the same field. and all entries are unique in an index. However, 2 indices may have common records with the same `Id`.  
My Question:- `Give me the list of all records which are present in both the indices.`

---

<div class="post-metadata">

**Author:** ![minhaj\_shakeel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/minhaj_shakeel/32/104387_2.png) [@minhaj\_shakeel](https://discuss.elastic.co/u/minhaj_shakeel)\
**Post date:** [April 18, 2022, 6:19am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/6 "2022-04-18T06:19:35Z")

</div>

My question is how to get records from the term aggregation? Like how to filter hits based on the aggregation?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [April 18, 2022, 6:28am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/7 "2022-04-18T06:28:17Z")

</div>

> [@minhaj\_shakeel](#):
>
> how to get records from the term aggregation

For low cardinality terms you’d use A nested top hits agg but for high cardinality you’d need to issue a second follow-up request.  
Your main problem is likely to be firstly correctly identifying the \>2 terms.

---

<div class="post-metadata">

**Author:** ![minhaj\_shakeel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/minhaj_shakeel/32/104387_2.png) [@minhaj\_shakeel](https://discuss.elastic.co/u/minhaj_shakeel)\
**Post date:** [April 18, 2022, 7:21am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/8 "2022-04-18T07:21:14Z")

</div>

I did not get your solution for higher cardinality terms? Can you please elaborate?

---

<div class="post-metadata">

**Author:** ![minhaj\_shakeel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/minhaj_shakeel/32/104387_2.png) [@minhaj\_shakeel](https://discuss.elastic.co/u/minhaj_shakeel)\
**Post date:** [April 18, 2022, 9:50am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/9 "2022-04-18T09:50:53Z")

</div>

In the nested top hits agg query, my main concern is how to get all the buckets without specifying the size.

---

<div class="post-metadata">

**Author:** ![anime\_lover](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anime_lover/32/100265_2.png) [@anime\_lover](https://discuss.elastic.co/u/anime_lover)\
**Post date:** [April 18, 2022, 11:32am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/10 "2022-04-18T11:32:13Z")

</div>

> [@minhaj\_shakeel](#):
>
> In the nested top hits agg query, my main concern is how to get all the buckets without specifying the size.

you can get all bucket by using pagination .For that you can use bucket sort or composite aggregation

---

<div class="post-metadata">

**Author:** ![minhaj\_shakeel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/minhaj_shakeel/32/104387_2.png) [@minhaj\_shakeel](https://discuss.elastic.co/u/minhaj_shakeel)\
**Post date:** [April 18, 2022, 11:33am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/11 "2022-04-18T11:33:56Z")

</div>

I am unable to find how to write my use case in composite aggregation. (with top\_hit). Can you point to the relevant doc?

---

<div class="post-metadata">

**Author:** ![anime\_lover](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anime_lover/32/100265_2.png) [@anime\_lover](https://discuss.elastic.co/u/anime_lover)\
**Post date:** [April 18, 2022, 11:51am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/12 "2022-04-18T11:51:32Z")

</div>

composite aggregation doesn't support ordering so use bucket sort to find top hit .I don't know from where i found this query but i think it might help you out.

```auto
GET /billdetail/_search
{
  "size": 0,
  "aggs": {
    "topData": {
      "terms": {
        "field": "billid.keyword",
        "size": 100
      },
      
      "aggs": {
        "paging": {
          "bucket_sort": {
            "sort": [
              {
                "_count": {
                  "order": "desc"
                }
              }
            ],
            "from": 0,
            "size": 10
          }
        }
      }
    }
  }
}

```

take it as a reference and also look about bucket sort in Elasticsearch documentation.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [April 18, 2022, 12:00pm UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/13 "2022-04-18T12:00:37Z")

</div>

> [@minhaj\_shakeel](#):
>
> I did not get your solution for higher cardinality terms? Can you please elaborate?

Put simply, if the data you want to link together is spread across multiple machines and there’s lots of individual keys to be linked on, then that’s a tough ask of any system.  
To offer fast analysis you need to prepare the data better by bringing related data closer together. Custom document “routing” can be used to keep related content on the same machine while the new “transforms” api can be used to copy related data into the same document.

---

<div class="post-metadata">

**Author:** ![minhaj\_shakeel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/minhaj_shakeel/32/104387_2.png) [@minhaj\_shakeel](https://discuss.elastic.co/u/minhaj_shakeel)\
**Post date:** [April 18, 2022, 12:43pm UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/14 "2022-04-18T12:43:46Z")

</div>

1 - why you have mentioned the inner size = 100 when we don't know how many buckets will be there

---

<div class="post-metadata">

**Author:** ![anime\_lover](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anime_lover/32/100265_2.png) [@anime\_lover](https://discuss.elastic.co/u/anime_lover)\
**Post date:** [April 19, 2022, 5:03am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/15 "2022-04-19T05:03:55Z")

</div>

oh sorry about that ...  
in my project i just needed top 100 so i used that ......  
It is recommended to keep it large i don't know exact value but in documentation it told to keep large ...  
According to my understanding it is like a ocean which collects all the bucket and sort that bucket ..  
And one more thing i didn't find any difference in result on keeping size to 100 and size to 1000 .

If you find my concept is wrong then feel free to correct it and give suggestions thank you.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 17, 2022, 5:04am UTC](https://discuss.elastic.co/t/getting-all-records-with-doc-count-2/302407/16 "2022-05-17T05:04:33Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
