# Kibana alternate way to remove duplicates or get precise Unique Count

**URL:** <https://discuss.elastic.co/t/kibana-alternate-way-to-remove-duplicates-or-get-precise-unique-count/174286>\
**Category:** Kibana\
**Created:** [March 28, 2019, 9:30am UTC](https://discuss.elastic.co/t/kibana-alternate-way-to-remove-duplicates-or-get-precise-unique-count/174286 "2019-03-28T09:30:21Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![groverjatin17](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/groverjatin17/32/34657_2.png) [@groverjatin17](https://discuss.elastic.co/u/groverjatin17)\
**Post date:** [March 28, 2019, 9:30am UTC](https://discuss.elastic.co/t/kibana-alternate-way-to-remove-duplicates-or-get-precise-unique-count/174286/1 "2019-03-28T09:30:22Z")

</div>

Hi All,

i have 60000 documents in an index. Many of these documents have same value for field "BookId".

NOTE:-There are 14000 unique BookId's just many duplicates because they have different values in other fields of other documents and creating 60000 total hits in an index.

I am creating a Bar Chart visualization with "Category" in X-axis and Unique count of "BookId" in Y-axis metric. But it produces wrong Unique count, upon some searching it says that it is for approximation and setting JSON to {"precision\_threshold" : 40000} would solve it. But it is still missing thousands of value.

If it is an approx value. how can I get the unique count/remove duplicates in my Bar Graph ?

Also, Can i filter out the unique in DISCOVER tab so it shows right count in hits?

---

<div class="post-metadata">

**Author:** ![christophilus](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christophilus/32/42991_2.png) [@christophilus](https://discuss.elastic.co/u/christophilus)\
**Post date:** [March 28, 2019, 2:48pm UTC](https://discuss.elastic.co/t/kibana-alternate-way-to-remove-duplicates-or-get-precise-unique-count/174286/2 "2019-03-28T14:48:26Z")

</div>

Above 40K, the results are still fuzzy, unfortunately, as documented [here](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-metrics-cardinality-aggregation.html). Since you have 60K records, I think you're still going to see fuzzy cardinality results.

We have a client-side scripting language (Kibana expressions), which you could probably put to use here, although it's not quite ready to go for the bar chart.

If you _really_ want precision, you may need to write a plugin that tallies things. Here's an example of getting a distinct count of "category" per "city" from a "pets" index, using JavaScript. You can test this locally by running Kibana like this: `yarn start --repl`. After Kibana boots, you can paste this code into the REPL (in your terminal), and then enter `clientDistinct()`, and you should see an accurate distinct count. You'll want to modify the query to actually select the fields / index you want.

```auto
async function clientDistinct(kbnServer) {
  const callCluster = kbnServer.server.plugins.elasticsearch.getCluster('admin').callWithInternalUser;
  const result = {};
  let from = 0;

  while (true) {
    // callCluster is a function which calls Elasticsearch, and may
    // not be exactly what you'd use...
    const { hits } = await callCluster('search', { 
      index: 'pets',
      body: {
        "from" : from,
        "_source" : {
          "includes" : [
            "category",
            "city"
          ],
          "excludes" : []
        },
        "sort" : [
          {
            "_doc" : {
              "order" : "asc"
            }
          }
        ]
      }
    });

    if (!hits || !hits.hits.length) {
      break;
    }

    from += hits.hits.length;

    // This does a distinct count of categories grouped by city
    hits.hits.forEach(({ _source }) => {
      const set = result[_source.city] || new Set();
      result[_source.city] = set;
      set.add(_source.category);
    });
  }

  // Returns something like: { newyork: 3, seattle: 55 }
  return Object.keys(result).reduce((acc, k) => {
    acc[k] = result[k].size;
    return acc;
  }, {});
}

```

---

<div class="post-metadata">

**Author:** ![christophilus](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christophilus/32/42991_2.png) [@christophilus](https://discuss.elastic.co/u/christophilus)\
**Post date:** [March 28, 2019, 2:57pm UTC](https://discuss.elastic.co/t/kibana-alternate-way-to-remove-duplicates-or-get-precise-unique-count/174286/3 "2019-03-28T14:57:23Z")

</div>

I should note that this is fairly trivial to do in Canvas:

```auto
essql query="SELECT city, category FROM pets"
| ply by=city fn={math "unique(category)"}

```

You'll need to modify that to query the index, and change`by=city` and `unique(category)` to be whatever columns you're working on.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 25, 2019, 2:57pm UTC](https://discuss.elastic.co/t/kibana-alternate-way-to-remove-duplicates-or-get-precise-unique-count/174286/4 "2019-04-25T14:57:38Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
