# ElasticSearch: DSL Query

**URL:** <https://discuss.elastic.co/t/elasticsearch-dsl-query/314767>\
**Category:** Elasticsearch\
**Created:** [September 20, 2022, 9:58am UTC](https://discuss.elastic.co/t/elasticsearch-dsl-query/314767 "2022-09-20T09:58:49Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![SalvoDM91](https://avatars.discourse-cdn.com/v4/letter/s/e19adc/32.png) [@SalvoDM91](https://discuss.elastic.co/u/SalvoDM91)\
**Post date:** [September 20, 2022, 9:58am UTC](https://discuss.elastic.co/t/elasticsearch-dsl-query/314767/1 "2022-09-20T09:58:49Z")

</div>

Hi guys, I am opening this topic because I have a problem with a large amount of data (14M).  
My dataset is composed as follows:

```auto
{"h":{"id":"AA001","process":"AK01","update-timestamp":1663665372171}}
{"h":{"id":"AA002","process":"AK01","update-timestamp":1663665372171}}
{"h":{"id":"AA003","process":"AK01","update-timestamp":1663665372171}}
{"h":{"id":"AA004","process":"AK01","update-timestamp":1663665372171}}
{"h":{"id":"AA001","process":"AK01","update-timestamp":1663665372172}}
{"h":{"id":"AA001","process":"AK01","update-timestamp":1663665372173}}

```

If my pipeline worked correctly, for each key (id and process) I should only have the most up-to-date update-timestamp.  
So I'm trying to count the duplicate values: knowing how many update-timestamps are associated with the same id - process.

I tried to do it with kibana, with a datatable but the volumes are too high and it goes in error.

Could you help me with dsl?

Thanks in advance!  
Salvo

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 18, 2022, 9:59am UTC](https://discuss.elastic.co/t/elasticsearch-dsl-query/314767/2 "2022-10-18T09:59:17Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
