# Getting all documents without duplicate values in Field

**URL:** <https://discuss.elastic.co/t/getting-all-documents-without-duplicate-values-in-field/195110>\
**Category:** Elasticsearch\
**Created:** [August 13, 2019, 9:37pm UTC](https://discuss.elastic.co/t/getting-all-documents-without-duplicate-values-in-field/195110 "2019-08-13T21:37:18Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![misiakj](https://avatars.discourse-cdn.com/v4/letter/m/bcef8e/32.png) [@misiakj](https://discuss.elastic.co/u/misiakj)\
**Post date:** [August 13, 2019, 9:37pm UTC](https://discuss.elastic.co/t/getting-all-documents-without-duplicate-values-in-field/195110/1 "2019-08-13T21:37:19Z")

</div>

I have an index that looks similar to the below (it really has a number of additional fields that aren't relevant for this post)

case\_id: int  
group: string  
unique\_id: case\_id + group  
value: int

for any one case\_id value I might have several documents, each with a different group. Each of these documents should have the same value in the "Value" field. Different case\_ids might also go through different sets of groups

For Example

case\_id: 1  
group: A  
value: 3

case\_id: 1  
group: B  
value: 3

case\_id: 1  
group: C  
value: 3

case\_id: 2  
group: A  
value: 1

case\_id: 2  
group: B  
value: 1

case\_id: 3  
group: D  
value: 5

I would like to Average the Value of each case\_id. But I would only like to calculate the value of each case\_id once. So for example with the above data set I want  
case\_id 1 = 3  
case\_id 2 = 1  
case\_id 3 = 5

(3 + 1 + 5) / 3 = 3. So I should get 3 as a result.

The problem is I can't just average all case\_id values of all the documents or I would get something like  
(3 + 3 + 3 + 1 + 1 + 5) ~= 2.6667

How can I first subset a group to just one document per case\_id (doesn't matter which document) and then average the value fields?

Thanks for any help!

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [August 14, 2019, 7:37am UTC](https://discuss.elastic.co/t/getting-all-documents-without-duplicate-values-in-field/195110/2 "2019-08-14T07:37:28Z")

</div>

hey,

you can try a terms aggregation (I think histogram might work as well) on the `case_id` field and then use a subaggregation to calculate the avg.

--Alex

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 11, 2019, 7:37am UTC](https://discuss.elastic.co/t/getting-all-documents-without-duplicate-values-in-field/195110/3 "2019-09-11T07:37:30Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
