# \[BUG?\] Wrong aggregated values shown in visualization

**URL:** <https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888>\
**Category:** Kibana\
**Created:** [September 18, 2017, 2:37pm UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888 "2017-09-18T14:37:03Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![pranav0091](https://avatars.discourse-cdn.com/v4/letter/p/cab0a1/32.png) [@pranav0091](https://discuss.elastic.co/u/pranav0091)\
**Post date:** [September 18, 2017, 2:37pm UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888/1 "2017-09-18T14:37:03Z")

</div>

I seem to have run across what is essentially a deal-breaking bug for my use case:  
When Visualizing an _ **Average** _ of a numeric field on the Y-axis versus a name on the X-axis the data shown differs substantially depending on whether or not I have filters enabled.

Here are the details:

- A few thousand (2000 - 20000) documents each having the same fields (1000-4000)
- Plotting a field (say "name" ) versus the "average" of another field
  - If x-axis has some terms on it, I get one value for the averages ( [393,940 in the example video](https://youtu.be/C_fN8qyf6NY?t=3s))
  - If I filter the X-axis terms (by clicking on it) I get another set of values ([466,667.2 in the example video](https://youtu.be/C_fN8qyf6NY?t=57s)).

**Why this is a major bug for me** :  
If I want to find an outlier (say, 5 lowest averages) then I cannot do so now, because what I see in the aggregated view often has no bearing with the actual aggregate for real data underneath. The issue is much worse when dealing with percentages.

I [have a dataset](https://drive.google.com/file/d/0B0uWf5JK6JyiMmdSaDZCX1pmOFk/view?usp=sharing) that can be used to reproduce the bug, and I have made [a video](https://youtu.be/C_fN8qyf6NY) showing it in action.

Please let me know if I am doing something incorrect, or if this issue supposed to me filed under ElasticSearch.

---

<div class="post-metadata">

**Author:** ![pranav0091](https://avatars.discourse-cdn.com/v4/letter/p/cab0a1/32.png) [@pranav0091](https://discuss.elastic.co/u/pranav0091)\
**Post date:** [September 18, 2017, 2:42pm UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888/2 "2017-09-18T14:42:53Z")

</div>

Some notes:

1. Easier to reproduce the bug when you have a large number of documents and fields.

2. ElasticSearch response seems to indicate that ALL the documents are _hits_, but in the `Average` aggregation, there are only a few hits (ie, `doc_count` doesnt represent all the available docs for that _field_ )

---

<div class="post-metadata">

**Author:** ![LeeDr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leedr/32/9289_2.png) [@LeeDr](https://discuss.elastic.co/u/LeeDr)\
**Post date:** [September 18, 2017, 9:42pm UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888/3 "2017-09-18T21:42:27Z")

</div>

It might be this issue;

[https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html#search-aggregations-bucket-terms-aggregation-approximate-counts](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html#search-aggregations-bucket-terms-aggregation-approximate-counts)

I'll look into it more, but that seems like it might be the issue.

Regards,  
Lee

---

<div class="post-metadata">

**Author:** ![pranav0091](https://avatars.discourse-cdn.com/v4/letter/p/cab0a1/32.png) [@pranav0091](https://discuss.elastic.co/u/pranav0091)\
**Post date:** [September 19, 2017, 6:52am UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888/4 "2017-09-19T06:52:13Z")

</div>

Thanks Lee. My reading of the page you linked to shows it as a performance feature - which is nice and often downright essential.

I'd like to make a case for a choice to override this on this an index-by-index granularity. The thing is, while I generate most of the Visualizations, some of the consumers (who are often not savvy) also create their own visualizations - and they see this _unexpected_ behaviour and are put off by the tool as _unreliable_. Having an index level option would let me ensure that they see what they expect.

**EDIT:**  
Rerunning the dataset with 1 shard seems to make the bug go away (I used 4 shards until now) - so it seems likely that what you said is indeed the reason for the bug.  
But I assume I am leaving a lot of performance on the table since my system has 4 threads available and a shard would use only 1? Would increasing copies help?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 19, 2017, 8:43am UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888/5 "2017-09-19T08:43:18Z")

</div>

Can you tell us a bit more about your use case? Are you using time-based indices? How large are your shards not that you are using a single primary shard?

---

<div class="post-metadata">

**Author:** ![pranav0091](https://avatars.discourse-cdn.com/v4/letter/p/cab0a1/32.png) [@pranav0091](https://discuss.elastic.co/u/pranav0091)\
**Post date:** [September 19, 2017, 9:38am UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888/6 "2017-09-19T09:38:05Z")

</div>

No I am not using time-based indices.

I have a few different indices (less than 10) - each of which originated as a single `.json` file. Presently the largest one has **20k rows** (documents) each with **~2k fields** (the same 2000 fields on each document) - that one is ~1.5GB when stored as a `.json` with one document per line.

The reason why I was using 4 shards was because I assumed from my reading of the docs ( **correct me if I am mistaken** ) that one CPU-thread can work on a shard and so 4 shards seemed to be a good match for 4 CPU threads.

Going forward I have one larger index in mind that would be **~70k documents** - each with the same **10-12k fields**. I'd assume thats going to be 15-20GB as a `.json`

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 19, 2017, 9:43am UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888/7 "2017-09-19T09:43:19Z")

</div>

> [@pranav0091](#):
>
> The reason why I was using 4 shards was because I assumed from my reading of the docs (correct me if I am mistaken) that one CPU-thread can work on a shard and so 4 shards seemed to be a good match for 4 CPU threads.

That is correct. Are you limited by CPU when you query? Is latency too high when you only use a single core?

---

<div class="post-metadata">

**Author:** ![pranav0091](https://avatars.discourse-cdn.com/v4/letter/p/cab0a1/32.png) [@pranav0091](https://discuss.elastic.co/u/pranav0091)\
**Post date:** [September 19, 2017, 10:00am UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888/8 "2017-09-19T10:00:33Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> That is correct. Are you limited by CPU when you query? Is latency too high when you only use a single core?

Not so far 🙂  
I was just trying to eke out as much performance as I could - the low hanging fruits.

Why I went with 4 shards at all?  
As an aside, when I began exploring Kibana+ ES I did make an attempt with a 15GB `.json` - 8 shards, 1 copy, **~70k documents** each with **~6500 (common) fields** spread on two machines (each with 4 cores, 8GB of JVM memory _each_, data on SSD, ping times of under 20ms between them) - that was not successful. I'd hit timeout or some other issue when either

1. Trying to create an Index in Kibana (after successfully posting to ES). OR
2. Trying to pick the index under New Visualization.

This made me concerned about performance, and therefore I dropped the number of fields massively and kept 4 shards on the same machine

---

<div class="post-metadata">

**Author:** ![pranav0091](https://avatars.discourse-cdn.com/v4/letter/p/cab0a1/32.png) [@pranav0091](https://discuss.elastic.co/u/pranav0091)\
**Post date:** [September 19, 2017, 7:25pm UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888/9 "2017-09-19T19:25:26Z")

</div>

Will increasing the number of _replicas_ helps ES parallelize its search when there the index was created with just one shard?

Edit:  
May I file an issue at github for the ability to override/set the shard\_size to enforce that all samples are read?

---

<div class="post-metadata">

**Author:** ![pranav0091](https://avatars.discourse-cdn.com/v4/letter/p/cab0a1/32.png) [@pranav0091](https://discuss.elastic.co/u/pranav0091)\
**Post date:** [September 21, 2017, 9:33pm UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888/10 "2017-09-21T21:33:33Z")

</div>

Bump

---

<div class="post-metadata">

**Author:** ![LeeDr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leedr/32/9289_2.png) [@LeeDr](https://discuss.elastic.co/u/LeeDr)\
**Post date:** [September 26, 2017, 2:55pm UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888/11 "2017-09-26T14:55:50Z")

</div>

I think this is a question that you might get the best response on from posting a question on the Elasticsearch forum.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 24, 2017, 2:56pm UTC](https://discuss.elastic.co/t/bug-wrong-aggregated-values-shown-in-visualization/100888/12 "2017-10-24T14:56:03Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
