# Aggregations slow after inserts/updates

**URL:** <https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352>\
**Category:** Elasticsearch\
**Created:** [July 19, 2019, 8:55am UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352 "2019-07-19T08:55:04Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![JF2018](https://avatars.discourse-cdn.com/v4/letter/j/edb3f5/32.png) [@JF2018](https://discuss.elastic.co/u/JF2018)\
**Post date:** [July 19, 2019, 8:55am UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352/1 "2019-07-19T08:55:04Z")

</div>

Hi

I have a query, where part of it is a terms-aggregation.  
The query takes \<500ms normally, but after inserts/updates to the index, the query takes  
~2-3000ms the first couple of times .

If I remove the aggregation part of the query, I do not see the same performance problems after inserts/updates.

I am running on elastic 1.7, but can see people experience the same problems on other versions:

> [@Slow aggregation queries while ingesting data](https://discuss.elastic.co/t/slow-aggregation-queries-while-ingesting-data/108212):
>
> I have a large aggregation query that becomes incredibly slow when I am updating my data. I am not saving the data to a tmp index (and then renaming it when it's done) but saving it directly to the index I'm querying. What are some ways to improve querying performance while indexing is occurring? What are the usual bottlenecks that I'm seeing here (possibly memory?)?

> [@Slow aggregation queries, only after data change (ES 2.3)](https://discuss.elastic.co/t/slow-aggregation-queries-only-after-data-change-es-2-3/66398/7):
>
> it helps a little but you have to look at the OS/Hardware at this point 24M records / 6 shards / 3 nodes = 1 million records each shard is searching. So, you will probably have some Load average and IO on each system that you should check as that will have the most impact on searchs The other is what your actual search is (Can you provide an example) but in general if your using lots of \* or doing \_all fields will slow you down Try reading this document before we go much further

The solution used in one of the threads, using filter-aggregation, is not possible for me, since the scoring is key for the query.

The threads I can find are from 2016 and 2017, so maybe someone has experienced it in the meantime, and found an explanation/solution?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 19, 2019, 9:18am UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352/2 "2019-07-19T09:18:47Z")

</div>

What kind of hardware do you have? Are you using `doc_values`?  
Can you upgrade? We are now on 7.2 and so many things happened in the last 3 or 4 years... Including in the JVM itself.

---

<div class="post-metadata">

**Author:** ![JF2018](https://avatars.discourse-cdn.com/v4/letter/j/edb3f5/32.png) [@JF2018](https://discuss.elastic.co/u/JF2018)\
**Post date:** [July 19, 2019, 9:33am UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352/3 "2019-07-19T09:33:10Z")

</div>

Sadly I can not upgrade at the current time, even though it is one of my biggest wishes 🙏  
We are using doc\_values yes, as far as I understand aggregation is not possible without using doc\_values?

We are running on instances with 4 virtual cores, 16GB ram with 7GB allocated to the JVM and 2TB discs

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 19, 2019, 9:35am UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352/4 "2019-07-19T09:35:06Z")

</div>

If you are not using `doc_values` (on disk) I think it's using `fielddata` then (in memory).  
Are you using spinning disks?

---

<div class="post-metadata">

**Author:** ![JF2018](https://avatars.discourse-cdn.com/v4/letter/j/edb3f5/32.png) [@JF2018](https://discuss.elastic.co/u/JF2018)\
**Post date:** [July 19, 2019, 10:04am UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352/5 "2019-07-19T10:04:38Z")

</div>

Okay, in the newest versions (7.2) I think doc\_values are necessary

[All fields which support doc values have them enabled by default. If you are sure that you don’t need to sort or aggregate on a field, or access the field value from a script, you can disable doc values in order to save disk space:](https://www.elastic.co/guide/en/elasticsearch/reference/7.2/doc-values.html)

But cant find it for the 1.7 version. But I guess doc-values should be the intended way to do it.

It is ssd, but over network, so a bit slower than if they were local

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 19, 2019, 10:20am UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352/6 "2019-07-19T10:20:37Z")

</div>

> [@JF2018](#):
>
> It is ssd, but over network, so a bit slower than if they were local

That's probably why it's slow. You should use local disks instead.

---

<div class="post-metadata">

**Author:** ![JF2018](https://avatars.discourse-cdn.com/v4/letter/j/edb3f5/32.png) [@JF2018](https://discuss.elastic.co/u/JF2018)\
**Post date:** [July 19, 2019, 10:38am UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352/7 "2019-07-19T10:38:03Z")

</div>

So the reason for it being slow, after data updates, should be that a cache is flushed, and thus it has to fill it again from disc, which is then slow because its over network?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 19, 2019, 11:11am UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352/8 "2019-07-19T11:11:23Z")

</div>

When you update/add data, you are writing new segments on disk. Also if segment merges needs to happen, more data then have to be read on disk.  
Which means that new search needs to read again new data from disk.

---

<div class="post-metadata">

**Author:** ![JF2018](https://avatars.discourse-cdn.com/v4/letter/j/edb3f5/32.png) [@JF2018](https://discuss.elastic.co/u/JF2018)\
**Post date:** [July 19, 2019, 11:15am UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352/9 "2019-07-19T11:15:36Z")

</div>

Makes good sense, I will try to change to a setup with local disks and test if that is enough to solve the problem.  
Thank you very much, for the help!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 19, 2019, 11:56am UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352/10 "2019-07-19T11:56:09Z")

</div>

And it's important that you think of upgrading.

---

<div class="post-metadata">

**Author:** ![JF2018](https://avatars.discourse-cdn.com/v4/letter/j/edb3f5/32.png) [@JF2018](https://discuss.elastic.co/u/JF2018)\
**Post date:** [July 24, 2019, 12:06pm UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352/11 "2019-07-24T12:06:02Z")

</div>

For future reference:

My problem was not that the disk wasn't local.  
I tried to change to local disks, without any improvements.

The problem however seemed to be, that the field I do my aggregation on, has a very high cardinality (\>1million).  
Thus the global ordinals where taking a long time to recompute after data changes.  
As default global ordinals are lazy loaded - that is on first search after changes.

By changing it to be eager-loaded, I pay on insert/refresh-time instead of search-time.  
In my case this is a fine solution, since there are no requirements to the time for inserts/updates.

This seems to solve my problems.

[global ordinals docs](https://www.elastic.co/guide/en/elasticsearch/reference/master/eager-global-ordinals.html)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 21, 2019, 12:06pm UTC](https://discuss.elastic.co/t/aggregations-slow-after-inserts-updates/191352/12 "2019-08-21T12:06:06Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
