# Indexing slowing down aggregations a lot

**URL:** <https://discuss.elastic.co/t/indexing-slowing-down-aggregations-a-lot/156486>\
**Category:** Elasticsearch\
**Created:** [November 13, 2018, 1:51pm UTC](https://discuss.elastic.co/t/indexing-slowing-down-aggregations-a-lot/156486 "2018-11-13T13:51:41Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![GerbenKD](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gerbenkd/32/50162_2.png) [@GerbenKD](https://discuss.elastic.co/u/GerbenKD)\
**Post date:** [November 13, 2018, 1:51pm UTC](https://discuss.elastic.co/t/indexing-slowing-down-aggregations-a-lot/156486/1 "2018-11-13T13:51:41Z")

</div>

We are running a search engine with an index of around 6M documents (~100GB) on a 3 node i3.xlarge managed AWS cluster. Our sharding and replicas are at the default settings, so 5 shards, one replica each. We are using ES version 6.3.1. The index is constantly updated by a crawler, roughly performing 1500 creates, updates and deletes a minute (all implemented with bulk queries).

As part of our autocomplete system we are running (simple) terms aggregations combined with edge n-grammed fields (pretty much the last example [here](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-edgengram-tokenizer.html), combined with a terms aggregation).  
We are aware that this is not the fastest implementation option for autocomplete, but we want to be able the handle a lot of specific contexts (so no completion suggester) and the response does not need to be ultra fast.

So with that said, running the constant indexing in the background more than doubles the query time of the aggregation query on average (from ~500ms to 1000ms roughly). However, it feels inconsistent, even with caching disabled, every once in a while there might be a fast response. Almost as if there is a background process that blocks the aggregation.

What can we do to increase the performance of this aggregation while still keeping the indexing running?

We have tried the recommended "one shard per node approach", with 1 master and two replicas (all on different i3.xlarge nodes), on a smaller 30GB test index. This only made the performance worse unfortunately (almost twice a slow).

Do we just throw more hardware at it? Is there a way to always make ES prioritize search requests over index/update/delete requests, maybe be adding even more replicas?

Thanks for any help and please ask for more details and clarification if needed.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 13, 2018, 7:56pm UTC](https://discuss.elastic.co/t/indexing-slowing-down-aggregations-a-lot/156486/2 "2018-11-13T19:56:15Z")

</div>

What does CPU usage look like on the nodes? Is there anything in the logs around GC being slow or frequent?

---

<div class="post-metadata">

**Author:** ![dimzak](https://avatars.discourse-cdn.com/v4/letter/d/8491ac/32.png) [@dimzak](https://discuss.elastic.co/u/dimzak)\
**Post date:** [November 14, 2018, 2:58pm UTC](https://discuss.elastic.co/t/indexing-slowing-down-aggregations-a-lot/156486/3 "2018-11-14T14:58:38Z")

</div>

Hi Christian!  
Adding more details about the autocomplete issue we are facing with Gerben 🙂

Besides sharding changes we also tried different instance families in our es cluster but no luck.  
Below is a screenshot of grafana showing cpu rate and gc information:

 ![cpu_gc_es_search](https://us1.discourse-cdn.com/elastic/original/3X/d/7/d71e1eea7e3d616b54825932f5dfc5619e3d717c.png)

And in [application logs](https://aws.amazon.com/blogs/big-data/viewing-amazon-elasticsearch-service-error-logs/) we got these gc `warnings`:

```auto
[2018-11-06T04:45:18,658][WARN][o.e.m.j.JvmGcMonitorService] [UfBO8nG] [gc][young][414178][1443] duration [3s], collections [1] __PATH__ [3.5s], total [3s] __PATH__ [32.8s], memory [309.2mb]->[250.7mb] __PATH__ [1015.6mb], all_pools {[young] [65.3mb]->[244.2kb] __PATH__ [66.5mb]}{[survivor] [7.8mb]->[8.3mb] __PATH__ [8.3mb]}{[old] [236.1mb]->[242.2mb] __PATH__ [940.8mb]}
[2018-11-06T04:45:18,658][WARN][o.e.m.j.JvmGcMonitorService] [UfBO8nG] [gc][414178] overhead, spent [3s] collecting in the last [3.5s]
[2018-11-06T04:51:22,904][WARN][o.e.m.j.JvmGcMonitorService] [UfBO8nG] [gc][young][414535][1444] duration [7.2s], collections [1] __PATH__ [8.1s], total [7.2s] __PATH__ [40s], memory [313.7mb]->[257.8mb] __PATH__ [1015.6mb], all_pools {[young] [63.2mb]->[579.2kb] __PATH__ [66.5mb]}{[survivor] [8.3mb]->[6.7mb] __PATH__ [8.3mb]}{[old] [242.2mb]->[250.5mb] __PATH__ [940.8mb]}
[2018-11-06T04:51:22,904][WARN][o.e.m.j.JvmGcMonitorService] [UfBO8nG] [gc][414535] overhead, spent [7.2s] collecting in the last [8.1s]
...
[2018-11-12T04:22:15,579][WARN][o.e.m.j.JvmGcMonitorService] [CDiRrfk] [gc][453860] overhead, spent [535ms] collecting in the last [1s]

```

Thanks a lot!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 14, 2018, 3:11pm UTC](https://discuss.elastic.co/t/indexing-slowing-down-aggregations-a-lot/156486/4 "2018-11-14T15:11:16Z")

</div>

It seems like you need more heap for that workload, so a larger instance type may help.

---

<div class="post-metadata">

**Author:** ![byronvoorbach](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/byronvoorbach/32/8283_2.png) [@byronvoorbach](https://discuss.elastic.co/u/byronvoorbach)\
**Post date:** [November 14, 2018, 3:14pm UTC](https://discuss.elastic.co/t/indexing-slowing-down-aggregations-a-lot/156486/5 "2018-11-14T15:14:09Z")

</div>

Maybe a bit off-topic, but maybe take a look at using highlights as a possible alternative for your aggregations for autocomplete 🙂 Depending on your data size + load this could give better performance + easier results.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 12, 2018, 3:14pm UTC](https://discuss.elastic.co/t/indexing-slowing-down-aggregations-a-lot/156486/6 "2018-12-12T15:14:11Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
