# Elasticsearch aggregation performance issue -- index allocation

**URL:** <https://discuss.elastic.co/t/elasticsearch-aggregation-performance-issue-index-allocation/52927>\
**Category:** Elasticsearch\
**Created:** [June 15, 2016, 11:11pm UTC](https://discuss.elastic.co/t/elasticsearch-aggregation-performance-issue-index-allocation/52927 "2016-06-15T23:11:46Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![sharon.c](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sharon.c/32/16076_2.png) [@sharon.c](https://discuss.elastic.co/u/sharon.c)\
**Post date:** [June 15, 2016, 11:11pm UTC](https://discuss.elastic.co/t/elasticsearch-aggregation-performance-issue-index-allocation/52927/1 "2016-06-15T23:11:46Z")

</div>

My company is using elasticsearch to transaction log indexing and aggregation. The data stream in elasticsearch from logstash at the speed of 1000-1500 messages per second. We currently have two data nodes to do both index and aggregation at the same time.

When aggregation task is big, we can see from marvel that elasticsearch node cpu load increases dramatically , and even indexing stops temporarily. Aggregation response time is long.

When we do aggregation without streaming data in elasticsearch (elasticsearch only has aggregation task not indexing task), then we can see the aggregation speed is much faster.

It looks like the indexing task and aggregation task has to be in different nodes to optimal performance.

Is that this a good solution if we allocate the **old indices** to the data node only **in charge of aggregation** , and **new indices** (indexing task is still going on) to the nodes only **in charge of indexing**?

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [June 16, 2016, 8:11am UTC](https://discuss.elastic.co/t/elasticsearch-aggregation-performance-issue-index-allocation/52927/2 "2016-06-16T08:11:56Z")

</div>

Hey,

the problem from the outside sounds less likely, that Elasticsearch cannot do aggregations while indexing, but rather that your nodes cannot handle the load of those two operations happen in parallel.

Before starting to optimize you should try to find out, where exactly the bottleneck is. When indexing is stopping, this could potentially be a garbage collection (check your logfiles for that) - which makes sense, if you aggregation creates a lot of buckets, then you require some memory for those. And CPU is required for an aggregation as well.

You could configure your indices (those you write and those you read) to be put on different nodes using [shard allocation filtering](https://www.elastic.co/guide/en/elasticsearch/reference/current/shard-allocation-filtering.html).

--Alex

---

<div class="post-metadata">

**Author:** ![sharon.c](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sharon.c/32/16076_2.png) [@sharon.c](https://discuss.elastic.co/u/sharon.c)\
**Post date:** [June 21, 2016, 6:30pm UTC](https://discuss.elastic.co/t/elasticsearch-aggregation-performance-issue-index-allocation/52927/3 "2016-06-21T18:30:46Z")

</div>

Thanks, Alex  
I used shard allocation filtering to set the ip address filtering like this  
PUT test/\_settings  
{  
"index.routing.allocation.include.\_ip": "192.168.2.\*"  
}  
It works pretty well, it reallocated the shards accordingly to the specified nodes.

So I further looked into similar setting of node.box\_type [https://www.elastic.co/blog/hot-warm-architecture](https://www.elastic.co/blog/hot-warm-architecture)

PUT test/\_settings  
{  
"index.routing.allocation.require.box\_type" : "warm"  
}

Elasticsearch response is {"acknowledged":true}, but it does not reallocate the index shards to the warm type nodes.  
Looks like after I set "node.box\_type: warm" in elasticsearch.yml, elasticsearch does not pick up the settings.

Does anyone have the same issue setting "warm" "hot" box\_type ?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 21, 2016, 6:44pm UTC](https://discuss.elastic.co/t/elasticsearch-aggregation-performance-issue-index-allocation/52927/4 "2016-06-21T18:44:31Z")

</div>

Do you have replicas configured,meaning that both nodes all data? If this is the case, I do not believe shard allocation awareness is going to help much as you only have 2 data nodes and indexing need to be performed on both primary and replica shards.

---

<div class="post-metadata">

**Author:** ![sharon.c](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sharon.c/32/16076_2.png) [@sharon.c](https://discuss.elastic.co/u/sharon.c)\
**Post date:** [June 21, 2016, 9:48pm UTC](https://discuss.elastic.co/t/elasticsearch-aggregation-performance-issue-index-allocation/52927/5 "2016-06-21T21:48:05Z")

</div>

I do not have replicas. The index setting is like this:

> 

curl -XPUT '[http://datanode2:9200/test](http://datanode2:9200/test)' -d '{  
"settings" : {  
"index" : {  
"number\_of\_shards" : 3 ,  
"number\_of\_replicas" : 0  
}  
}  
}'

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:41pm UTC](https://discuss.elastic.co/t/elasticsearch-aggregation-performance-issue-index-allocation/52927/6 "2017-07-05T22:41:37Z")

</div>


