# Very slow aggregation performance for trivial aggs

**URL:** <https://discuss.elastic.co/t/very-slow-aggregation-performance-for-trivial-aggs/24964>\
**Category:** Elasticsearch\
**Created:** [July 6, 2015, 4:09pm UTC](https://discuss.elastic.co/t/very-slow-aggregation-performance-for-trivial-aggs/24964 "2015-07-06T16:09:24Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![haizaar](https://avatars.discourse-cdn.com/v4/letter/h/3d9bf3/32.png) [@haizaar](https://discuss.elastic.co/u/haizaar)\
**Post date:** [July 6, 2015, 4:09pm UTC](https://discuss.elastic.co/t/very-slow-aggregation-performance-for-trivial-aggs/24964/1 "2015-07-06T16:09:24Z")

</div>

Good day,

I have ES 1.5.2, deployed over 15 nodes. Each node has 40Gb RAM, two 2.5Ghz CPU cores and a single SSD drive.

I want to store some "sensor data" for each machine _m_ and its sensors _s_. In each doc I capture machine\_id (m), sensor id and sub\_ids (s, s1, s2), timestamp (ts), and 4 values (v1..v4).

I've created 24 indexes with 15 shards each using the following mapping:

```
{ "sample": {
    "_all": {"enabled": False},
    "properties": {
        "m": {
            "type": "string",
            "index": "not_analyzed",
            "doc_values": True,
        },
        "s": {
            "type": "string",
            "index": "not_analyzed",
            "doc_values": True,
        },
        "ts": {
            "type": "date",
            "doc_values": True,
        },
        "ss1": {
            "type": "string",
            "index": "not_analyzed",
            "doc_values": True,
        },
        "ss1": {
            "type": "string",
            "index": "not_analyzed",
            "doc_values": True,
        },
        "v1": {
            "type": "long",
            "doc_values": True,
            "index": "no",
        },
        "v2": {
            "type": "long",
            "doc_values": True,
            "index": "no",
        },
        "v3": {
            "type": "long",
            "doc_values": True,
            "index": "no",
        },
        "v4": {
            "type": "long",
            "doc_values": True,
            "index": "no",
        }
    }
}
}

```

I've indexed 600 million docs into each index (using routing by _m_ field). My indexing speed goes up to 200,000 docs/sec which is very satisfying, but the aggregation performance is very slow.

I've created a single alias for all of these indices. Now consider this simple aggregation:

```
curl http://example.com:9200/myalias/_search?search_type=count -d {
    "aggs": {
        "1_min_date": {
            "min": {
                "field": "ts"
            }
        },
        "2_max_date": {
            "max": {
                "field": "ts"
            }
        }
    }
}

```

I run it against my idle cluster and it **took about 40 seconds to execute**. During execution, about half of the I see all of the nodes CPU is busy 100%, then only a single node is 100% CPU busy for the rest of the time probably trying to aggregate the results.

Why is it so slow? `ts` field is indexed, so from my understanding, finding minimal value for it is a matter of O(1) operation on each segment of every shard. I.e. the complexity should be O(number\_of\_shards). I have 360 shards, to its 360 lookups and then sorting an array with 360 members.

What am I doing wrong?

Thank you.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 7, 2015, 8:17am UTC](https://discuss.elastic.co/t/very-slow-aggregation-performance-for-trivial-aggs/24964/2 "2015-07-07T08:17:47Z")

</div>

> ts field is indexed ... finding minimal value for it is a matter of O(1) operation on each segment of every shard

What you are describing is theoretically the best speed attainable if you ignore everything else aggregations are designed to do.  
Let's look at what is missing from your example:

1. No query to subset the data (e.g. "find all records with IP address X")
2. No parent aggregation (e.g. a terms agg to group by "machine" field value)
3. No filtering (e.g. an alias restricting access to docs via a filter)

Now consider that the max/min etc "leaf" aggregations are designed to sit at the end of this chain of filters and look up values for each doc produced by this stream. There's no cheating or shortcuts in these implementations that look in the inverted index for the last value in the global terms enum - they always expect to be processing a stream of (potentially filtered) doc IDs and each value for each doc needs to be retrieved to determine the max. So your cost will be O(n) where n is the number of docs on a shard. Obviously when you are using DocValues OS-level file system caching is going to be a big help.

---

<div class="post-metadata">

**Author:** ![haizaar](https://avatars.discourse-cdn.com/v4/letter/h/3d9bf3/32.png) [@haizaar](https://discuss.elastic.co/u/haizaar)\
**Post date:** [July 16, 2015, 5:49pm UTC](https://discuss.elastic.co/t/very-slow-aggregation-performance-for-trivial-aggs/24964/3 "2015-07-16T17:49:07Z")

</div>

Mark,  
thanks for the detailed answer.

I understand the ElasticSearch way of doing this. Adding in some filterting and time-range bucketing that mimic my real application-to-be more closely, I brought query execution time below 400ms. The execution looks CPU bound still. I guess I'll get more speed if I'll run on cores with higher clocks.

I'm curious however if there is any efficient way to find min/max values over a large data sets in ElasticSearch?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 16, 2015, 6:16pm UTC](https://discuss.elastic.co/t/very-slow-aggregation-performance-for-trivial-aggs/24964/4 "2015-07-16T18:16:03Z")

</div>

You can get some of the info out of the field stats api  
[https://www.elastic.co/guide/en/elasticsearch/reference/current/search-field-stats.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-field-stats.html)

---

<div class="post-metadata">

**Author:** ![haizaar](https://avatars.discourse-cdn.com/v4/letter/h/3d9bf3/32.png) [@haizaar](https://discuss.elastic.co/u/haizaar)\
**Post date:** [July 19, 2015, 3:44pm UTC](https://discuss.elastic.co/t/very-slow-aggregation-performance-for-trivial-aggs/24964/5 "2015-07-19T15:44:15Z")

</div>

This is for ES 1.6 and on, right?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:00am UTC](https://discuss.elastic.co/t/very-slow-aggregation-performance-for-trivial-aggs/24964/6 "2017-07-06T00:00:32Z")

</div>


