# Sampler aggregation fails to optimize queries

**URL:** <https://discuss.elastic.co/t/sampler-aggregation-fails-to-optimize-queries/201946>\
**Category:** Elasticsearch\
**Created:** [October 2, 2019, 11:56am UTC](https://discuss.elastic.co/t/sampler-aggregation-fails-to-optimize-queries/201946 "2019-10-02T11:56:43Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![eladempow](https://avatars.discourse-cdn.com/v4/letter/e/958977/32.png) [@eladempow](https://discuss.elastic.co/u/eladempow)\
**Post date:** [October 2, 2019, 11:56am UTC](https://discuss.elastic.co/t/sampler-aggregation-fails-to-optimize-queries/201946/1 "2019-10-02T11:56:43Z")

</div>

I need to calculate some metrics for a dashboard view of a ~30gb index.  
As I understand it, the sampler aggregation can be used to perform faster calculations on a small sample of the data, but the performance I get is abysmal even for a very small sample size, which does not make sense and defeats the purpose of using the sampler aggregation.

Example (runs for 6 seconds):

```
GET index_prefix*/_search?size=0
{
  "aggs": {
    "sample": {
      "sampler": {
        "shard_size": 10
      },
      "aggs": {
        "last_month_events": {
          "filter": {
            "range": {
              "@timestamp": {
                "gte": "now-30d"
              }
            }
          }
        }
      }
    }
  }
}

```

Am I missing something here? Is it possible to achieve good performance for this query?

Thanks

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [October 4, 2019, 9:45am UTC](https://discuss.elastic.co/t/sampler-aggregation-fails-to-optimize-queries/201946/2 "2019-10-04T09:45:59Z")

</div>

The sampler aggregation gets the best-scoring docs. In your example request you have no query so there is no notion of "best" - it just iterates over _all_ docs in the index hoping to find the highest scoring docs (they will all score "1" in your example).  
You then filter this sample by your date range.

It would make more sense to use the search index and put your range criteria in the `query` part of the request. This would mean we'd only iterate over docs that match the criteria.

---

<div class="post-metadata">

**Author:** ![eladempow](https://avatars.discourse-cdn.com/v4/letter/e/958977/32.png) [@eladempow](https://discuss.elastic.co/u/eladempow)\
**Post date:** [October 6, 2019, 8:06am UTC](https://discuss.elastic.co/t/sampler-aggregation-fails-to-optimize-queries/201946/4 "2019-10-06T08:06:34Z")

</div>

Sorry, this is the correct query (a uniform sample of documents):

```
GET index_prefix*/_search?size=0
{
  "query": {
    "function_score": {
      "random_score": {}
    }
  },
  "aggs": {
    "sample": {
      "sampler": {
        "shard_size": 10
      },
      "aggs": {
        "last_month_events": {
          "filter": {
            "range": {
              "@timestamp": {
                "gte": "now-30d"
              }
            }
          }
        }
      }
    }
  }
}

```

The performance is still bad.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [October 6, 2019, 8:30am UTC](https://discuss.elastic.co/t/sampler-aggregation-fails-to-optimize-queries/201946/5 "2019-10-06T08:30:55Z")

</div>

My point re the date criteria being in the query part of the clause still stands.

---

<div class="post-metadata">

**Author:** ![eladempow](https://avatars.discourse-cdn.com/v4/letter/e/958977/32.png) [@eladempow](https://discuss.elastic.co/u/eladempow)\
**Post date:** [October 6, 2019, 8:46am UTC](https://discuss.elastic.co/t/sampler-aggregation-fails-to-optimize-queries/201946/7 "2019-10-06T08:46:02Z")

</div>

Using a simple count query by date range is still too slow (a few seconds).

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [October 7, 2019, 9:58am UTC](https://discuss.elastic.co/t/sampler-aggregation-fails-to-optimize-queries/201946/8 "2019-10-07T09:58:40Z")

</div>

What version of elasticsearch are you running?

---

<div class="post-metadata">

**Author:** ![eladempow](https://avatars.discourse-cdn.com/v4/letter/e/958977/32.png) [@eladempow](https://discuss.elastic.co/u/eladempow)\
**Post date:** [October 7, 2019, 10:34am UTC](https://discuss.elastic.co/t/sampler-aggregation-fails-to-optimize-queries/201946/9 "2019-10-07T10:34:23Z")

</div>

ELK 7.2.1  
(As far as I understand, the function score is calculated for the entire index during the aggregation, which takes a long time. Sampling a few documents uniformly should be faster.)

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [October 7, 2019, 10:35am UTC](https://discuss.elastic.co/t/sampler-aggregation-fails-to-optimize-queries/201946/10 "2019-10-07T10:35:49Z")

</div>

Thanks.

> [@eladempow](#):
>
> Using a simple count query by date range is still too slow (a few seconds).

Can you share this query JSON?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 4, 2019, 10:35am UTC](https://discuss.elastic.co/t/sampler-aggregation-fails-to-optimize-queries/201946/11 "2019-11-04T10:35:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
