# Optimizing using sample aggregation

**URL:** <https://discuss.elastic.co/t/optimizing-using-sample-aggregation/199934>\
**Category:** Elasticsearch\
**Created:** [September 18, 2019, 7:30am UTC](https://discuss.elastic.co/t/optimizing-using-sample-aggregation/199934 "2019-09-18T07:30:36Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Yonatan\_Omer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yonatan_omer/32/42450_2.png) [@Yonatan\_Omer](https://discuss.elastic.co/u/Yonatan_Omer)\
**Post date:** [September 18, 2019, 7:30am UTC](https://discuss.elastic.co/t/optimizing-using-sample-aggregation/199934/1 "2019-09-18T07:30:37Z")

</div>

I have an index where each doc contains all user's-weekly events.  
each event has messages.txt , which contain whole line strings, some very long (mapping bellow).  
The following query takes 15 seconds to return, the "sampler" aggregation does not help.  
Is there a way to limit the number of documents which are sent into the agg?

```auto
GET /users_weekly_events/_search
{
  "size": 0,
  "query": {...},
  "aggs": {
      "sample": {
          "sampler": {
             "shard_size": 5
          },
          "aggs": {
              "keywords": {
                  "terms": {
                    "field": "messages.txt.keyword",
                    "size": 5                 
                  }
              }
          }
      }
   }  
}

mappings of the text field

          "messages" : {
            "properties" : {
              "txt" : {
                "type" : "text",
                "fields" : {
                  "keyword" : {
                    "type" : "keyword",
                    "ignore_above" : 256
                  }
                }
              }
            }
          },

```

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 18, 2019, 9:47am UTC](https://discuss.elastic.co/t/optimizing-using-sample-aggregation/199934/2 "2019-09-18T09:47:00Z")

</div>

I have a number of questions:

What is the query? How many docs in the index? How many nodes?  
Isn't the `top_hits` aggregation on it's own more appropriate than the sampler/terms agg combo you're using here?

---

<div class="post-metadata">

**Author:** ![Yonatan\_Omer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yonatan_omer/32/42450_2.png) [@Yonatan\_Omer](https://discuss.elastic.co/u/Yonatan_Omer)\
**Post date:** [September 18, 2019, 10:30am UTC](https://discuss.elastic.co/t/optimizing-using-sample-aggregation/199934/3 "2019-09-18T10:30:26Z")

</div>

I have around 200k docs, 5 shards.  
Each doc contains an array with up to 10k strings (weekly events of this user)+ some stats (about 2 MB). the @timestamp is the beginning of the week time window.  
The query is different for every search, the query along perform very fast.  
So you suggest that I should use the 'top\_hits' to return the newest using the @timestamp fields?

```auto
GET /users_weekly_events/_search
{
  "size": 0,
  "query": {
    "bool": {
      "must": [ 
           {"term": {"osVersion": 10}, {"term": {"foo": "bar"} }
      ]
    }
  },
  "aggs": {
    "event_buckets": {
      "terms": {
        "field": "messages.txt.keyword",
        "include": "some-prefix-of-the-event.*",
        "size": 5
      }
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![Yonatan\_Omer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yonatan_omer/32/42450_2.png) [@Yonatan\_Omer](https://discuss.elastic.co/u/Yonatan_Omer)\
**Post date:** [September 18, 2019, 11:59am UTC](https://discuss.elastic.co/t/optimizing-using-sample-aggregation/199934/4 "2019-09-18T11:59:26Z")

</div>

Could you please show how to use the `top_hits` agg as a parent agg for the "event\_buckets" agg in my example above?  
My documents contain a `@timestamp` field, i could sort by that, but ideally would pick just random top

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 18, 2019, 12:42pm UTC](https://discuss.elastic.co/t/optimizing-using-sample-aggregation/199934/5 "2019-09-18T12:42:47Z")

</div>

Thanks for the additional info.

> [@Yonatan\_Omer](#):
>
> So you suggest that I should use the 'top\_hits'

What I don't like about large text fields as keywords is the index overheads and the arbitrary loss of data for those strings exceeding your `ignore_above` setting.  
It's hard to know what solution to suggest without a full grasp of what business question you're trying to answer

---

<div class="post-metadata">

**Author:** ![Yonatan\_Omer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yonatan_omer/32/42450_2.png) [@Yonatan\_Omer](https://discuss.elastic.co/u/Yonatan_Omer)\
**Post date:** [September 18, 2019, 1:29pm UTC](https://discuss.elastic.co/t/optimizing-using-sample-aggregation/199934/6 "2019-09-18T13:29:59Z")

</div>

I agree that long strings are an issue. I think the `top_hits` may help, but i need to clarify the syntax.  
my use case is as follows:  
event lines from log files are bucketed using their common string prefix.  
for each user (\_doc) i'm storing an array these event-id's, and an array of full events.  
The \_doc contain all the events during the passed week.  
Event ID's are used for significant-terms aggs ("find unusual events for a subset of all users"),which is working very well.  
What i'm trying to achieve is, for a given event-id, return the top 5 occurrences of the full message.

So for event ID `"CURL Failed with err code:"`

I would get : `["CURL Failed with err code:404", "CURL Failed with err code:123"...]`

The following query does the job, only that it takes too long.  
Is it possible to use `top_hits` to limit the number of docs which are sent to the second agg ?

```auto
GET /users_weekly_events/_search
{
  "size": 0,
  "query": {
    "bool": {
      "must": [ 
           {"term": {"osVersion": 10}, {"term": {"foo": "bar"} }
      ]
    }
  },
  "aggs": {
    "event_buckets": {
      "terms": {
        "field": "messages.txt.keyword",
        "include": "some-prefix-of-the-event.*",
        "size": 5
      }
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 19, 2019, 8:09am UTC](https://discuss.elastic.co/t/optimizing-using-sample-aggregation/199934/7 "2019-09-19T08:09:33Z")

</div>

I expect a more useful strategy might be to avoid aggregations based on keyword fields with large strings and instead use hashed versions of these strings. Obviously users will not be able to understand these values so you'd have to issue a second query to get the related full-text but it does mean you'd be dealing with shorter strings

---

<div class="post-metadata">

**Author:** ![Yonatan\_Omer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yonatan_omer/32/42450_2.png) [@Yonatan\_Omer](https://discuss.elastic.co/u/Yonatan_Omer)\
**Post date:** [September 19, 2019, 8:34am UTC](https://discuss.elastic.co/t/optimizing-using-sample-aggregation/199934/8 "2019-09-19T08:34:19Z")

</div>

Our current implementation is working very well, we now only wish to tune this feature. is it possible to use the output of `top_hits` as an input to the next aggregation in the pipeline?

And yes, I do consider to use hashes and store the long strings in another index, but in our use case it's not trivial:  
200k \_docs (users) each \_doc has 1k of `eventIDs`'s and 10k of distinct `messages`.  
`eventID` is a keyword and is always the prefix of each full `message`.

eventID used to query: `find unusual eventID for users having eventId=X`  
messages used to make `match_phrase` queries. messages.keyword are used to aggregate distinct messages.

I tried nested aggs, but we reached the 10000 nested objects limit...it was also very very slow

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 19, 2019, 8:37am UTC](https://discuss.elastic.co/t/optimizing-using-sample-aggregation/199934/9 "2019-09-19T08:37:05Z")

</div>

> [@Yonatan\_Omer](#):
>
> is it possible to use the output of `top_hits` as an input to the next aggregation in the pipeline?

No. Top hits is used as a leaf-node, generally to give more detail on the parent buckets discovered.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 17, 2019, 8:37am UTC](https://discuss.elastic.co/t/optimizing-using-sample-aggregation/199934/10 "2019-10-17T08:37:06Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
