# Sampler aggregation performance vs 2 queries

**URL:** <https://discuss.elastic.co/t/sampler-aggregation-performance-vs-2-queries/112690>\
**Category:** Elasticsearch\
**Created:** [December 20, 2017, 5:52pm UTC](https://discuss.elastic.co/t/sampler-aggregation-performance-vs-2-queries/112690 "2017-12-20T17:52:15Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![adam3](https://avatars.discourse-cdn.com/v4/letter/a/dec6dc/32.png) [@adam3](https://discuss.elastic.co/u/adam3)\
**Post date:** [December 20, 2017, 5:52pm UTC](https://discuss.elastic.co/t/sampler-aggregation-performance-vs-2-queries/112690/1 "2017-12-20T17:52:15Z")

</div>

Ok so I am trying to increase performance of an Elasticsearch application.

For this case there is a single Elasticsearch node (non-sharded) with about 7M docs (index is like 40G).

We were doing a query to get top docs (like 1000) and then another query to do aggregation on some fields on those docs filter ids values (those 1000 docs) and then aggregate mutual\_information for those terms.

I had thought that doing a sampler aggregation would help by doing 1 query instead of 2 and that this would speed things up, but not seeing that currently.

Ok so here's the orig query:  
{  
"from" : 0,  
"size" : 1000,  
"query" : {  
"bool" : {  
"should" : [  
{  
"match" : {  
"name" : {  
"query" : "workout",  
"operator" : "AND",  
"prefix\_length" : 0,  
"max\_expansions" : 50,  
"fuzzy\_transpositions" : true,  
"lenient" : false,  
"zero\_terms\_query" : "NONE",  
"boost" : 1.0  
}  
}  
}  
],  
"disable\_coord" : false,  
"adjust\_pure\_negative" : true,  
"boost" : 1.0  
}  
},  
"min\_score" : 7.0,  
"\_source" : false,  
"sort" : [  
{  
"\_score" : {  
"order" : "desc"  
}  
},  
{  
"mau" : {  
"order" : "desc"  
}  
}  
]  
}

followed by:  
{  
"from" : 0,  
"size" : 1000,  
"query" : {  
"bool" : {  
"filter" : [  
{  
"ids" : {  
"type" : [],  
"values" : [  
"AWBNMzCn5eVrMnnV89Iw",  
.  
.  
. (lots of these)  
],  
"boost" : 1.0  
}  
}  
],  
"disable\_coord" : false,  
"adjust\_pure\_negative" : true,  
"boost" : 1.0  
}  
},  
"sort" : [  
{  
"\_score" : {  
"order" : "desc"  
}  
},  
{  
"mau" : {  
"order" : "desc"  
}  
}  
],  
"aggregations" : {  
"tracks" : {  
"significant\_terms" : {  
"field" : "tracks.raw",  
"size" : 1000,  
"min\_doc\_count" : 2,  
"shard\_min\_doc\_count" : 0,  
"mutual\_information" : {  
"include\_negatives" : false,  
"background\_is\_superset" : true  
}  
}  
}  
}  
}

vs:  
{  
"from" : 0,  
"size" : 0,  
"query" : {  
"bool" : {  
"should" : [  
{  
"match" : {  
"name" : {  
"query" : "workout",  
"operator" : "AND",  
"prefix\_length" : 0,  
"max\_expansions" : 50,  
"fuzzy\_transpositions" : true,  
"lenient" : false,  
"zero\_terms\_query" : "NONE",  
"boost" : 1.0  
}  
}  
}  
],  
"disable\_coord" : false,  
"adjust\_pure\_negative" : true,  
"boost" : 1.0  
}  
},  
"min\_score" : 7.0,  
"\_source" : false,  
"sort" : [  
{  
"\_score" : {  
"order" : "desc"  
}  
},  
{  
"mau" : {  
"order" : "desc"  
}  
}  
],  
"aggregations" : {  
"tracks" : {  
"sampler" : {  
"shard\_size" : 1000  
},  
"aggregations" : {  
"tracks" : {  
"significant\_terms" : {  
"field" : "tracks.raw",  
"size" : 1000,  
"min\_doc\_count" : 2,  
"shard\_min\_doc\_count" : 2,  
"mutual\_information" : {  
"include\_negatives" : false,  
"background\_is\_superset" : true  
}  
}  
}  
}  
}  
}  
}

Does that make sense?

Thanks,

Adam

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [December 20, 2017, 8:55pm UTC](https://discuss.elastic.co/t/sampler-aggregation-performance-vs-2-queries/112690/2 "2017-12-20T20:55:19Z")

</div>

I would have thought sampler agg should be faster. Maybe the bulk of the time is spent in the expensive loop that is common to both which is looking up the background frequency of all the terms found in matching docs. How many terms are in each doc?

---

<div class="post-metadata">

**Author:** ![adam3](https://avatars.discourse-cdn.com/v4/letter/a/dec6dc/32.png) [@adam3](https://discuss.elastic.co/u/adam3)\
**Post date:** [December 20, 2017, 9:44pm UTC](https://discuss.elastic.co/t/sampler-aggregation-performance-vs-2-queries/112690/3 "2017-12-20T21:44:02Z")

</div>

Mark,

There are a bunch of terms per doc.

I guess there are about 65 per document.

I can see some hot threads if that helps...

```
9/10 snapshots sharing following 29 elements
   org.elasticsearch.search.aggregations.bucket.significant.GlobalOrdinalsSignificantTermsAggregator.buildAggregation(GlobalOrdinalsSignificantTermsAggregator.java:104)
   org.elasticsearch.search.aggregations.bucket.significant.GlobalOrdinalsSignificantTermsAggregator$WithHash.buildAggregation(GlobalOrdinalsSignificantTermsAggregator.java:158)
   org.elasticsearch.search.aggregations.AggregatorFactory$MultiBucketAggregatorWrapper.buildAggregation(AggregatorFactory.java:147)
   org.elasticsearch.search.aggregations.bucket.DeferringBucketCollector$WrappedAggregator.buildAggregation(DeferringBucketCollector.java:96)
   org.elasticsearch.search.aggregations.bucket.BucketsAggregator.bucketAggregations(BucketsAggregator.java:116)
   org.elasticsearch.search.aggregations.bucket.sampler.SamplerAggregator.buildAggregation(SamplerAggregator.java:171)
   org.elasticsearch.search.aggregations.AggregationPhase.execute(AggregationPhase.java:139)
   org.elasticsearch.search.query.QueryPhase.execute(QueryPhase.java:114)
   org.elasticsearch.indices.IndicesService.lambda$loadIntoContext$16(IndicesService.java:1108)
   org.elasticsearch.indices.IndicesService$$Lambda$1799/1783108370.accept(Unknown Source)
   org.elasticsearch.indices.IndicesService.lambda$cacheShardLevelResult$18(IndicesService.java:1189)
   org.elasticsearch.indices.IndicesService$$Lambda$1803/1750113915.get(Unknown Source)
   org.elasticsearch.indices.IndicesRequestCache$Loader.load(IndicesRequestCache.java:160)
   org.elasticsearch.indices.IndicesRequestCache$Loader.load(IndicesRequestCache.java:143)
   org.elasticsearch.common.cache.Cache.computeIfAbsent(Cache.java:398)
   org.elasticsearch.indices.IndicesRequestCache.getOrCompute(IndicesRequestCache.java:116)
   org.elasticsearch.indices.IndicesService.cacheShardLevelResult(IndicesService.java:1195)
   org.elasticsearch.indices.IndicesService.loadIntoContext(IndicesService.java:1107)
   org.elasticsearch.search.SearchService.loadOrExecuteQueryPhase(SearchService.java:245)
   org.elasticsearch.search.SearchService.executeQueryPhase(SearchService.java:261)
   org.elasticsearch.action.search.SearchTransportService$6.messageReceived(SearchTransportService.java:331)
   org.elasticsearch.action.search.SearchTransportService$6.messageReceived(SearchTransportService.java:328)
   org.elasticsearch.transport.RequestHandlerRegistry.processMessageReceived(RequestHandlerRegistry.java:69)
   org.elasticsearch.transport.TransportService$7.doRun(TransportService.java:618)
   org.elasticsearch.common.util.concurrent.ThreadContext$ContextPreservingAbstractRunnable.doRun(ThreadContext.java:613)
   org.elasticsearch.common.util.concurrent.AbstractRunnable.run(AbstractRunnable.java:37)
   java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)
   java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)
   java.lang.Thread.run(Thread.java:748)

```

Thanks so much,

Adam

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [December 20, 2017, 11:24pm UTC](https://discuss.elastic.co/t/sampler-aggregation-performance-vs-2-queries/112690/4 "2017-12-20T23:24:32Z")

</div>

There are a number of execution modes when it comes to gathering terms. Try the "map" execution hint. See [https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-significantterms-aggregation.html#\_execution\_hint\_2](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-significantterms-aggregation.html#_execution_hint_2)

---

<div class="post-metadata">

**Author:** ![adam3](https://avatars.discourse-cdn.com/v4/letter/a/dec6dc/32.png) [@adam3](https://discuss.elastic.co/u/adam3)\
**Post date:** [December 21, 2017, 2:50pm UTC](https://discuss.elastic.co/t/sampler-aggregation-performance-vs-2-queries/112690/5 "2017-12-21T14:50:05Z")

</div>

Mark,

That tweak seems to have had a tremendous positive effect.

Thanks so much!!

Adam

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 18, 2018, 2:50pm UTC](https://discuss.elastic.co/t/sampler-aggregation-performance-vs-2-queries/112690/6 "2018-01-18T14:50:08Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
