# Top K results in RDD (beginner question)

**URL:** <https://discuss.elastic.co/t/top-k-results-in-rdd-beginner-question/911>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [May 19, 2015, 3:20pm UTC](https://discuss.elastic.co/t/top-k-results-in-rdd-beginner-question/911 "2015-05-19T15:20:29Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![tal](https://avatars.discourse-cdn.com/v4/letter/t/54ee81/32.png) [@tal](https://discuss.elastic.co/u/tal)\
**Post date:** [May 19, 2015, 3:20pm UTC](https://discuss.elastic.co/t/top-k-results-in-rdd-beginner-question/911/1 "2015-05-19T15:20:29Z")

</div>

Hey,  
I just started using ES and spark, and i'm trying to get only the k best results by score, for a really simple query on simple documents (just id and content text)  
How can i populate the RDD with only the top K results for that query? (i tried using the size option but since it's just page size and it seems that the rdd creation runs over all pages i still get all the results)

it would be great if anyone can point me in the right direction,  
is it top\_hits ? (even though i don't need any bucketing or anything like that?)

thanks!

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [May 26, 2015, 5:12am UTC](https://discuss.elastic.co/t/top-k-results-in-rdd-beginner-question/911/2 "2015-05-26T05:12:59Z")

</div>

Aggregations (like `top_hits`) are not yet supported by the Elasticsearch Hadoop connector. Most likely they will be added in 2.2 as the 2.1 version already contains plenty of features and is on its way out.  
Potentially one can do `top_hits` in Spark but in an inefficient way as it will pull all the data from Elasticsearch first.

---

<div class="post-metadata">

**Author:** ![tal](https://avatars.discourse-cdn.com/v4/letter/t/54ee81/32.png) [@tal](https://discuss.elastic.co/u/tal)\
**Post date:** [May 26, 2015, 11:55am UTC](https://discuss.elastic.co/t/top-k-results-in-rdd-beginner-question/911/3 "2015-05-26T11:55:39Z")

</div>

I see, thanks for the answer!

---

<div class="post-metadata">

**Author:** ![tal](https://avatars.discourse-cdn.com/v4/letter/t/54ee81/32.png) [@tal](https://discuss.elastic.co/u/tal)\
**Post date:** [May 31, 2015, 2:10pm UTC](https://discuss.elastic.co/t/top-k-results-in-rdd-beginner-question/911/4 "2015-05-31T14:10:01Z")

</div>

Hey - a quick follow up:  
I'd like to sort by score in my spark code, for that i need the scores in my rdd, but I get it without metadata, how can i specify in the query that i would like to keep the \_score field in my results?

thanks!

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [June 1, 2015, 8:45am UTC](https://discuss.elastic.co/t/top-k-results-in-rdd-beginner-question/911/5 "2015-06-01T08:45:35Z")

</div>

The connector doesn't return any relevant score since the it relies on `scan-and-scroll` [1], that is it pulls the results as they appear from each shard. Having a score means the results need to be _globally_ sorted which kills parallelism.  
This will be addressed with the aggregation part where instead of "fan"-ing out the call, the connector will only make one global one.

[1] [https://www.elastic.co/guide/en/elasticsearch/guide/master/scan-scroll.html](https://www.elastic.co/guide/en/elasticsearch/guide/master/scan-scroll.html)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:28pm UTC](https://discuss.elastic.co/t/top-k-results-in-rdd-beginner-question/911/6 "2017-07-06T13:28:21Z")

</div>


