# Sort MLT query results: spark + scala

**URL:** <https://discuss.elastic.co/t/sort-mlt-query-results-spark-scala/307782>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [June 21, 2022, 4:35pm UTC](https://discuss.elastic.co/t/sort-mlt-query-results-spark-scala/307782 "2022-06-21T16:35:45Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![nobre0](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nobre0/32/98640_2.png) [@nobre0](https://discuss.elastic.co/u/nobre0)\
**Post date:** [June 21, 2022, 4:35pm UTC](https://discuss.elastic.co/t/sort-mlt-query-results-spark-scala/307782/1 "2022-06-21T16:35:45Z")

</div>

Hi,

I'm trying to run the 'More Like This' (MLT) query using the [apache spark connector](https://github.com/elastic/elasticsearch-hadoop#apache-spark).  
The problem is that the result is not sorted by computed MLT score. I think it is related to the `sort=_doc` parameter added in the [query builder](https://github.com/elastic/elasticsearch-hadoop/blob/4a518581b7ee909f89172190ff6683c0054303d9/mr/src/main/java/org/elasticsearch/hadoop/rest/SearchRequestBuilder.java#L206).

the code is the following:

```auto
val localSpark = SparkSession
    .builder()
    .appName("teste")
    .config("spark.es.nodes", "localhost")
    .config("spark.es.port", "9200")
    .config("es.mapping.id", "id")
    .config("es.write.operation", "upsert")
    .config("spark.es.nodes.wan.only", "true") 
    .config("es.scroll.size", 15)
    .master("local").getOrCreate()

val query = """{"query" : {"more_like_this": { "fields": ["text"], "like": [{"_index": "documents", "_id": "1234"}]}}}"""

val df = localSpark.read.format("org.elasticsearch.spark.sql").option("query", query).option("pushdown", "true").load("documents")

```

setup:

- java: 1.8
- spark : 3.1.0
- scala: 2.12.12
- "elasticsearch-spark-30" % "8.2.2"

---

<div class="post-metadata">

**Author:** ![Keith\_Massey](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/keith_massey/32/83666_2.png) [@Keith\_Massey](https://discuss.elastic.co/u/Keith_Massey)\
**Post date:** [June 22, 2022, 7:34pm UTC](https://discuss.elastic.co/t/sort-mlt-query-results-spark-scala/307782/2 "2022-06-22T19:34:31Z")

</div>

I was able to reproduce this problem. The default sort of `_doc` makes sense, since that is the most efficient way for a scroll to pull back data. But I thought that maybe adding a sort field to it like this would work:

```auto
val query = """{"sort":"_score","query" : {"more_like_this": { "fields": ["text"], "like": [{"_index": "documents", "_id": "1234"}],"min_term_freq": 1,"min_doc_freq": 1}}}"""

```

Unfortunately it looks like that sort is silently ignored and the results are still ordered by `_doc`.. It looks like a bug. You can probably sort the results by `_score` on the spark side, but that is not going to perform as well if you have a very large amount of data.

---

<div class="post-metadata">

**Author:** ![nobre0](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nobre0/32/98640_2.png) [@nobre0](https://discuss.elastic.co/u/nobre0)\
**Post date:** [June 22, 2022, 8:59pm UTC](https://discuss.elastic.co/t/sort-mlt-query-results-spark-scala/307782/3 "2022-06-22T20:59:02Z")

</div>

In my case, getting all results and then sorting by `_score ` is impracticable.  
I will open an issue since it is probably a bug.

Thanks for your response.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 20, 2022, 8:59pm UTC](https://discuss.elastic.co/t/sort-mlt-query-results-spark-scala/307782/4 "2022-07-20T20:59:23Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
