# How can I rerank query results by using Euclidean distance on fields have datatype is vector in elasticsearch?

**URL:** <https://discuss.elastic.co/t/how-can-i-rerank-query-results-by-using-euclidean-distance-on-fields-have-datatype-is-vector-in-elasticsearch/191319>\
**Category:** Elasticsearch\
**Created:** [July 19, 2019, 5:53am UTC](https://discuss.elastic.co/t/how-can-i-rerank-query-results-by-using-euclidean-distance-on-fields-have-datatype-is-vector-in-elasticsearch/191319 "2019-07-19T05:53:16Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![gia\_huy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gia_huy/32/50507_2.png) [@gia\_huy](https://discuss.elastic.co/u/gia_huy)\
**Post date:** [July 19, 2019, 5:53am UTC](https://discuss.elastic.co/t/how-can-i-rerank-query-results-by-using-euclidean-distance-on-fields-have-datatype-is-vector-in-elasticsearch/191319/1 "2019-07-19T05:53:17Z")

</div>

I have a question about `Elasticsearch`. Namely, I have some data about **embedding vectors** _(dense vector)_ and their corresponding **string tokens** from a algorithm using K-Means to map them from high-dimensionality vector space into smaller subspace (text format) for full-text search engine Elasticsearch to fast query _(Similarity searching)_.

And then I will get the results from Elasticsearch query phase to rescore (or rerank) it with Euclidean distance.

But this rescoring phase seems not working, results after **rescoring** lose similarities from query phase.

Here is my _request body (json)_ for `query` and `rescore` with Elasticsearh:

```auto
	request_body_1 = {
		"size": s,
		"query": {
			"function_score": {
				"functions": string_tokens_body,
				"score_mode": "sum",
				"boost_mode": "replace"
			}
		},
		"rescore": {
			"window_size": r, # Get top-r results from query phase for rescoring with Eucliean distance.
			"query": {
				"rescore_query": {
					"function_score": {
						"script_score": {
							"script": {
								"lang": "painless",
								"source": """
									def sum = 0.0 ;
									for (def index = 0; index < params['_source']['embedding_vector'].length; index++) {
										sum += Math.pow(params.query_vector[index] - doc['embedding_vector'][index], 2);
									}
									return(Math.sqrt(sum));
								""",
								"params": {
									"query_vector": query_vector.tolist() # numpy array not working here.
								}
							}
						},
						"boost_mode": "replace"
					}
				},
				"query_weight": 0, # Remove scores from query phase.
				"rescore_query_weight": 1 # Just calculate scores according to *rescoring phase*.
			}
		}
	}

```

Here is an example my document for indexing to Elasticsearch:

```auto
{
"index": "my_project",
"type": "_doc",
"id": 1,
"source": {
"embedding_vector": [1.12, 2.24, 3,34, 4,45],
"other_field": "other_datatypes"
}
}

```

How can I solve this problem ?

Thanks in advance for any reply of you.

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [July 23, 2019, 2:26pm UTC](https://discuss.elastic.co/t/how-can-i-rerank-query-results-by-using-euclidean-distance-on-fields-have-datatype-is-vector-in-elasticsearch/191319/2 "2019-07-23T14:26:05Z")

</div>

From 7.3, we have `cosineSimilarity` [function](https://www.elastic.co/guide/en/elasticsearch/reference/master/query-dsl-script-score-query.html#vector-functions) available for a special field type `dense_vector`. For 7.4 `l1norm` and `l2norm` (euclidean distance) will be available as well.

If you want to use euclidean distance in the elasticsearch before that, then indeed you need to design a script something like you are doing it. One thing to note here is that it is incorrect to access `doc['embedding_vector'][index]` in script, if your field `embedding_vector` is a simple numeric field. Even if you index its values as an array in your json, inside the index the values will be stored as multiple values in a sorted way. Thus, for example, `doc['embedding_vector'][3]` will return you `4` instead of your expected `34`. For the correct behaviour, you can instead parse the source as you are doing it: `params['_source']['embedding_vector'][index]`, but this will be slower .

About your specific question about rescoring, can you elaborate more what did you mean by "results after rescoring lose similarities from query phase"? Does it mean that it looks like rescoring phase is not applied at all?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 20, 2019, 2:26pm UTC](https://discuss.elastic.co/t/how-can-i-rerank-query-results-by-using-euclidean-distance-on-fields-have-datatype-is-vector-in-elasticsearch/191319/3 "2019-08-20T14:26:12Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
