# Why "knn\_query" doesn’t have a separate k parameter?

**URL:** <https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727>\
**Category:** Elasticsearch\
**Tags:** vector-search\
**Created:** [May 3, 2024, 5:08pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727 "2024-05-03T17:08:39Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![alliswell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alliswell/32/117514_2.png) [@alliswell](https://discuss.elastic.co/u/alliswell)\
**Post date:** [May 3, 2024, 5:08pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/1 "2024-05-03T17:08:39Z")

</div>

Thanks elastic team, for adding [knn\_query](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-knn-query.html). Now with knn\_query it is possible to use `function_score` or `script_score` and influence the rank of vector search result and not solely rely on the cosine similarity score. This was not possible using [top\_level\_knn\_section](https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html).

However, I have a question regarding the design of `knn_query`: Why is there no separate `k` parameter? Managing kNN searches without a distinct `k` parameter seems both inconvenient and potentially less accurate.

**Drawbacks of Coupling `k` with `size`:**

1. **Vector Search Systems:** For pure vector\_search system if I want to retieve `k` results and using `size` and `from` parameter to perform pagination. I have to set my `num_candidates` and `number_of_shards` such that `num_candidates*number_of_shards` is close (or equal) to `k`. This setup disrupts the typical operation of HNSW, where `num_candidates` (equivalent to efSearch) is kept higher than `k` to enhance the accuracy of vector matches. This (num\_candidates\>k) setting compensates for potential accuracy losses due to the greedy and local nature of node traversal (during search operation) in HNSW.

2. **Hybrid Search Systems:** In most systems I've observed, there's a limit on the maximum vector matching candidates, controlled by `k`. Pagination is then handled with the `from` and `size` parameters. But, with `k` parameter coupled with `size` parameter, then as explained in above point, we need to use `num_candidates` and `number_of_shards` as a proxy to limit the vector matching candidates.

3. **Aggregation Results:** As [documented](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-knn-query.html#knn-query-aggregations), the final results from aggregations contain `num_candidates * number_of_shards` documents, which may not be ideal.

The potential for `knn_query` to enhance searches with `function_score`, `script_score`, and sub-searches is significant. I'm curious about the rationale behind not supporting a separate `k` parameter in `knn_query`, especially since we are already collecting a "num\_candidates" number of results from each shard. It seems feasible for the coordinator or manager node to simply prune the list to `k`.

Could you provide some insights into this design choice?

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [May 3, 2024, 8:01pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/2 "2024-05-03T20:01:57Z")

</div>

Thanks for your interesting questions and providing your insights.

- The decision to use `size` for `knn` query comes with how Elasticsearch queries are organized. They all return max `size` number of results. And `knn` query being one of them, operates the same way – returns `size` results.

- Not sure I understood your point about pagination (why k needs to be close to `num_candidates*number_of_shards `?) . I can see how deeper pagination, you need to increase `num_candidates` to make sure its' value \> `from + size`. Overall pagination with knn search is tricky, and the best way to implement it to retrieve all results you want and do pagination on your application level.

- About hybrid search, that's a fair point to allow the control of vector matching candidates. But here control can be done with different ways: 1) boosting 2)applying minimum `similarity` parameter for `knn` query . Are these controls not enough?

- Aggregation results, that's a fair point, and if your aggregations results need to reflect the precise counts of documents that are returned in a request that you can use the top level knn search. Very often we found that it is not needed though.

---

<div class="post-metadata">

**Author:** ![alliswell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alliswell/32/117514_2.png) [@alliswell](https://discuss.elastic.co/u/alliswell)\
**Post date:** [May 6, 2024, 3:50pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/3 "2024-05-06T15:50:47Z")

</div>

Thanks for looking into it @mayya🙏

> About hybrid search, that's a fair point to allow the control of vector matching candidates. But here control can be done with different ways: 1) boosting 2)applying minimum similarity parameter for knn query . Are these controls not enough?

We do have boosting and similarity threshold parameter. The good thing about vector similarity score (in our case cosine similarity) that, unlike BM25 score, its value reflects absolute degree of relevance. However, tuning similarity score is bit tricky because, it is very dependent on underline distribution of \<query, document\> score. Since in vector space, everyone is neighbor of everyone, and has a hedge, I believe that we should also restrict the vector matching candidates using `k`

(Above is my main concern, Below points are my secondary concern)

> Not sure I understood your point about pagination (why k needs to be close to num\_candidates\*number\_of\_shards ?).

What I mean by this, suppose the business requirement is we only want to fetch at max "k" vector matching candidates. The only way to do this in [knn-query](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-knn-query.html) is by adjusting `num_candidates` and `number_of_shards`. Such that num\_candidates\*number\_of\_shards is equal to k. (Strictly talking from HNSW perspective, using `num_candidates` without `k` could lead to drop in accuracy of matching candidates, let me know if you would like me to elaborate more on this).

> The decision to use size for knn query comes with how Elasticsearch queries are organized. They all return max size number of results. And knn query being one of them, operates the same way – returns size results.

This make sense. But I feel like in order to keep the process consistent, we are manipulating the intrinsic behavior of kNN search (whose job is to fetch k nearest neighbour). I think `size` and `k` could co-exist independently. Suppose I'm doing only knn-query, first we can retrieve k-nearest neighbor (lets call it `results`) and then we can return results[:size] document. In case of pagination (results[from: from+size]).

> I can see how deeper pagination, you need to increase num\_candidates to make sure its' value \> from + size. Overall pagination with knn search is tricky, and the best way to implement it to retrieve all results you want and do pagination on your application level.

Pagination is important for e-commerce search. Doing it at the application level might not be feasible. Additionally, we would be foregoing all the optimizations involved in pagination done by elasticsearch team.

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [May 6, 2024, 7:56pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/4 "2024-05-06T19:56:52Z")

</div>

> I believe that we should also restrict the vector matching candidates using `k`

Restricting vector matches to `k` size is a valid use-case and need. We can consider that, but remember the `k` size will still be per shard basis.

> business requirement is we only want to fetch at max "k" vector matching candidates

You should not use `knn_query` for this then. Indeed, limiting `num_candidates` could lead to a drop of accuracy.

> I think `size` and `k` could co-exist independently. Suppose I'm doing only knn-query, first we can retrieve k-nearest neighbor (lets call it `results` ) and then we can return results[:size] document. In case of pagination (results[from: from+size]).

I guess you can do pagination, if your `k` is big enough to cover all pages.

---

<div class="post-metadata">

**Author:** ![alliswell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alliswell/32/117514_2.png) [@alliswell](https://discuss.elastic.co/u/alliswell)\
**Post date:** [May 6, 2024, 8:32pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/5 "2024-05-06T20:32:17Z")

</div>

> [@mayya](#):
>
> Restricting vector matches to `k` size is a valid use-case and need. We can consider that, but remember the `k` size will still be per shard basis.

After taking `k` most similar results from each shard, would it further merges the result from all shard and then takes global top `k` matches? (similar to how it is done in [top-level-knn-search](https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#tune-approximate-knn-for-speed-accuracy))

```auto
To gather results, the kNN search API finds a `num_candidates` number of
approximate nearest neighbor candidates on each shard.
The search computes the similarity of these candidate vectors to the query
vector, selecting the `k` most similar results from each shard.
The search then merges the results from each shard to return the global top `k` nearest neighbors.

```

---

<div class="post-metadata">

**Author:** ![alliswell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alliswell/32/117514_2.png) [@alliswell](https://discuss.elastic.co/u/alliswell)\
**Post date:** [May 6, 2024, 9:18pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/6 "2024-05-06T21:18:35Z")

</div>

> [@mayya](#):
>
> > business requirement is we only want to fetch at max "k" vector matching candidates
> 
> You should not use `knn_query` for this then. Indeed, limiting `num_candidates` could lead to a drop of accuracy.

Isn't this is how majority of the hybrid search-system is designed, where we define max number of the vector matching candidates to be included. For example here is the wording of from [knn-query-in-hybrid-search-doc](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-knn-query.html#knn-query-in-hybrid-search)

```auto
Knn query can be used as a part of hybrid search, where knn query is
combined with other lexical queries. For example, the query below finds
documents with title matching mountain lake, and combines them with
the top 10 documents that have the closest image vectors to the query_vector. 

```

Query:

```auto
POST my-image-index/_search
{
  "size" : 3,
  "query": {
    "bool": {
      "should": [
        {
          "match": {
            "title": {
              "query": "mountain lake",
              "boost": 1
            }
          }
        },
        {
          "knn": {
            "field": "image-vector",
            "query_vector": [-5, 9, -12],
            "num_candidates": 10,
            "boost": 2
          }
        }
      ]
    }
  }
}

```

In above example, even though it says:

```auto
combines them with the top 10 documents that have the closest image vectors to the query_vector.

```

But actually it is combining `10*number_of_shards` documents. Here, if we could have provided support for `k` parameter, by setting its value as `10` we could have achieve what is mentioned in document.

- assuming `k` in knn-query, is changed to work in similar way as mentioned in [top\_level\_knn\_section](https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html).

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [May 7, 2024, 10:30am UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/7 "2024-05-07T10:30:43Z")

</div>

Indeed each shard combines its own 10 nearest vectors with 10 lexical matches to obtain the shard's 10 results. Then all shards send their 10 results to the coordinator to combine the results from all shards to obtain the global 10 results. So overall, indeed, `10*number_of_shards` of nearest vectors were considered..

> After taking `k` most similar results from each shard, would it further merges the result from all shard and then takes global top `k` matches?

For the standalone knn query, yes, that's what you get. But for hybrid search, that's may not be the case, you can get more that `k` nearest vectors is your size is big.

Will it still be useful to have a separate `k` parameter, if it is applicable only per shard base?

---

<div class="post-metadata">

**Author:** ![alliswell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alliswell/32/117514_2.png) [@alliswell](https://discuss.elastic.co/u/alliswell)\
**Post date:** [May 7, 2024, 8:19pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/8 "2024-05-07T20:19:02Z")

</div>

> [@mayya](#):
>
> Indeed each shard combines its own 10 nearest vectors with 10 lexical matches to obtain the shard's 10 results. Then all shards send their 10 results to the coordinator to combine the results from all shards to obtain the global 10 results. So overall, indeed, `10*number_of_shards` of nearest vectors were considered..

Here you mean shard collects results for each queries (eg: lexical, knn) and merges them first at shard level, and then sends them to coordinator node, which will merge each shard results and generates final result.

For examples, assume num\_candidates=2 and number\_of\_shards=2:

**Shard 1:**

- Lexical search results: [d1:10, d2:2] # each entry in the list is "document\_id:score"

- KNN search results (num\_candidates=2): [d1:0.1, d3:0.5]

**Shard 2:**

- Lexical search results: [d5:100, d6:50]

- KNN search results (num\_candidates=2): [d7:0.1, d8:0.5]

**Final Result:**

- Combined and sorted by the coordinator node: [d5:100, d6:50, d1:10.1, d2:2, d3: 0.5, d7: 0.1, d8: 0.5]

While for the top-level-knn it first finds a `num_candidates` number of  
approximate nearest neighbor candidates on each shard. Then select the `k` best results from each shard. Then merges the results from each shard to return the global top `k` nearest neighbors.

Is my understanding correct here?

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [May 8, 2024, 10:43am UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/9 "2024-05-08T10:43:52Z")

</div>

Yes, it is correct, as knn query operates on a shard level, even if we introduce `k` parameter, it will be applied on a shard level.  
As in your example, even `k=2`, globally we got 4 knn results.

* * *

The top level knn search has an extra phase where it first collects the global `k` results, this ensures that globally we always get `k` results regardless of number of shards.

---

<div class="post-metadata">

**Author:** ![alliswell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alliswell/32/117514_2.png) [@alliswell](https://discuss.elastic.co/u/alliswell)\
**Post date:** [May 8, 2024, 12:31pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/10 "2024-05-08T12:31:29Z")

</div>

Thanks for the explaining how merging works differently for knn-query and top-level-knn.

Coming back to

> Will it still be useful to have a separate k parameter, if it is applicable only per shard base?

## short response;

I think we should still keep support for separate `k` parameter. (even when at shard level). We should keep it for the very same reason why we have separate `k` parameter in top-level-knn. Because it not only helps during selecting top-k at coordinator node. But it is also use to select top-k within each shard.

## not so short response;

For a moment lets keep Elasticsearch and sharding out of picture and lets see how search for most nearest-neighbours for given query\_vector is performed in HNSW graph.

The HNSW graph consists of multiple layers (refer attached figure).

**Step 1:** The search begins by selecting a random entry point (current\_vertex) at the highest layer.

**Step 2:** Calculate the distance between the query\_vector and the current\_vertex, as well as all neighbors of the current\_vertex.

- **Case 1:** If the closest vertex is a neighbor of the current\_vertex, greedily select that vertex and make it the current\_vertex. If the current layer is not lowest layer goto step-2, othwerwise goto step-3

- **Case 2:** If the closest vertex is the current\_vertex and it’s not the lowest layer, move down one level and goto Step 2. If the current layer is the lowest layer, proceed to Step 3.

**Step 3:** Ultimately, the current\_vertex at the lowest layer is identified as the nearest neighbor for the given query\_vector.

 ![hnsw_search](https://us1.discourse-cdn.com/elastic/original/3X/f/d/fdf291633afc4ff3657ea6820f572c05ba620892.jpeg)

### reason

During the whole search process, after randomly selecting entry point we employed “greedy” strategy to select next current\_vertex. And at each step it is using only local information to make decision. Due to this the selected vertex might not be the closet neighbour to query\_vector. To overcome this HNSW suggests following:

Instead of finding only one nearest neighbour on each layer, the efSearch (i.e num\_candidates) closest nearest neighbours to the query\_vector are found and each of these neighbours is used as the entry point on the next layer.

reference =\> [link](https://towardsdatascience.com/similarity-search-part-4-hierarchical-navigable-small-world-hnsw-2aad4fe87d37)

Returning to the context of Elasticsearch, each shard has its own HSNW graph. Suppose the goal is to find the 10 nearest neighbors from each shard. We should use more than 10 entry points (i.e., num\_candidates/efSearch should be \> 10) to ensure more precise vector matches. To enable this, we should add support for k (which would be set to 10 in this case).

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [May 9, 2024, 1:51pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/11 "2024-05-09T13:51:57Z")

</div>

Thanks for your explanation. Good to know that `k` parameter on a shard level will still be useful, I will bring this to the team for discussion.

Currently `efSearch` is `num_candidates`, and `k` is `size` for `knn` query. For a standalone `knn` query, that's all you need. Only when you use hybrid search, I can see how it makes sense to have a separate `k` parameter.

---

<div class="post-metadata">

**Author:** ![alliswell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alliswell/32/117514_2.png) [@alliswell](https://discuss.elastic.co/u/alliswell)\
**Post date:** [May 9, 2024, 2:26pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/12 "2024-05-09T14:26:48Z")

</div>

> [@mayya](#):
>
> Only when you use hybrid search, I can see how it makes sense to have a separate `k` parameter.

Right, this is useful for hybrid search.  
thanks @mayya 🙏. Will look forward to more updates on this.

---

<div class="post-metadata">

**Author:** ![alliswell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alliswell/32/117514_2.png) [@alliswell](https://discuss.elastic.co/u/alliswell)\
**Post date:** [May 9, 2024, 3:10pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/13 "2024-05-09T15:10:36Z")

</div>

Just to confirm that we are on same page:

So after introducing separate k parameter in `knn-query`, now we will have `k*number_of_shards` total vector matching candidates, and this is because lexical and knn merging is happening at shard level (unlike `top-level-knn-section` where it happens at coordinator level)

Also, I was thinking is elastic planning to introduce function\_score/script\_score in `top-level-knn-section`? So that ranking of vector matching candidates could be influenced by some business logic (eg: click, purchase rate, etc) and not just rank it on the basis of similarity.

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [May 9, 2024, 5:00pm UTC](https://discuss.elastic.co/t/why-knn-query-doesn-t-have-a-separate-k-parameter/358727/14 "2024-05-09T17:00:08Z")

</div>

I've created an [issue](https://github.com/elastic/elasticsearch/issues/108473) for adding `k` parameter for `knn` query.

For the top level knn search, please submit a separate discuss topic or github issue.
