# The use of \`size\` and \`from\` parameter in a \`knn\` search

**URL:** <https://discuss.elastic.co/t/the-use-of-size-and-from-parameter-in-a-knn-search/343962>\
**Category:** Elasticsearch\
**Tags:** vector-search\
**Created:** [September 27, 2023, 9:17am UTC](https://discuss.elastic.co/t/the-use-of-size-and-from-parameter-in-a-knn-search/343962 "2023-09-27T09:17:45Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![sbruinsje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sbruinsje/32/108568_2.png) [@sbruinsje](https://discuss.elastic.co/u/sbruinsje)\
**Post date:** [September 27, 2023, 9:17am UTC](https://discuss.elastic.co/t/the-use-of-size-and-from-parameter-in-a-knn-search/343962/1 "2023-09-27T09:17:45Z")

</div>

When doing a `knn` search there is a parameter `k` which specifies the number of best matching documents to return (approximately). You can also use the `size` parameter to determine the number of documents to return. Is there ever any reason why `k` and `size` would not be set to the same value?

And what about the `from` parameter, it can be used with knn searches, but I suppose that how it works is that if you have a `from` value of 10 then elasticsearch will compute the k+10 best matching documents and then return them all except the first 10. Is that correct? So the higher the `from` value the heavier the computation will be?

My questions are related to this query that I am using:

```auto
GET /author_index/_search
{
  "knn": {
    "field": "textEmbedding",
    "k": 20,
    "num_candidates": 1000,
    "query_vector": [...],
    "filter": {
      "bool": {
        "must": {
          "exists": {
            "field": "contactId"
          }
        }
      }
    }
  },
  "from": 0,
  "size": 20
}

```

---

<div class="post-metadata">

**Author:** ![Priscilla\_Parodi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/priscilla_parodi/32/43047_2.png) [@Priscilla\_Parodi](https://discuss.elastic.co/u/Priscilla_Parodi)\
**Post date:** [October 1, 2023, 2:00am UTC](https://discuss.elastic.co/t/the-use-of-size-and-from-parameter-in-a-knn-search/343962/2 "2023-10-01T02:00:28Z")

</div>

Hello @sbruinsje,

Good question.

As you mentioned, kNN search adds `k` matching documents to the search. So, if you set `k=20`, then it finds 20 matches. Then `size` takes the top scoring documents from these and returns them, it only tells how many hits should be returned in the response. In this case, considering your query, it will return the 20 matches.

It's a little confusing when you're doing only kNN search. But it fits with how the search API is designed.

The `from` parameter is the same idea but defining the number of hits to skip, it’s good to paginate search results.

Note: `from` and `size` parameters are optional.

---

<div class="post-metadata">

**Author:** ![Priscilla\_Parodi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/priscilla_parodi/32/43047_2.png) [@Priscilla\_Parodi](https://discuss.elastic.co/u/Priscilla_Parodi)\
**Post date:** [October 1, 2023, 2:03am UTC](https://discuss.elastic.co/t/the-use-of-size-and-from-parameter-in-a-knn-search/343962/3 "2023-10-01T02:03:05Z")

</div>

Now, to gather results, kNN search finds a `num_candidates` number of approximate nearest neighbor candidates on each shard. Elasticsearch collects `num_candidates` results from each shard, then merges them to find the top `k` results. So, if what you want is faster searches you can decrease `num_candidates`, but at the cost of potentially less accurate results.

---

<div class="post-metadata">

**Author:** ![BenTrent](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bentrent/32/33915_2.png) [@BenTrent](https://discuss.elastic.co/u/BenTrent)\
**Post date:** [October 2, 2023, 12:11pm UTC](https://discuss.elastic.co/t/the-use-of-size-and-from-parameter-in-a-knn-search/343962/4 "2023-10-02T12:11:07Z")

</div>

> [@sbruinsje](#):
>
> , but I suppose that how it works is that if you have a `from` value of 10 then elasticsearch will compute the k+10 best matching documents and then return them all except the first 10. Is that correct? So the higher the `from` value the heavier the computation will be?

This is not how it works.

We only look directly at the top `k` and `num_candidates` as of version 8.11 and earlier. I am unsure if this behavior will ever change.

If `k` is larger than `size`, we still search that many hits, but you will only retrieve `size`. In the scenario where `from>0` & `k` \> `from + size`, the results will indeed skip the first `from` elements and only return `size`.

But, we do NOT adjust `k` and `num_candidates` with those provided parameters.

---

<div class="post-metadata">

**Author:** ![sbruinsje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sbruinsje/32/108568_2.png) [@sbruinsje](https://discuss.elastic.co/u/sbruinsje)\
**Post date:** [October 4, 2023, 8:13am UTC](https://discuss.elastic.co/t/the-use-of-size-and-from-parameter-in-a-knn-search/343962/5 "2023-10-04T08:13:24Z")

</div>

Thanks for the replies!

I see, when combining a regular query with a knn search the combined results can be more than `k`, so `size` will then determine the actual returned number of results. In that case it may make sense to pick a different value for either of those.

And my conclusion about the computational costs: `from` doesn't affect the computational cost. The query/knn clauses are executed as is and sliced up by size/from at the very end.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 1, 2023, 8:13am UTC](https://discuss.elastic.co/t/the-use-of-size-and-from-parameter-in-a-knn-search/343962/6 "2023-11-01T08:13:41Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
