# Elasticsearch 8.17 → 8.19 upgrade: kNN now eagerly defaults "k = size", breaking aggregation-only queries

**URL:** <https://discuss.elastic.co/t/elasticsearch-8-17-8-19-upgrade-knn-now-eagerly-defaults-k-size-breaking-aggregation-only-queries/384628>\
**Category:** Elasticsearch\
**Tags:** vector-search\
**Created:** [January 19, 2026, 4:50pm UTC](https://discuss.elastic.co/t/elasticsearch-8-17-8-19-upgrade-knn-now-eagerly-defaults-k-size-breaking-aggregation-only-queries/384628 "2026-01-19T16:50:37Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![shekhar\_k](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shekhar_k/32/140791_2.png) [@shekhar\_k](https://discuss.elastic.co/u/shekhar_k)\
**Post date:** [January 19, 2026, 4:50pm UTC](https://discuss.elastic.co/t/elasticsearch-8-17-8-19-upgrade-knn-now-eagerly-defaults-k-size-breaking-aggregation-only-queries/384628/1 "2026-01-19T16:50:37Z")

</div>

I’m upgrading Elasticsearch from **8.17.3** to **8.19.10** and ran into a behavioural change with kNN + aggregations that breaks an existing use case.

What worked in **8.17.3:**

We use knn inside the query DSL (bool.must) together with size: 0, because we don’t need hits, only the aggregations.

In **8.17.3** , omitting k allowed aggregations to effectively run over all kNN candidates ( **num\_candidates** ) across shards, as long as they passed a similarity / min\_score threshold.  
This let us:

- keep response size small (size: 0)
- run large nested aggregations
- avoid artificially bounding results to top-k

What breaks in **8.19.10:**  
In **8.19.10** , the same query returns **empty aggregation buckets**.  
After investigation, it appears that:

- k is now **eagerly defaulted to size**
- with size: 0, this effectively becomes k = 0
- aggregations then run over **zero kNN hits**

Setting an explicit k fixes emptiness, but introduces a **hard top-k bound** (e.g. k ≤ 10k), which changes semantics for us:

- our previous queries aggregated over all candidates above a similarity threshold
- now they are strictly bounded to top-k neighbours

Our use case is **aggregation-heavy** (nested + reverse\_nested) and the kNN stage is only meant to define the candidate set, not to limit results to top-k.  
In practice, the number of documents above the similarity threshold can vary from **1k to \>1M** , and we need aggregations to reflect that set.

> **Question**

1. Is this behaviour change intentional (possibly related to “eager defaulting of k”)?

2. Is there a supported way in 8.19+ to:

> **Sample Snippet**

```auto
{
  "query": {
    "function_score": {
      "boost_mode": "replace",
      "functions": [
        {
          "script_score": {
            "script": {
              "source": "_score / params.total_boost",
              "params": {
                "total_boost": 1
              }
            }
          }
        }
      ],
      "min_score": 0.5,
      "query": {
        "bool": {
          "filter": [],
          "must": [
            {
              "knn": {
                "field": "embeddings_768_bgebase",
                "query_vector": [],
                "num_candidates": 3500,
                "boost": 1,
                "similarity": 0
              }
            }
          ]
        }
      }
    }
  },
  "aggs": {
    "total_hits_bucket": {
      "filter": {
        "match_all": {}
      },
      "aggs": {
        "score_filters": {
          "range": {
            "ranges": [
              {
                "from": 0.6
              }
            ],
            "script": {
              "source": "_score"
            }
          }
        }
      }
    }
  },
  "from": 0,
  "size": 0
}

```

---

<div class="post-metadata">

**Author:** ![Carlos\_D](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/carlos_d/32/126245_2.png) [@Carlos\_D](https://discuss.elastic.co/u/Carlos_D)\
**Post date:** [January 28, 2026, 6:18pm UTC](https://discuss.elastic.co/t/elasticsearch-8-17-8-19-upgrade-knn-now-eagerly-defaults-k-size-breaking-aggregation-only-queries/384628/2 "2026-01-28T18:18:17Z")

</div>

Hi @shekhar_k , sorry for the late reply.

> our previous queries aggregated over all candidates above a similarity threshold

To clarify, that is not what happened before. Aggregations are done over the top k documents from each shard - it just happened that when not specified, `k` defaulted to `num_candidates` before.

According to the [docs](https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-knn-query#knn-query-aggregations):

> `knn` query calculates aggregations on top `k` documents from each shard. Thus, the final results from aggregations contain `k * number_of_shards` documents.

Now that `k` defaults to `size`, you should be able to get the same results by specifying `k` as `num_candidates` to get the same behavior.

Hopefully that makes sense, and you can get back to your previous behavior!

---

<div class="post-metadata">

**Author:** ![Thijsvdp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thijsvdp/32/146696_2.png) [@Thijsvdp](https://discuss.elastic.co/u/Thijsvdp)\
**Post date:** [June 26, 2026, 11:03am UTC](https://discuss.elastic.co/t/elasticsearch-8-17-8-19-upgrade-knn-now-eagerly-defaults-k-size-breaking-aggregation-only-queries/384628/3 "2026-06-26T11:03:36Z")

</div>

Hi @Carlos_D ,

It seems actually to work different when kNN is defined at the root search level versus when you use the kNN inside the query section. Below two examples (we use 420 shards, and Elasticsearch 8.19.17).

**kNN at root level**

```auto
{
  "knn": {
    "field": "embeddings_768_bgebase",
    "query_vector": [...],
    "k": 1000,
    "num_candidates": 10000
  },
  "size": 0,
  "aggs": {
    "cnt": {
      "value_count": {
        "field": "family_id"
      }
    }
  }
}

// Response
{
  "took": 455,
  "timed_out": false,
  "_shards": {
    "total": 420,
    "successful": 420,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 1000,
      "relation": "eq"
    },
    "max_score": null,
    "hits": []
  },
  "aggregations": {
    "cnt": {
      "value": 1000
    }
  }
}

```

**kNN inside query**

```auto
{
  "query": {
    "knn": {
    "field": "embeddings_768_bgebase",
    "query_vector": [...],
    "k": 1000,
    "num_candidates": 10000
  }
  },
  "size": 0,
  "aggs": {
    "cnt": {
      "value_count": {
        "field": "family_id"
      }
    }
  }
}

// Response
{
  "took": 87,
  "timed_out": false,
  "_shards": {
    "total": 420,
    "successful": 420,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 10000,
      "relation": "gte"
    },
    "max_score": null,
    "hits": []
  },
  "aggregations": {
    "cnt": {
      "value": 420000
    }
  }
}

```

Findings:

- When kNN is used inside the root search, it seems that aggregations happen on the top k results. Defined by k, when k is set.
- When kNN is used inside the query clause, it seems that aggregations happen on the top k results per shard. This is now indeed consistent with how it worked before.

Is this behavior intentional, or a bug? Because I can see a use case to only aggregate on the top k results as well as the use case that @shekhar_k describes (which is more relevant when the similarity is provided).

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![BenTrent](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bentrent/32/33915_2.png) [@BenTrent](https://discuss.elastic.co/u/BenTrent)\
**Post date:** [June 26, 2026, 2:17pm UTC](https://discuss.elastic.co/t/elasticsearch-8-17-8-19-upgrade-knn-now-eagerly-defaults-k-size-breaking-aggregation-only-queries/384628/4 "2026-06-26T14:17:13Z")

</div>

> [@Thijsvdp](#):
>
> - When kNN is used inside the root search, it seems that aggregations happen on the top k results. Defined by k, when k is set.

Yep, its a "global top-k".

> When kNN is used inside the query clause, it seems that aggregations happen on the top k results per shard. This is now indeed consistent with how it worked before.

When used in the query clause, it is indeed per shard and by design.

---

<div class="post-metadata">

**Author:** ![Thijsvdp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thijsvdp/32/146696_2.png) [@Thijsvdp](https://discuss.elastic.co/u/Thijsvdp)\
**Post date:** [June 26, 2026, 2:49pm UTC](https://discuss.elastic.co/t/elasticsearch-8-17-8-19-upgrade-knn-now-eagerly-defaults-k-size-breaking-aggregation-only-queries/384628/5 "2026-06-26T14:49:39Z")

</div>

Alright thanks for the quick response!

---

<div class="post-metadata">

**Author:** ![Kishor\_Kharbas](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kishor_kharbas/32/143102_2.png) [@Kishor\_Kharbas](https://discuss.elastic.co/u/Kishor_Kharbas)\
**Post date:** [June 26, 2026, 7:11pm UTC](https://discuss.elastic.co/t/elasticsearch-8-17-8-19-upgrade-knn-now-eagerly-defaults-k-size-breaking-aggregation-only-queries/384628/6 "2026-06-26T19:11:15Z")

</div>

Replying to original question and adding to Carlos's comment. The limit for `num_candidates` has been `10K`, which is same as the limit for `K`. So explicitly setting K value will bring back the previous behavior.  
I fail to understand how are \> 10K documents evaluated for aggregation in 8.17 as per OP's comment.
