# Performance Issue with KNN + Filter on Large Index (v8.12)

**URL:** <https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279>\
**Category:** Elasticsearch\
**Tags:** vector-search\
**Created:** [April 18, 2025, 8:44am UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279 "2025-04-18T08:44:16Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Saleh\_AbuAli](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saleh_abuali/32/141444_2.png) [@Saleh\_AbuAli](https://discuss.elastic.co/u/Saleh_AbuAli)\
**Post date:** [April 18, 2025, 8:44am UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/1 "2025-04-18T08:44:16Z")

</div>

Hi,

I’m currently using KNN search in Elasticsearch (version 8.12) with the following setup:

- Query vector dimension: 384
- Index size: 200M+ documents
- k = 100, num\_candidates = 200
- Quantized vectors using 'int8\_hnsw' instead of float
- Preloaded vector fields: 'vex' and 'veq'

I referred to this documentation for optimization:  
🔗 [Tune approximate kNN search | Elastic Docs](https://www.elastic.co/docs/deploy-manage/production-guidance/optimize-performance/approximate-knn-search)

The main issue arises when **I apply filters in the KNN query** — the performance significantly degrades, and in some cases, I encounter **response timeout errors (\>60 seconds)**.

Is there any way to improve the performance of KNN search when using filters on such a large dataset? Any tuning recommendations, settings, or known limitations I should consider?

Thanks in advance!

---

<div class="post-metadata">

**Author:** ![Kathleen\_DeRusso](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kathleen_derusso/32/132039_2.png) [@Kathleen\_DeRusso](https://discuss.elastic.co/u/Kathleen_DeRusso)\
**Post date:** [April 18, 2025, 12:18pm UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/2 "2025-04-18T12:18:39Z")

</div>

Hey there @Saleh_AbuAli ,

It would be interesting to see more about specifically what filters you are applying, and how many documents they are matching. Also [profiling](https://www.elastic.co/docs/reference/elasticsearch/rest-apis/search-profile) the query could be interesting.

Without more information about your queries, I can offer some general advice:

- Is the size of your results page 100? If not, consider decreasing `k` to something smaller.
- Consider upgrading to a newer version of Elasticsearch. Since 8.12 there have been several performance optimizations that speed up performance of vector search.

---

<div class="post-metadata">

**Author:** ![Saleh\_AbuAli](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saleh_abuali/32/141444_2.png) [@Saleh\_AbuAli](https://discuss.elastic.co/u/Saleh_AbuAli)\
**Post date:** [April 18, 2025, 12:33pm UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/3 "2025-04-18T12:33:54Z")

</div>

Hi @Kathleen_DeRusso , **Thank you for your response.**  
The filters are restrictive, but the filtered documents still number more than 500,000.  
I apply some custom filtering on the returned results based on the score — that's why I set **K** to 100. After applying my custom filtration, I usually keep only 10 documents. Reducing **K** might lower the number of relevant results returned.  
Number of shards =40 and total number of segments =181.  
Could you please let me know what optimizations in the newer version help improve the performance of vector search?  
Example filter:

````auto
 "filter": [ {
        "range": {
          "publicationYear": {
            "gte": "2010",
            "lte": "2022"
          }
        }
      },
      {
        "bool": {
          "should": [
            {
              "terms": {
                "HostedInVenue.Venue.PublishedByInstitution.Institution.@id.pkg": [
                  "112321dssa"
                ]
              }
            }
          ]
        }
      } ```
````

---

<div class="post-metadata">

**Author:** ![Kathleen\_DeRusso](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kathleen_derusso/32/132039_2.png) [@Kathleen\_DeRusso](https://discuss.elastic.co/u/Kathleen_DeRusso)\
**Post date:** [April 18, 2025, 12:54pm UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/4 "2025-04-18T12:54:09Z")

</div>

This [blog](https://www.elastic.co/search-labs/blog/vector-search-improvements) is a bit outdated since it talks about 8.15.0 and we just released 8.18.0/9.0.0, but it talks about some of the enhancements we made. Since then we have also introduced and GA'd [BBQ](https://www.elastic.co/blog/whats-new-elastic-search-9-0-0) quantization as well as several other optimizations in Lucene and Elasticsearch.

---

<div class="post-metadata">

**Author:** ![Saleh\_AbuAli](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saleh_abuali/32/141444_2.png) [@Saleh\_AbuAli](https://discuss.elastic.co/u/Saleh_AbuAli)\
**Post date:** [April 18, 2025, 1:07pm UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/5 "2025-04-18T13:07:31Z")

</div>

Thanks for your help. I'm currently using int8\_hnsw quantization. Are there any recommended steps or any advice on what to add or change in the mapping or anything else?  
the mapping for the vector is  
"@vector": {  
"type": "dense\_vector",  
"dims": 384,  
"index": true,  
"similarity": "cosine",  
"index\_options": {  
"type": "int8\_hnsw",  
"m": 16,  
"ef\_construction": 100  
}  
}

---

<div class="post-metadata">

**Author:** ![Kathleen\_DeRusso](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kathleen_derusso/32/132039_2.png) [@Kathleen\_DeRusso](https://discuss.elastic.co/u/Kathleen_DeRusso)\
**Post date:** [April 18, 2025, 1:29pm UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/6 "2025-04-18T13:29:10Z")

</div>

Actually, a colleague pointed out to me that 40 shards is likely a culprit. For that many vectors you probably need a lot of RAM (maybe 80 GB)? How much RAM do you have?

Also, if you need to save space on your existing deployment you could consider switching to `int4_hnsw` which will halve the RAM requirements.

---

<div class="post-metadata">

**Author:** ![Saleh\_AbuAli](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saleh_abuali/32/141444_2.png) [@Saleh\_AbuAli](https://discuss.elastic.co/u/Saleh_AbuAli)\
**Post date:** [April 18, 2025, 1:58pm UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/7 "2025-04-18T13:58:48Z")

</div>

These are the details for the nodes I have, with information for each of them. Please note that I have another index on the same nodes. Can you please let me know if this amount of RAM is sufficient, or if I need to increase it?

| **Node Name** | **Role** | **Total RAM (GB)** | **Used RAM (GB)** | **Used %** | **Heap Used %** | **JVM Heap (GB)** |
| --- | --- | --- | --- | --- | --- | --- |
| `es-data-2` | Data | 66.57 | 60.12 | 90% | 36% | 34.36 |
| `es-master-0` | Master | 4.00 | 2.99 | 70% | 45% | 2.15 |
| `es-master-1` | Master | 4.00 | 2.85 | 66% | 57% | 2.15 |
| `es-data-5` | Data | 66.57 | 60.31 | 91% | 70% | 34.36 |
| `es-data-3` | Data | 66.57 | 60.37 | 91% | 75% | 34.36 |
| `es-data-0` | Data | 66.57 | 60.27 | 91% | 43% | 34.36 |
| `es-master-2` | Master | 4.00 | 2.79 | 65% | 35% | 2.15 |
| `es-data-1` | Data | 66.57 | 60.37 | 91% | 57% | 34.36 |
| `es-data-4` | Data | 66.57 | 60.36 | 91% | 39% | 34.36 |

### **Document Distribution** :

| Node Name | Total Documents |
| --- | --- |
| es-data-0 | 38,274,701 |
| es-data-1 | 0 |
| es-data-2 | 38,457,734 |
| es-data-3 | 38,351,625 |
| es-data-4 | 38,342,771 |
| es-data-5 | 38,401,115 |

---

<div class="post-metadata">

**Author:** ![Kathleen\_DeRusso](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kathleen_derusso/32/132039_2.png) [@Kathleen\_DeRusso](https://discuss.elastic.co/u/Kathleen_DeRusso)\
**Post date:** [April 18, 2025, 2:29pm UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/8 "2025-04-18T14:29:00Z")

</div>

Here's a good [reference](https://www.elastic.co/docs/deploy-manage/production-guidance/optimize-performance/approximate-knn-search#_ensure_data_nodes_have_enough_memory) on sizing.

Remember that for HNSW, we have to store everything in memory, so the back of the envelope formula is `num_vectors * (num_dimensions + 4)` - that's without replicas and also by itself without other indices on the same nodes. So by itself, ballpark around 80 GB. Assuming that your data nodes here are only 1 replica, that might be sufficient on its own but it really depends on what else is there besides the vectors. The [analyze disk usage API](https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-indices-disk-usage) could also be helpful here.

Beyond that it would be interesting to know what type of filters you're sending in that result in timeouts.

---

<div class="post-metadata">

**Author:** ![Saleh\_AbuAli](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saleh_abuali/32/141444_2.png) [@Saleh\_AbuAli](https://discuss.elastic.co/u/Saleh_AbuAli)\
**Post date:** [April 18, 2025, 2:48pm UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/9 "2025-04-18T14:48:55Z")

</div>

Thanks @Kathleen_DeRusso I tried to run the `disk_usage` query, but I got a timeout error every time, and there's no way to control the timeout value.

---

<div class="post-metadata">

**Author:** ![Saleh\_AbuAli](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saleh_abuali/32/141444_2.png) [@Saleh\_AbuAli](https://discuss.elastic.co/u/Saleh_AbuAli)\
**Post date:** [April 18, 2025, 3:00pm UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/10 "2025-04-18T15:00:45Z")

</div>

Here are sample of the filters that I applied :

````auto
{"knn":[{"field":"@vector","query_vector”:[queryvector,]”k”:100,"num_candidates":200,
"filter"
:[{"range":{"publicationYear":{"gte":"2010","lte":"2022"}}}],
{"bool":
{"must_not":[{"terms":{"HostedIn.Venue.@id":["data1","data2"]}}],
{"should":[{"terms":{"HostedIn.Venue.Publish.Institution.@id":["data3","data4"]}},{"terms":{"HostedIn.Venue.@id":["data6"]}}]}}]}]
}```
````

---

<div class="post-metadata">

**Author:** ![Kathleen\_DeRusso](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kathleen_derusso/32/132039_2.png) [@Kathleen\_DeRusso](https://discuss.elastic.co/u/Kathleen_DeRusso)\
**Post date:** [April 18, 2025, 5:30pm UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/11 "2025-04-18T17:30:17Z")

</div>

Nothing stands out to me as extremely expensive in the filters - if I had seen something like `function_score` in the filter I would say it needs to be optimized, but you're only doing `terms` and `range` filters. I suppose you could experiment and see if one of them is the root cause of the timeout (IDK, maybe caching or something at play there if your documents are frequently changing and these filters are hitting a ton of documents)..

You're not doing other expensive query operations like function score queries or aggs when you're timing out are you?

You could also look at how many segments you're searching and whether a [force merge](https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-indices-flush-2) helps.

You could try either increasing memory/nodes or going down to int4 quantization if you feel memory might be an issue based on the above information.

Also, going back to the smaller `k` that would definitely be faster, especially if your application didn't have to do custom post-filtering.

---

<div class="post-metadata">

**Author:** ![Bob\_Penrod](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bob_penrod/32/116702_2.png) [@Bob\_Penrod](https://discuss.elastic.co/u/Bob_Penrod)\
**Post date:** [April 25, 2025, 5:34pm UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/12 "2025-04-25T17:34:16Z")

</div>

Just a lurker on this thread, with a smaller data set 😄 — If RAM weren't so much a constraint, what sort of performance benefit might we expect from using `dot_product` instead of `cosine` similarity, provided we normalize our vectors?

---

<div class="post-metadata">

**Author:** ![Kathleen\_DeRusso](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kathleen_derusso/32/132039_2.png) [@Kathleen\_DeRusso](https://discuss.elastic.co/u/Kathleen_DeRusso)\
**Post date:** [April 28, 2025, 12:48pm UTC](https://discuss.elastic.co/t/performance-issue-with-knn-filter-on-large-index-v8-12/377279/13 "2025-04-28T12:48:33Z")

</div>

Newer versions of Elasticsearch (8.12+) will do that normalization for you under the hood! You can find out more in the [PR](https://github.com/elastic/elasticsearch/pull/99445).
