# Pre-filter KNN vector search

**URL:** <https://discuss.elastic.co/t/pre-filter-knn-vector-search/341585>\
**Category:** Elasticsearch\
**Tags:** vector-search\
**Created:** [August 24, 2023, 3:21pm UTC](https://discuss.elastic.co/t/pre-filter-knn-vector-search/341585 "2023-08-24T15:21:18Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![veesahni](https://avatars.discourse-cdn.com/v4/letter/v/8e7dd6/32.png) [@veesahni](https://discuss.elastic.co/u/veesahni)\
**Post date:** [August 24, 2023, 3:21pm UTC](https://discuss.elastic.co/t/pre-filter-knn-vector-search/341585/1 "2023-08-24T15:21:18Z")

</div>

I'm trying to understand how to best implement a pre-filtered KNN vector search. To be clear: I'd like to pre-filter the total number of docs down to a smaller set, and then run a vector search that leverages a vector index over the results.

Documentation describes a [Filtered KNN Search](https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#knn-search-filter-example) with the following note: "The filter is applied **during** the approximate kNN search to ensure that `k` matching documents are returned."

The **during** phrase is unclear.

- Does this mean that it will apply a per-shared pre-filter?
- Or does this mean that it will do a KNN search on the all the data in the shard and then filter before the share returns results.. possibly getting more results if too many got filtered out?  
If a filter is applied to a KNN search, Is the underlying vector index still being used or it is reverting to brute force?

---

<div class="post-metadata">

**Author:** ![BenTrent](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bentrent/32/33915_2.png) [@BenTrent](https://discuss.elastic.co/u/BenTrent)\
**Post date:** [August 24, 2023, 7:01pm UTC](https://discuss.elastic.co/t/pre-filter-knn-vector-search/341585/2 "2023-08-24T19:01:59Z")

</div>

Hey @veesahni !

Apologies it was unclear. We can update the verbiage.

> - Does this mean that it will apply a per-shared pre-filter?

It means that it is a pre-filter applied while searching the individual HNSW graphs. So, it is a true pre-filter while still only searching for the kNN within the HNSW graph.

> Or does this mean that it will do a KNN search on the all the data in the shard and then filter before the share returns results..

No, it does not do that, it would be very expensive to do that.

The only time we dynamically switch to brute force is when the filter set is very restrictive (less than `num_candidates`).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 21, 2023, 7:02pm UTC](https://discuss.elastic.co/t/pre-filter-knn-vector-search/341585/3 "2023-09-21T19:02:05Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
