# Semantic search on more than 10k documents

**URL:** <https://discuss.elastic.co/t/semantic-search-on-more-than-10k-documents/343362>\
**Category:** Elasticsearch\
**Tags:** vector-search\
**Created:** [September 19, 2023, 12:25pm UTC](https://discuss.elastic.co/t/semantic-search-on-more-than-10k-documents/343362 "2023-09-19T12:25:53Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Denis\_Stefan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/denis_stefan/32/125801_2.png) [@Denis\_Stefan](https://discuss.elastic.co/u/Denis_Stefan)\
**Post date:** [September 19, 2023, 12:25pm UTC](https://discuss.elastic.co/t/semantic-search-on-more-than-10k-documents/343362/1 "2023-09-19T12:25:53Z")

</div>

Hello.

I am currently developing a semantic search solution and I have to work with more than 10k documents (more than the maximum number of candidates which is 10k for the kNN algorithm). I am trying to find a solution to search in whole index, not just 10k documents. Also I am interested in the way in which the 10k documents are selected for kNN.

So basically the question is how can I execute a semantic search on an index which has more than 10k documents and how the 10k candidates are selected?

Thank you.

---

<div class="post-metadata">

**Author:** ![Carlos\_D](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/carlos_d/32/126245_2.png) [@Carlos\_D](https://discuss.elastic.co/u/Carlos_D)\
**Post date:** [September 19, 2023, 1:59pm UTC](https://discuss.elastic.co/t/semantic-search-on-more-than-10k-documents/343362/2 "2023-09-19T13:59:30Z")

</div>

Hi @Denis_Stefan !

When performing kNN search, each index shard will perform an approximate neighbours search. Each of these searches will at most consider `num_candidates` as the top results. Thus, `num_candidates` does limit the total number of documents searched on each shard.

knn query uses a data structure ([HSNW](https://www.elastic.co/blog/introducing-approximate-nearest-neighbor-search-in-elasticsearch-8-0)) for making it easy to find the nearest neighbours to the vector that represents the search query. This is built to avoid having to compute the similarity to all vectors and retrieve the best possible candidates.

However, this is _approximate_. It offers a tradeoff between speed and precision, as we depend on the `num_candidates` searched. Increasing the `num_candidates` will increase precision (as more candidates will be considered) at the cost of search speed.

Once num\_candidates have been calculated on each shard, the top `k` will be selected on each shard. After that, the query will return the top `k` from all the shard results combined.

Keep in mind that you can also perform [exact kNN](https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#exact-knn) search so you can compare precision, recall and speed of the two approaches.

---

<div class="post-metadata">

**Author:** ![Denis\_Stefan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/denis_stefan/32/125801_2.png) [@Denis\_Stefan](https://discuss.elastic.co/u/Denis_Stefan)\
**Post date:** [September 19, 2023, 3:12pm UTC](https://discuss.elastic.co/t/semantic-search-on-more-than-10k-documents/343362/3 "2023-09-19T15:12:34Z")

</div>

Following your explanation, I fail to see the difference between the combination of k: 1 and num\_candidates: 1 and k:1 and num\_candidates:10 if all documents are taken into account. Shouldn't both combinations return same result? Could you provide a practical example?

---

<div class="post-metadata">

**Author:** ![Carlos\_D](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/carlos_d/32/126245_2.png) [@Carlos\_D](https://discuss.elastic.co/u/Carlos_D)\
**Post date:** [September 20, 2023, 7:44am UTC](https://discuss.elastic.co/t/semantic-search-on-more-than-10k-documents/343362/4 "2023-09-20T07:44:47Z")

</div>

Hi @Denis_Stefan !

My apologies - my response was incorrect in terms of total documents searched. I have edited it to avoid confusion to future readers.

On your example, `num_candidates: 1` will effectively search a single nearest neighbour on each shard, vs using `num_candidates: 10` that will look for 10 . So both examples would return the top candidate, but in the first case just 1 document will be considered.

We believe that approximate knn offers a good, tunable tradeoff for speed and precision. The HNSW data structure is specifically built for this use case.

However, you will need to experiment with `num_candidates` to find your speed / precision balance. Also, remember that you can use [exact knn](https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#exact-knn) for comparison or as your knn query method.

I hope this helps!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 18, 2023, 7:45am UTC](https://discuss.elastic.co/t/semantic-search-on-more-than-10k-documents/343362/5 "2023-10-18T07:45:32Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
