# ANN Search is super slow

**URL:** <https://discuss.elastic.co/t/ann-search-is-super-slow/344863>\
**Category:** Elasticsearch\
**Tags:** vector-search\
**Created:** [October 12, 2023, 2:37am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863 "2023-10-12T02:37:20Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![spliter2157](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spliter2157/32/126460_2.png) [@spliter2157](https://discuss.elastic.co/u/spliter2157)\
**Post date:** [October 12, 2023, 2:37am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/1 "2023-10-12T02:37:20Z")

</div>

Hello There,

Hello, I have a question regarding Elasticsearch vector search. Our vector index has 768 dimensions, and it contains 25,000,000 documents split into two indices. Here are the results from `/_cat/indices`:

```auto
health status index uuid pri rep docs.count docs.deleted store.size pri.store.size
green open kr-documents 8bCjsGAxQfC8CPPwnTKBCg 32 0 9587847 698158 182.8gb 182.8gb
green open us-documents Es86NloYSVK7xMN7XyzPZg 32 0 17183616 3021606 680.7gb 680.7gb

```

Here is vector property

```auto
"vector": {
    "type": "dense_vector",
    "dims": 768,
    "index": true,
    "similarity": "cosine"
}

```

I've executed the Python code below for vector search, but even after waiting for more than a minute, the search results are not returned, and I receive a ConnectionTimedOut error. I need to retry 6-8 times to get the search results.

```auto
try:
    es_resp: dict = ES_MODULE.search(
        index=target_index,
        query=query,
        knn={
            "field": "vector",
            "query_vector": embedding,
            "k": kwargs.get("k", 200), # It should be larger than 200
            "num_candidates": 200,
            "similarity": 0.8,
        },
        size=size,
        source=["patent_number", "country", "vector"],
    )
except ConnectionTimeout:
    LOGGER.warning(msg={"message": f"Connection Timed out(TRIED {retry_count} / 10)"})
else:
    break

```

I will also provide additional information related to the nodes. Each node has 32GB of RAM.

- Server 1
  - master
  - data\_content
  - data\_content
  - coordinate\_only

- Server 2
  - data\_content

The reason for specifying the RAM of each node as 32GB is that I assumed the required RAM capacity is 72GB according to the formula below, and there are three data\_content nodes. Therefore, I thought that each node would need 24GB.  
`num_vectors * 4 * (num_dimensions + 12)` [refer](https://www.elastic.co/guide/en/elasticsearch/reference/current/tune-knn-search.html)

If you could assist me with this issue, it would be greatly appreciated.

---

<div class="post-metadata">

**Author:** ![spliter2157](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spliter2157/32/126460_2.png) [@spliter2157](https://discuss.elastic.co/u/spliter2157)\
**Post date:** [October 12, 2023, 3:47am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/2 "2023-10-12T03:47:07Z")

</div>

Oh each nodes is operating on Docker Container

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 12, 2023, 6:37am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/3 "2023-10-12T06:37:32Z")

</div>

Which version is it?

---

<div class="post-metadata">

**Author:** ![spliter2157](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spliter2157/32/126460_2.png) [@spliter2157](https://discuss.elastic.co/u/spliter2157)\
**Post date:** [October 12, 2023, 7:01am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/4 "2023-10-12T07:01:39Z")

</div>

It's a very basic thing, but I forgot it. My apologies.  
It's 8.8.1

---

<div class="post-metadata">

**Author:** ![Thijsvdp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thijsvdp/32/146696_2.png) [@Thijsvdp](https://discuss.elastic.co/u/Thijsvdp)\
**Post date:** [October 12, 2023, 7:20am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/5 "2023-10-12T07:20:45Z")

</div>

I have been doing something similar lately. I have faced the same issues you are facing and found a solution to it by following the documentation on tuning for kNN performance.

A quick summary of what I am doing:

- Roughly 90.000.000 docs,
- Each docs has a Dense Vector of size 768,
- 5 nodes (32vCPU, 128GB each)

A couple of things that really made a difference:

- Make sure you have sufficient RAM ([Analyze index disk usage API | Elasticsearch Guide [8.10] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-disk-usage.html)). Also make sure you have some spare RAM for other processes. Also your system by default uses 50% for heap I believe. So my guess is that your 24GB per shard exceeds this threshold.
- Use forcemerge to reduce segments (I used 2 per shard). Merging segments really helped in query speeds! Also there are downsides to having few shards, so maybe finding some balance would be good here.
- Preloading the kNN index into memory ([Preloading data into the file system cache | Elasticsearch Guide [8.9] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/8.9/preload-data-to-file-system-cache.html)). This really helped a lot as well! But make sure when you do this the index can fit into your RAM.
- When you have sufficient RAM and have set up preloading the kNN index into memory, restart the cluster and rerun your experiment. Check IO on your cluster, when doing. When the kNN index does NOT fit into memory, you will see a high read IO.

There are a couple of challenges that lie ahead when you do these things, I noticed:

- (Heavy) indexing into this index will increase segments again making your queries slow again! I have not found a proper way to deal with this.
- Heavy operations on the index, like expensive queries or heavy indexing on the cluster (even in another index), may push out the kNN index from memory, making it terribly slow again! Yesterday I have started a thread on this: [Manage heavy indexing in kNN indexes](https://discuss.elastic.co/t/manage-heavy-indexing-in-knn-indexes/344840)

Hope this helps a bit!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 12, 2023, 8:34am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/6 "2023-10-12T08:34:08Z")

</div>

As there are a lot of performance improvements, could you upgrade to 8.10.3 and test that again?

---

<div class="post-metadata">

**Author:** ![spliter2157](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spliter2157/32/126460_2.png) [@spliter2157](https://discuss.elastic.co/u/spliter2157)\
**Post date:** [October 12, 2023, 9:09am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/7 "2023-10-12T09:09:48Z")

</div>

@dadoonet Thank you for your reply, i tried it. but still slow..

But, Increasing the RAM capacity to 64GB significantly alleviated the symptoms. It seems like @Thijsvdp hypothesis was correct! I will give it a try.  
Appreciate That!

---

<div class="post-metadata">

**Author:** ![Thijsvdp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thijsvdp/32/146696_2.png) [@Thijsvdp](https://discuss.elastic.co/u/Thijsvdp)\
**Post date:** [October 12, 2023, 10:34am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/8 "2023-10-12T10:34:09Z")

</div>

We are on 8.9.1, has there been a significant improvement in 8.10.3 vs 8.9.1 in terms of vector search?

---

<div class="post-metadata">

**Author:** ![Thijsvdp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thijsvdp/32/146696_2.png) [@Thijsvdp](https://discuss.elastic.co/u/Thijsvdp)\
**Post date:** [October 12, 2023, 10:35am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/9 "2023-10-12T10:35:38Z")

</div>

What is the number of shards and the number of segments currently in your setup?

---

<div class="post-metadata">

**Author:** ![spliter2157](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spliter2157/32/126460_2.png) [@spliter2157](https://discuss.elastic.co/u/spliter2157)\
**Post date:** [October 12, 2023, 10:51am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/10 "2023-10-12T10:51:37Z")

</div>

I specified 32 shards per node, and I didn't specify segments separately. Just checked, and we have 657 segments in "us-documents" and 73 segments in "kr-documents."

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 12, 2023, 11:03am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/11 "2023-10-12T11:03:02Z")

</div>

I saw that in the release notes: [What’s new in 8.10 | Elasticsearch Guide [8.10] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/8.10/release-highlights.html#enable_parallel_knn_search_across_segments)

---

<div class="post-metadata">

**Author:** ![Thijsvdp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thijsvdp/32/146696_2.png) [@Thijsvdp](https://discuss.elastic.co/u/Thijsvdp)\
**Post date:** [October 12, 2023, 11:16am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/12 "2023-10-12T11:16:46Z")

</div>

Alright thanks. So you can still try to preload the kNN index, that may help a lot I think. And if you want to still speed up you may decrease segments, but this can have negative effects as described above.

That said, I also noticed you use `cosine` distance. I would recommend using `dot` product. If you make sure you insert normalized vectors, then they are equivalent. You would then avoid doing the normalization on each search request over and over again.

---

<div class="post-metadata">

**Author:** ![Thijsvdp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thijsvdp/32/146696_2.png) [@Thijsvdp](https://discuss.elastic.co/u/Thijsvdp)\
**Post date:** [October 12, 2023, 11:18am UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/13 "2023-10-12T11:18:13Z")

</div>

Oh that is great! I was actually waiting for this feature, but must have missed it. Thanks a lot!

---

<div class="post-metadata">

**Author:** ![Andrew\_Mora](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrew_mora/32/125622_2.png) [@Andrew\_Mora](https://discuss.elastic.co/u/Andrew_Mora)\
**Post date:** [October 12, 2023, 12:53pm UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/14 "2023-10-12T12:53:00Z")

</div>

> [@spliter2157](#):
>
> Hello There,
> 
> Hello, I have a question regarding Elasticsearch vector search. Our vector index has 768 dimensions, and it contains 25,000,000 documents split into two indices. Here are the results from `/_cat/indices`:
> 
> ```auto
> health status index uuid pri rep docs.count docs.deleted store.size pri.store.size
> green open kr-documents 8bCjsGAxQfC8CPPwnTKBCg 32 0 9587847 698158 182.8gb 182.8gb
> green open us-documents Es86NloYSVK7xMN7XyzPZg 32 0 17183616 3021606 680.7gb 680.7gb
> 
> ```
> 
> Here is vector property
> 
> ```auto
> "vector": {
> "type": "dense_vector",
> "dims": 768,
> "index": true,
> "similarity": "cosine"
> }
> 
> ```
> 
> I've executed the Python code below for vector search, but even after waiting for more than a minute, the search results are not returned, and I receive a ConnectionTimedOut error. I need to retry 6-8 times to get the search results.
> 
> ```auto
> try:
> es_resp: dict = ES_MODULE.search(
> index=target_index,
> query=query,
> knn={
> "field": "vector",
> "query_vector": embedding,
> "k": kwargs.get("k", 200), # It should be larger than 200
> "num_candidates": 200,
> "similarity": 0.8,
> },
> size=size,
> source=["patent_number", "country", "vector"],
> )
> except ConnectionTimeout:
> LOGGER.warning(msg={"message": f"Connection Timed out(TRIED {retry_count} / 10)"})
> else:
> break
> 
> ```
> 
> I will also provide additional information related to the nodes. Each node has 32GB of RAM.
> 
> - Server 1
> - master
> - data\_content
> - data\_content
> - coordinate\_only
> 
> - Server 2
> - data\_content
> 
> The reason for specifying the RAM of each node as 32GB is that I assumed the required RAM capacity is 72GB according to the formula below, and there are three data\_content nodes. Therefore, I thought that each node would need 24GB.  
> `num_vectors * 4 * (num_dimensions + 12)` [refer](https://www.elastic.co/guide/en/elasticsearch/reference/current/tune-knn-search.html)
> 
> If you could assist me with this issue, it would be greatly appreciated.

It seems like you're encountering ConnectionTimedOut errors in your Elasticsearch vector search. Given your data size and setup, optimizing the Elasticsearch cluster for improved performance, possibly increasing the RAM per node, and tweaking the timeout settings might help resolve this issue. AC Football Cases.

---

<div class="post-metadata">

**Author:** ![spliter2157](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spliter2157/32/126460_2.png) [@spliter2157](https://discuss.elastic.co/u/spliter2157)\
**Post date:** [October 25, 2023, 1:38pm UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/15 "2023-10-25T13:38:02Z")

</div>

Setting up pre-load as you suggested was incredibly helpful. A query that used to take 2-3 minutes now executes in under 1 second. You saved my ass. Thank you

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 22, 2023, 1:38pm UTC](https://discuss.elastic.co/t/ann-search-is-super-slow/344863/16 "2023-11-22T13:38:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
