# Weird behavior of dot\_product similarity on dense\_vector field

**URL:** <https://discuss.elastic.co/t/weird-behavior-of-dot-product-similarity-on-dense-vector-field/373007>\
**Category:** Elasticsearch\
**Tags:** vector-search\
**Created:** [January 9, 2025, 3:27pm UTC](https://discuss.elastic.co/t/weird-behavior-of-dot-product-similarity-on-dense-vector-field/373007 "2025-01-09T15:27:50Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![elaj](https://avatars.discourse-cdn.com/v4/letter/e/2acd7d/32.png) [@elaj](https://discuss.elastic.co/u/elaj)\
**Post date:** [January 9, 2025, 3:27pm UTC](https://discuss.elastic.co/t/weird-behavior-of-dot-product-similarity-on-dense-vector-field/373007/1 "2025-01-09T15:27:50Z")

</div>

Hello,

TLDR: What's the "indexation" difference between different similarities for `dense_vector` field?

I have an index with filed `dense_vector` defined with `similarity: cosine`.  
Now, I want to experiment with `similarity:dot_product`.

For that, I added two additional fields to my mappings, which will be populated via the ingest pipeline.

1. Original field `feature1_vector` with `similarity: cosine` -\> embeddings come from processor ` tag: inference 1` - calls custom trained ML model, hosted in ES
2. New field `feature2_vector_inference` with `similarity: dot_product` -\> embeddings come from processor ` tag: inference 2`, which is precisely the same as the previous one
3. New field `feature3_vector_copied` with `similarity: dot_product` -\> embeddings come from `set` processor, which copies them from the first field `feature1_vector`

**Mapping** :

```auto
"feature1_vector": { -- ORIGINAL FIELD
          "type": "dense_vector",
          "dims": 384,
          "index": true,
          "similarity": "cosine",
          "index_options": {
            "type": "int8_hnsw",
            "m": 16,
            "ef_construction": 100
          }
        },
"feature2_vector_inference": { -- NEW FIELD TO BE POPULATED BY INFERENCE PROCESSOR
          "type": "dense_vector",
          "dims": 384,
          "index": true,
          "similarity": "dot_product",
          "index_options": {
            "type": "int8_hnsw",
            "m": 16,
            "ef_construction": 100
          }
        },
"feature3_vector_copied": { -- NEW FIELD TO BE POPULATED BY SET PROCESSOR
          "type": "dense_vector",
          "dims": 384,
          "index": true,
          "similarity": "dot_product",
          "index_options": {
            "type": "int8_hnsw",
            "m": 16,
            "ef_construction": 100
          }
        }

```

**Ingest pipeline (shortened version):**

```auto
[ {
      "inference": {
        "tag": "inference 1",
        "model_id": "candidate_a",
        "input_output": [
          {
            "input_field": "feature1",
            "output_field": "feature1_vector"
          }
        ],
        "ignore_failure": false,
        "on_failure": [...]
      }
    },
    {
      "inference": {
        "tag": "inference 2",
        "model_id": "candidate_a",
        "input_output": [
          {
            "input_field": "feature1",
            "output_field": "feature2_vector_inference"
          }
        ],
        "ignore_failure": false,
        "on_failure": [...]
      }
    },
    {
      "set": {
        "field": "feature3_vector_copied",
        "copy_from": "feature1_vector"
      }
    }]

```

**After reindexing some documents, I see all three fields containing precisely the same embeddings (I reference to `_source` returned to my `knn` query).**

My questions:

1. Is that expected?
2. If so, why can't I select `similarity` during query time but only in index time?
3. Copying embeddings from `feature1_vector` to `feature3_vector_copied` with different similarities shouldn't simply fail?

Maybe someone will shed some light 🙂  
Thanks!

---

<div class="post-metadata">

**Author:** ![Kathleen\_DeRusso](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kathleen_derusso/32/132039_2.png) [@Kathleen\_DeRusso](https://discuss.elastic.co/u/Kathleen_DeRusso)\
**Post date:** [January 9, 2025, 4:08pm UTC](https://discuss.elastic.co/t/weird-behavior-of-dot-product-similarity-on-dense-vector-field/373007/2 "2025-01-09T16:08:04Z")

</div>

Starting with 8.12, cosine automatically normalizes vectors, and the dot product calculation is used out of the box as a performance enhancement. You can read more details in the [blog](https://www.elastic.co/blog/whats-new-elasticsearch-platform-8-12-0) or the [PR](https://github.com/elastic/elasticsearch/pull/99445) if you're interested.

---

<div class="post-metadata">

**Author:** ![elaj](https://avatars.discourse-cdn.com/v4/letter/e/2acd7d/32.png) [@elaj](https://discuss.elastic.co/u/elaj)\
**Post date:** [January 10, 2025, 11:54am UTC](https://discuss.elastic.co/t/weird-behavior-of-dot-product-similarity-on-dense-vector-field/373007/3 "2025-01-10T11:54:31Z")

</div>

Thank you for your reply.

May I have two follow-up questions?

1. Why keep both options `cosine` and `dot_product` for the `dense_vector.similarity` field if cosine internally uses `dot_product` for computing similarity (which is more efficient)? Are there other cases when defining `cosine` or `dot_products` makes actually a difference?
2. Would indexation be faster if I define explicitly in index `similarity: dot_product`?

Thank you 🙂

---

<div class="post-metadata">

**Author:** ![Kathleen\_DeRusso](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kathleen_derusso/32/132039_2.png) [@Kathleen\_DeRusso](https://discuss.elastic.co/u/Kathleen_DeRusso)\
**Post date:** [January 10, 2025, 1:10pm UTC](https://discuss.elastic.co/t/weird-behavior-of-dot-product-similarity-on-dense-vector-field/373007/4 "2025-01-10T13:10:36Z")

</div>

We would not remove a GA feature without warning to maintain backward compatibility. I think you'd have to benchmark to see if you saw a real difference in your models.
