# Failure on document\_parsing\_exception - dot\_product similarity on dense\_vector index field

**URL:** <https://discuss.elastic.co/t/failure-on-document-parsing-exception-dot-product-similarity-on-dense-vector-index-field/346718>\
**Category:** Elasticsearch\
**Tags:** vector-search\
**Created:** [November 8, 2023, 2:45pm UTC](https://discuss.elastic.co/t/failure-on-document-parsing-exception-dot-product-similarity-on-dense-vector-index-field/346718 "2023-11-08T14:45:35Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![ORipalta](https://avatars.discourse-cdn.com/v4/letter/o/e9a140/32.png) [@ORipalta](https://discuss.elastic.co/u/ORipalta)\
**Post date:** [November 8, 2023, 2:45pm UTC](https://discuss.elastic.co/t/failure-on-document-parsing-exception-dot-product-similarity-on-dense-vector-index-field/346718/1 "2023-11-08T14:45:36Z")

</div>

I'm trying to use dot-plot similarity on Elasticsearch. But after creating the index, the data/rows fail to load due to `document_parsing_exception`.

The error message that is being returned is `failed to parse: The [dot_product] similarity can only be used with unit-length vectors.`.

Not sure why this is happening, I've double checked the dimensions and attempted to cast each vector point as `float16` and `float32` but still won't work. For reference this is how one of my vector points looks like: `-0.29541016`

The model I'm using to create the embeddings is `msmarco-distilbert-base-dot-prod-v3` and I am running my code in Python.

This is the index I currently have.

```auto
index_mapping = {
    "properties": {
        "title": {
            "type": "text"
        },
        "abstract": {
            "type": "text"
        },
        "doc_len": {
            "type": "long"
        },
        "vector": {
            "type": "dense_vector",
            "dims": 768,
            "index": True,
            "similarity": "dot_product"
        }
    }
}

```

Any help will be greatly appreciated.

---

<div class="post-metadata">

**Author:** ![Carlos\_D](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/carlos_d/32/126245_2.png) [@Carlos\_D](https://discuss.elastic.co/u/Carlos_D)\
**Post date:** [November 8, 2023, 3:19pm UTC](https://discuss.elastic.co/t/failure-on-document-parsing-exception-dot-product-similarity-on-dense-vector-index-field/346718/2 "2023-11-08T15:19:25Z")

</div>

Hi @ORipalta !

For using `dot_product` similarity, all your vectors must be of unit length (see [dense\_vector parameters](https://www.elastic.co/guide/en/elasticsearch/reference/current/dense-vector.html#dense-vector-params)):

> When `element_type` is `float` , all vectors must be unit length, including both document and query vectors.

You're probably missing [normalization](https://www.khanacademy.org/computing/computer-programming/programming-natural-simulations/programming-vectors/a/vector-magnitude-normalization) for your vectors.  
In case you don't have normalized vectors, you can use `cosine` similarity, which will be less efficient.

---

<div class="post-metadata">

**Author:** ![ORipalta](https://avatars.discourse-cdn.com/v4/letter/o/e9a140/32.png) [@ORipalta](https://discuss.elastic.co/u/ORipalta)\
**Post date:** [November 8, 2023, 4:59pm UTC](https://discuss.elastic.co/t/failure-on-document-parsing-exception-dot-product-similarity-on-dense-vector-index-field/346718/3 "2023-11-08T16:59:56Z")

</div>

Thank you Carlos for your response!

I'm a bit new to embeddings so I just want to ask some followup questions if that's okay.

What does exactly mean "of unit length"?

I've also read the normalisation article which was very useful. The idea that I'm getting is that I have to normalise every single vector point inside the vector array using:

```auto
magnitude = magnitude(vector array)

for each single vector point in the vector array:
    single vector = single vector / magnitude

```

---

<div class="post-metadata">

**Author:** ![Carlos\_D](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/carlos_d/32/126245_2.png) [@Carlos\_D](https://discuss.elastic.co/u/Carlos_D)\
**Post date:** [November 8, 2023, 5:58pm UTC](https://discuss.elastic.co/t/failure-on-document-parsing-exception-dot-product-similarity-on-dense-vector-index-field/346718/4 "2023-11-08T17:58:18Z")

</div>

"of unit length" means the magnitude for each vector is 1. That's the end result of normalization - every vector will have magnitude 1.

The process you mention for normalization is the one you should take.

Keep in mind that you can use `cosine` similarity without the need for normalizing your vectors. That might be good for a first approach, and you'll skip the normalization process.

---

<div class="post-metadata">

**Author:** ![Thijsvdp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thijsvdp/32/146696_2.png) [@Thijsvdp](https://discuss.elastic.co/u/Thijsvdp)\
**Post date:** [November 8, 2023, 11:10pm UTC](https://discuss.elastic.co/t/failure-on-document-parsing-exception-dot-product-similarity-on-dense-vector-index-field/346718/5 "2023-11-08T23:10:05Z")

</div>

It is important to note though that if you are able to do vector normalization during the preprocessing steps it is preferable. As per documentation cosine similarity is slower than dot-product, because cosine similarity normalizes the vectors on the fly.

---

<div class="post-metadata">

**Author:** ![ORipalta](https://avatars.discourse-cdn.com/v4/letter/o/e9a140/32.png) [@ORipalta](https://discuss.elastic.co/u/ORipalta)\
**Post date:** [November 13, 2023, 9:58pm UTC](https://discuss.elastic.co/t/failure-on-document-parsing-exception-dot-product-similarity-on-dense-vector-index-field/346718/6 "2023-11-13T21:58:00Z")

</div>

Does it matter what similarity method do you use? I've read that some models are specifically trained with `cosine` and others with `dot-plot`. Therefore, is it safe to assume that one should be using the respective similarity method for optimal results?

On top of this I've also read this in a blog:

> Blockquote _Also,_ _models tuned for cosine-similarity will prefer the retrieval of short documents, while models tuned for dot-product will prefer the retrieval of longer documents__._

Is this something that holds true across all or most models?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 11, 2023, 9:59pm UTC](https://discuss.elastic.co/t/failure-on-document-parsing-exception-dot-product-similarity-on-dense-vector-index-field/346718/7 "2023-12-11T21:59:01Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
