# Ngrams or dense vectors for similarity between arrays?

**URL:** <https://discuss.elastic.co/t/ngrams-or-dense-vectors-for-similarity-between-arrays/338929>\
**Category:** Elasticsearch\
**Tags:** vector-search\
**Created:** [July 21, 2023, 8:07am UTC](https://discuss.elastic.co/t/ngrams-or-dense-vectors-for-similarity-between-arrays/338929 "2023-07-21T08:07:30Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![kgeographer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kgeographer/32/26217_2.png) [@kgeographer](https://discuss.elastic.co/u/kgeographer)\
**Post date:** [July 21, 2023, 8:07am UTC](https://discuss.elastic.co/t/ngrams-or-dense-vectors-for-similarity-between-arrays/338929/1 "2023-07-21T08:07:30Z")

</div>

I have an index field "names" holding an array of place names. My goal is an ES query that sends an array of place names and returns the most similar docs based on all the names in both the query array and the index arrays. If there is _any_ exact match between names, that doc should score highest. As I understand the ngram tokenizer, it produces a bag of tokens, which would not rank candidate matches properly in many cases - e.g. two names are quite similar to an index doc, but 4 others in the array are not. So I think the key is somehow merging name-to-names comparisons. Any ideas for how to approach this welcome - are pre-computed dense vectors a possible approach?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 18, 2023, 8:07am UTC](https://discuss.elastic.co/t/ngrams-or-dense-vectors-for-similarity-between-arrays/338929/2 "2023-08-18T08:07:54Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
