# Custom TF-IDF implementation

**URL:** <https://discuss.elastic.co/t/custom-tf-idf-implementation/326872>\
**Category:** Elasticsearch\
**Created:** [March 2, 2023, 3:32pm UTC](https://discuss.elastic.co/t/custom-tf-idf-implementation/326872 "2023-03-02T15:32:29Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Karel\_Haerens1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/karel_haerens1/32/113590_2.png) [@Karel\_Haerens1](https://discuss.elastic.co/u/Karel_Haerens1)\
**Post date:** [March 2, 2023, 3:32pm UTC](https://discuss.elastic.co/t/custom-tf-idf-implementation/326872/1 "2023-03-02T15:32:29Z")

</div>

I'm trying to implement a custom TF-IDF-like algorithm with scripted similarity, my current approach:

The way term frequency is determined is custom, these values are precalculated and stored in the records as lists. as an example:

{  
"my\_text": "the apple falls",  
"my\_text\_counts": [10, 5, 7]  
}

in this document the array of numbers represents the term frequency for each word in the textfield.

I now want the similarity score to be the sum of the inverses of these values, if the corresponding word is in the query.

e.g. the query "the apple" would yield  
(1 \* 1/10 + 1 \* 1/5 + 0 \* 1 / 7)  
and "the falls"  
(1 \* 1/10 + 0 \* 1/5 + 1 \* 1 / 7)

After a lot of searching through the documentation I'm starting to think this is impossible with the current scripted similarity context.

Any tips or advice would be welcome

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 30, 2023, 3:32pm UTC](https://discuss.elastic.co/t/custom-tf-idf-implementation/326872/2 "2023-03-30T15:32:35Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
