# Using painless scripting to re-implement BM25 scoring

**URL:** https://discuss.elastic.co/t/using-painless-scripting-to-re-implement-bm25-scoring/195802
**Category:** Elasticsearch
**Tags:** painless
**Created:** [August 19, 2019, 7:06pm UTC](https://discuss.elastic.co/t/using-painless-scripting-to-re-implement-bm25-scoring/195802 "2019-08-19T19:06:45Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![dpappas](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dpappas/32/52571_2.png) [@dpappas](https://discuss.elastic.co/u/dpappas)
#### Post date: [August 19, 2019, 7:06pm UTC](https://discuss.elastic.co/t/using-painless-scripting-to-re-implement-bm25-scoring/195802/1 "2019-08-19T19:06:45Z")

</div>

Hello everyone

I have indexed a few hundreds of millions of data into ElasticSearch using the default parameters of b and k1. It is prohibitive for me to reindex the data however i would like to optimize the parameters b and k1 of bm25 for better scoring.  
To my understanding there are some functions in Painless scripting that could compute/fetch tf and idf scores of a token in a document.  
Could you please reproduce BM25 in Painless scripting so that i could tune the b and k1 parameters?

Thank you all in advance  
Dimitris

---

<div class="post-metadata">

### Author: ![rjernst](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rjernst/32/6363_2.png) [@rjernst](https://discuss.elastic.co/u/rjernst)
#### Post date: [August 19, 2019, 10:10pm UTC](https://discuss.elastic.co/t/using-painless-scripting-to-re-implement-bm25-scoring/195802/2 "2019-08-19T22:10:15Z")

</div>

In Elasticsearch the module implementing textual scoring is called similarity. There isn't a need to write a painless script, since b and k1 can be [customized](https://www.elastic.co/guide/en/elasticsearch/reference/current/index-modules-similarity.html#index-modules-similarity). However if you are intent on using painless, you can write a [similarity script](https://www.elastic.co/guide/en/elasticsearch/reference/current/index-modules-similarity.html#scripted_similarity).

---

<div class="post-metadata">

### Author: ![dpappas](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dpappas/32/52571_2.png) [@dpappas](https://discuss.elastic.co/u/dpappas)
#### Post date: [August 19, 2019, 10:28pm UTC](https://discuss.elastic.co/t/using-painless-scripting-to-re-implement-bm25-scoring/195802/3 "2019-08-19T22:28:55Z")

</div>

Unfortunately this will not work.  
As i mentioned i do not want to reindex.  
I have hundreds of millions of data.  
i just need to optimize the values of b and k1.  
I just want to replicate bm25 with painless scripting so to just change thw values of b and k1 as needed.

Thank you again

---

<div class="post-metadata">

### Author: ![rjernst](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rjernst/32/6363_2.png) [@rjernst](https://discuss.elastic.co/u/rjernst)
#### Post date: [August 19, 2019, 11:27pm UTC](https://discuss.elastic.co/t/using-painless-scripting-to-re-implement-bm25-scoring/195802/4 "2019-08-19T23:27:16Z")

</div>

The similarity can be changed on an existing index, no reindexing necessary. What do you see implying reindexing would be needed?

---

<div class="post-metadata">

### Author: ![dpappas](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dpappas/32/52571_2.png) [@dpappas](https://discuss.elastic.co/u/dpappas)
#### Post date: [August 20, 2019, 9:44am UTC](https://discuss.elastic.co/t/using-painless-scripting-to-re-implement-bm25-scoring/195802/5 "2019-08-20T09:44:50Z")

</div>

I thought that "PUT /index ..." is a creation of an index.  
Dont i need to re-index the data once i change the mapping/settings ?

---

<div class="post-metadata">

### Author: ![rjernst](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rjernst/32/6363_2.png) [@rjernst](https://discuss.elastic.co/u/rjernst)
#### Post date: [August 21, 2019, 12:16am UTC](https://discuss.elastic.co/t/using-painless-scripting-to-re-implement-bm25-scoring/195802/6 "2019-08-21T00:16:33Z")

</div>

While some settings cannot be changed (eg `number_of_shards`) many settings can be, like the configured similarity. Use the [update settings api](https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-update-settings.html).

---

<div class="post-metadata">

### Author: ![rjernst](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rjernst/32/6363_2.png) [@rjernst](https://discuss.elastic.co/u/rjernst)
#### Post date: [August 21, 2019, 12:17am UTC](https://discuss.elastic.co/t/using-painless-scripting-to-re-implement-bm25-scoring/195802/7 "2019-08-21T00:17:45Z")

</div>

> Dont i need to re-index the data once i change the mapping/settings ?

Since the similarity settings are not baked into the index, changing these parameters does not requiring reindexing.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [September 18, 2019, 12:17am UTC](https://discuss.elastic.co/t/using-painless-scripting-to-re-implement-bm25-scoring/195802/8 "2019-09-18T00:17:49Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
