# To Interchange the Similarity Algorithm

**URL:** <https://discuss.elastic.co/t/to-interchange-the-similarity-algorithm/13354>\
**Category:** Elasticsearch\
**Created:** [August 27, 2013, 8:20pm UTC](https://discuss.elastic.co/t/to-interchange-the-similarity-algorithm/13354 "2013-08-27T20:20:07Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Fabio\_Figueiredo](https://avatars.discourse-cdn.com/v4/letter/f/a698b9/32.png) [@Fabio\_Figueiredo](https://discuss.elastic.co/u/Fabio_Figueiredo)\
**Post date:** [August 27, 2013, 8:20pm UTC](https://discuss.elastic.co/t/to-interchange-the-similarity-algorithm/13354/1 "2013-08-27T20:20:07Z")

</div>

Hi people,

I've read in ElasticSearch docs [1] that an index is bounded to a  
Similarity Algorithm (like BM25, DFR, TF/IDF or IB).

However, I'd like to know if would it be possible to specify a Similarity  
Algorithm at runtime because I have some evidences that to interchange the  
algorithm can bring better results if we use a machine learning technique  
to weight the score of each Similarity Algorithm based on the  
cirscunstances.

In other words, would it be possible, given a sample (ex.: first 500  
documents returned by BM25), to choose and run other Similarity Algorithm_s_just over those 500 documents instead of processing the whole index again?

[1] [http://www.elasticsearch.org/guide/reference/index-modules/similarity/](http://www.elasticsearch.org/guide/reference/index-modules/similarity/)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)\
**Post date:** [August 27, 2013, 8:48pm UTC](https://discuss.elastic.co/t/to-interchange-the-similarity-algorithm/13354/2 "2013-08-27T20:48:23Z")

</div>

you can configure similarity per field not per index. Yet, statistics for  
certain algos are written at index time so you can't change it at runtime  
you would need to index the field multiple times.

simon

On Tuesday, August 27, 2013 10:20:07 PM UTC+2, Fábio Figueiredo wrote:

> Hi people,
> 
> I've read in Elasticsearch docs [1] that an index is bounded to a  
> Similarity Algorithm (like BM25, DFR, TF/IDF or IB).
> 
> However, I'd like to know if would it be possible to specify a Similarity  
> Algorithm at runtime because I have some evidences that to interchange the  
> algorithm can bring better results if we use a machine learning technique  
> to weight the score of each Similarity Algorithm based on the  
> cirscunstances.
> 
> In other words, would it be possible, given a sample (ex.: first 500  
> documents returned by BM25), to choose and run other Similarity Algorithm\*  
> s\* just over those 500 documents instead of processing the whole index  
> again?
> 
> [1] [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/similarity/)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Israel\_Ekpo](https://avatars.discourse-cdn.com/v4/letter/i/49beb7/32.png) [@Israel\_Ekpo](https://discuss.elastic.co/u/Israel_Ekpo)\
**Post date:** [August 28, 2013, 1:40pm UTC](https://discuss.elastic.co/t/to-interchange-the-similarity-algorithm/13354/3 "2013-08-28T13:40:05Z")

</div>

Have you thought about putting the same documents in multiple indices (with  
different similarity algorithms) and then figure out which algorithm works  
best?

I am not sure how many documents you have but it is something to think  
about.

Though some of a computation is done at query time, there is a good chunk  
of preparatory work has to be done at index time for each algorithm so I  
don't think you can change it on the fly like the way you are proposing.

Take a look at this documentation [1] to get a better understanding of the  
inner workings of the Similarity feature

[1]  
[http://lucene.apache.org/core/4\_4\_0/core/org/apache/lucene/search/similarities/Similarity.html](http://lucene.apache.org/core/4_4_0/core/org/apache/lucene/search/similarities/Similarity.html)

[2]  
[http://lucene.apache.org/core/4\_4\_0/core/org/apache/lucene/search/similarities/TFIDFSimilarity.html](http://lucene.apache.org/core/4_4_0/core/org/apache/lucene/search/similarities/TFIDFSimilarity.html)

_Author and Instructor for the Upcoming Book and Lecture Series_  
_Massive Log Data Aggregation, Processing, Searching and Visualization with  
Open Source Software_  
_[http://massivelogdata.com](http://massivelogdata.com)_

On 27 August 2013 16:48, simonw [simon.willnauer@elasticsearch.com](mailto:simon.willnauer@elasticsearch.com) wrote:

> you can configure similarity per field not per index. Yet, statistics for  
> certain algos are written at index time so you can't change it at runtime  
> you would need to index the field multiple times.
> 
> simon
> 
> On Tuesday, August 27, 2013 10:20:07 PM UTC+2, Fábio Figueiredo wrote:
> 
> > Hi people,
> > 
> > I've read in Elasticsearch docs [1] that an index is bounded to a  
> > Similarity Algorithm (like BM25, DFR, TF/IDF or IB).
> > 
> > However, I'd like to know if would it be possible to specify a Similarity  
> > Algorithm at runtime because I have some evidences that to interchange the  
> > algorithm can bring better results if we use a machine learning technique  
> > to weight the score of each Similarity Algorithm based on the  
> > cirscunstances.
> > 
> > In other words, would it be possible, given a sample (ex.: first 500  
> > documents returned by BM25), to choose and run other Similarity Algorithm  
> > _s_ just over those 500 documents instead of processing the whole index  
> > again?
> > 
> > [1] [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/**guide/reference/index-modules/)\*\*  
> > similarity/[http://www.elasticsearch.org/guide/reference/index-modules/similarity/](http://www.elasticsearch.org/guide/reference/index-modules/similarity/)
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:19am UTC](https://discuss.elastic.co/t/to-interchange-the-similarity-algorithm/13354/4 "2017-07-06T02:19:17Z")

</div>


