# Evaluating your search system's performance

**URL:** <https://discuss.elastic.co/t/evaluating-your-search-systems-performance/17064>\
**Category:** Elasticsearch\
**Created:** [April 17, 2014, 4:23pm UTC](https://discuss.elastic.co/t/evaluating-your-search-systems-performance/17064 "2014-04-17T16:23:19Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Andrew\_O\_Brien](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrew_o_brien/32/1642_2.png) [@Andrew\_O\_Brien](https://discuss.elastic.co/u/Andrew_O_Brien)\
**Post date:** [April 17, 2014, 4:23pm UTC](https://discuss.elastic.co/t/evaluating-your-search-systems-performance/17064/1 "2014-04-17T16:23:19Z")

</div>

What do you generally do to evaluate your search system's performance? Do  
you use a metrics based approach where they can compare how changes to  
scoring, analysis, or similarities effect hits in a quantitative way? Or  
something more manual?

Going through Intro to Information Retrieval[http://nlp.stanford.edu/IR-book/html/htmledition/evaluation-of-ranked-retrieval-results-1.html](http://nlp.stanford.edu/IR-book/html/htmledition/evaluation-of-ranked-retrieval-results-1.html),  
I see they pay a lot of attention to this and I can see the advantage of  
having that kind of feedback loop, but I haven't heard too many cases of  
this being used in practice.

For my own system, I've been looking to implement bpref (PDF; see Chapter  
3.1) [http://trec.nist.gov/pubs/trec16/appendices/measures.pdf](http://trec.nist.gov/pubs/trec16/appendices/measures.pdf), since I  
have fairly incomplete knowledge of which documents are relevant/irrelevant  
for my queries. I found that it would be helpful to be able to run a query  
and give some expected documents as parameters and just get back the ranks  
of those (I suppose I could implement this as a scan, but it would be nice  
to avoid the traffic). Any similar experiences?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/8c97dce1-1c77-47c8-8f2c-9488b2af4eaa%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/8c97dce1-1c77-47c8-8f2c-9488b2af4eaa%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [April 17, 2014, 5:19pm UTC](https://discuss.elastic.co/t/evaluating-your-search-systems-performance/17064/2 "2014-04-17T17:19:04Z")

</div>

Maybe MLR (machine based learning for ranking) is of some interest, if you  
do not know much about your document relevancy.

I use BM25 Okapi. For library catalogs, I have "document zones" like  
subject headings, title, author, identifiers and other supplemental texts  
like abstracts. All searches are on very short fields, fortunately. I am  
surrounded by librarians who are very skeptical that Elasticsearch can find  
"all the documents" they are looking for, they know what "relevancy" is. In  
the future I want to extend the catalog by linked open data, a real  
challenge for relevancy.

So BM25F [BM25 / BM25F Scoring · Issue #2388 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/2388) would  
be nice to have to tune document zone features to get the linkages into  
account. For now, I use a bit of field boosting and document boosting.

Jörg

On Thu, Apr 17, 2014 at 6:23 PM, Andrew O'Brien [obrien.andrew@gmail.com](mailto:obrien.andrew@gmail.com)wrote:

> What do you generally do to evaluate your search system's performance? Do  
> you use a metrics based approach where they can compare how changes to  
> scoring, analysis, or similarities effect hits in a quantitative way? Or  
> something more manual?
> 
> Going through Intro to Information Retrieval[http://nlp.stanford.edu/IR-book/html/htmledition/evaluation-of-ranked-retrieval-results-1.html](http://nlp.stanford.edu/IR-book/html/htmledition/evaluation-of-ranked-retrieval-results-1.html),  
> I see they pay a lot of attention to this and I can see the advantage of  
> having that kind of feedback loop, but I haven't heard too many cases of  
> this being used in practice.
> 
> For my own system, I've been looking to implement bpref (PDF; see Chapter  
> 3.1) [http://trec.nist.gov/pubs/trec16/appendices/measures.pdf](http://trec.nist.gov/pubs/trec16/appendices/measures.pdf), since I  
> have fairly incomplete knowledge of which documents are relevant/irrelevant  
> for my queries. I found that it would be helpful to be able to run a query  
> and give some expected documents as parameters and just get back the ranks  
> of those (I suppose I could implement this as a scan, but it would be nice  
> to avoid the traffic). Any similar experiences?
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/8c97dce1-1c77-47c8-8f2c-9488b2af4eaa%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/8c97dce1-1c77-47c8-8f2c-9488b2af4eaa%40googlegroups.com)[https://groups.google.com/d/msgid/elasticsearch/8c97dce1-1c77-47c8-8f2c-9488b2af4eaa%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/8c97dce1-1c77-47c8-8f2c-9488b2af4eaa%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoEKcr04z1%3D4fgzZOKJYeB4VoHWvPbWSgPZQizyz6nnPig%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoEKcr04z1%3D4fgzZOKJYeB4VoHWvPbWSgPZQizyz6nnPig%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:35am UTC](https://discuss.elastic.co/t/evaluating-your-search-systems-performance/17064/3 "2017-07-06T01:35:10Z")

</div>



---

<div class="post-metadata">

**Author:** ![cbuescher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cbuescher/32/60402_2.png) [@cbuescher](https://discuss.elastic.co/u/cbuescher)\
**Post date:** [April 23, 2018, 11:57am UTC](https://discuss.elastic.co/t/evaluating-your-search-systems-performance/17064/4 "2018-04-23T11:57:31Z")

</div>



---

<div class="post-metadata">

**Author:** ![cbuescher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cbuescher/32/60402_2.png) [@cbuescher](https://discuss.elastic.co/u/cbuescher)\
**Post date:** [April 23, 2018, 12:05pm UTC](https://discuss.elastic.co/t/evaluating-your-search-systems-performance/17064/5 "2018-04-23T12:05:13Z")

</div>

> [@Andrew\_O\_Brien](#):
>
> For my own system, I've been looking to implement bpref, since I have fairly incomplete knowledge of which documents are relevant/irrelevant for my queries.

Hi @Andrew_O_Brien,

I just found this post from over four years ago while searching for topics related to ranking evaluation, since we just recently introduced an (still experimental) [ranking evaluation API](https://www.elastic.co/guide/en/elasticsearch/reference/6.2/search-rank-eval.html) in Elasticsearch and I'm looking for other mertrics we might want to support. Did you end up implementing bpref? Did it works/was it useful? What challenges did you meet and was it necesarry to scroll all hits for a quiery until you found all the "relevant" documents or is it possible to use this measure only on a window of the top N hits.  
Please ignore me if this topic is no longer of interest to you, otherwise I'd be interested in your findings.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 4, 2022, 4:01am UTC](https://discuss.elastic.co/t/evaluating-your-search-systems-performance/17064/6 "2022-11-04T04:01:25Z")

</div>


