# Sparse vector vs rank features. Which one?

**URL:** <https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772>\
**Category:** Elasticsearch\
**Created:** [January 13, 2020, 7:32am UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772 "2020-01-13T07:32:19Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![snakeztc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/snakeztc/32/60686_2.png) [@snakeztc](https://discuss.elastic.co/u/snakeztc)\
**Post date:** [January 13, 2020, 7:32am UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/1 "2020-01-13T07:32:19Z")

</div>

I am trying to implement a customized search function that will rank documents based on sparse features. Concretely, each document will have a list of sparse features, e.g.

doc\_1 -\> {a: 1.1, b: 0.2}  
doc\_2 -\> {a: 0.2, z: 0.3, zz: 1.2} ...

Now at the query stage, the query will have a list of the sparse dimension that appears, e.g. [a, zz, b]. My scoring function is for each doc is simply added up all the value of terms that appear in both the query and the document. Take the above example,  
doc\_1\_score = 0.2 + 1.2  
doc\_2\_score = 1.1 + 0.2

My question is what is the most efficient way to implement this, I have two ideas now.

1. using rank features to save all the values in documents, and then using a list of "should" to get the score
2. save the features as sparse vectors, then using a dot\_product to get the final score.

Which one do you think will be more efficient (memory & speed). Is there better way to accomplish this, e.g. using inverted index? Thank you!

---

<div class="post-metadata">

**Author:** ![snakeztc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/snakeztc/32/60686_2.png) [@snakeztc](https://discuss.elastic.co/u/snakeztc)\
**Post date:** [January 13, 2020, 10:07pm UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/2 "2020-01-13T22:07:07Z")

</div>

Any one whom can help?

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [January 14, 2020, 9:34pm UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/3 "2020-01-14T21:34:25Z")

</div>

Hello there,  
we have deprecated `sparse_vector` datatype and you should not use it anymore. We did not see a good adoption for it.

Using `rank_features` seems to be a good alternative if you don't have that many features in your query to have a reasonable boolean query.

---

<div class="post-metadata">

**Author:** ![snakeztc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/snakeztc/32/60686_2.png) [@snakeztc](https://discuss.elastic.co/u/snakeztc)\
**Post date:** [January 15, 2020, 5:29am UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/4 "2020-01-15T05:29:00Z")

</div>

> [@mayya](#):
>
> Using `rank_features` seems to be a good alternative if you don't have that many features in your query to have a reasonable boolean query.

Thanks for the reply! What is considered as reasonable for boolean query features? Usually I will have 20-100 features in the query. Is that considered okay? Thanks!

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [January 15, 2020, 9:25pm UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/5 "2020-01-15T21:25:40Z")

</div>

20-100 features sounds reasonable. There is a [limit](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-settings.html) on the maximum number of clauses within a boolean query that should not be exceeded.

---

<div class="post-metadata">

**Author:** ![snakeztc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/snakeztc/32/60686_2.png) [@snakeztc](https://discuss.elastic.co/u/snakeztc)\
**Post date:** [January 17, 2020, 6:46pm UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/6 "2020-01-17T18:46:23Z")

</div>

Cool! Thanks.

(Newbie here) do you think this type of Boolean query that I am Interested can scale and maintain high speed for large index? I have more than 10-50 million documents in the index.

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [January 20, 2020, 12:02pm UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/7 "2020-01-20T12:02:33Z")

</div>

10-50 million docs is not a very big collection, but the performance of a query depends on many factors: for a single clause how many docs contain a particular feature, how many clauses you have in total etc. A good thing with rank features query is that it can efficiently skip non-competitive documents if you just need top N docs and don't need total hit count.

So the best advice for you is to index your collection and test the performance of your queries yourself.

---

<div class="post-metadata">

**Author:** ![snakeztc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/snakeztc/32/60686_2.png) [@snakeztc](https://discuss.elastic.co/u/snakeztc)\
**Post date:** [January 21, 2020, 7:36pm UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/8 "2020-01-21T19:36:23Z")

</div>

Thank you! I tested myself and the speed is not bad. Is there a place I can read more about how rank features are implemented? E.g what kind of data structure and how it skip non competitive docs etc.

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [January 22, 2020, 5:38pm UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/9 "2020-01-22T17:38:30Z")

</div>

@snakeztc  
About rank features, you can read more [here](https://www.elastic.co/guide/en/elasticsearch/reference/master/query-dsl-rank-feature-query.html)

We also have a [blog](https://www.elastic.co/blog/faster-retrieval-of-top-hits-in-elasticsearch-with-block-max-wand) devoted to the topic of skipping non-competitive hits.

Further details how rank features field is implemented can be found in Lucene code [here](https://github.com/apache/lucene-solr/blob/master/lucene/core/src/java/org/apache/lucene/document/FeatureField.java)

---

<div class="post-metadata">

**Author:** ![snakeztc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/snakeztc/32/60686_2.png) [@snakeztc](https://discuss.elastic.co/u/snakeztc)\
**Post date:** [February 8, 2020, 6:21am UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/10 "2020-02-08T06:21:56Z")

</div>

Hi I have implemented the rank\_feature as your suggested and it works great in general. I am using log-based rank\_feature with scaling factor = 1. I.e. y = log(x + 1)

one thing I notice that, the return score sometime is different from the actual correct score, if I manually compute it. In general the overall ranking is correct, however, there are cases where one doc should have slightly higher score, but it actually gets a lower score. For example, the return score and actual score rank list may look like:  
(\_score, my\_score)  
8.0079155 7.726183013850095  
7.4573565 7.240700372005419  
5.7250347 5.718247028450972  
5.582849 5.3363395329703245  
5.35503 5.120882132596499  
5.3481665 5.350607821185878  
5.231105 5.068864868915258  
5.190113 5.007839149255898  
4.9003787 4.8552734990413615

Do you have ideas does this happen? Is it because of the 9 bit precision or it's because of the skipping algorithm that does some approximation?

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [February 10, 2020, 8:54pm UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/11 "2020-02-10T20:54:13Z")

</div>

> In general the overall ranking is correct

Have you experienced a case where the ranking was incorrect? It looks from the result you submitted that the ranking is incorrect. We would appreciate if you share a reproducible test case that results in wrong ranking.

> Do you have ideas does this happen? Is it because of the 9 bit precision or it's because of the skipping algorithm that does some approximation?

Skipping algorithm should not be at blame here. The approximation happens because of 9 bit precision and converting Math.log result to a `float` score.

---

<div class="post-metadata">

**Author:** ![snakeztc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/snakeztc/32/60686_2.png) [@snakeztc](https://discuss.elastic.co/u/snakeztc)\
**Post date:** [February 11, 2020, 6:47pm UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/12 "2020-02-11T18:47:20Z")

</div>

> [@mayya](#):
>
> Skipping algorithm should not be at blame here. The approximation happens because of 9 bit precision and converting Math.log result to a `float` score.

Yes. There are cases where the ranking is incorrect as I shown from the example above. My index is pretty big and what's the best way for me to share a reproducible test?

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [February 11, 2020, 7:02pm UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/13 "2020-02-11T19:02:04Z")

</div>

@snakeztc  
The best way would be to create an issue in [https://github.com/elastic/elasticsearch](https://github.com/elastic/elasticsearch).  
`rank_feature` query scores don't depend on other documents, so showing just these few documents that produce the wrong ranking would be very helpful. Thank you

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 10, 2020, 7:02pm UTC](https://discuss.elastic.co/t/sparse-vector-vs-rank-features-which-one/214772/14 "2020-03-10T19:02:05Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
