# Rank based on rarity of a field value

**URL:** <https://discuss.elastic.co/t/rank-based-on-rarity-of-a-field-value/323546>\
**Category:** Elasticsearch\
**Created:** [January 19, 2023, 8:09pm UTC](https://discuss.elastic.co/t/rank-based-on-rarity-of-a-field-value/323546 "2023-01-19T20:09:50Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![pheeria](https://avatars.discourse-cdn.com/v4/letter/p/7cd45c/32.png) [@pheeria](https://discuss.elastic.co/u/pheeria)\
**Post date:** [January 19, 2023, 8:09pm UTC](https://discuss.elastic.co/t/rank-based-on-rarity-of-a-field-value/323546/1 "2023-01-19T20:09:50Z")

</div>

Hi 🖖

I'd like to know how can I rank lower items, which have fields that are frequently appearing among the results.  
Say, we have a similar result set:

```auto
"name": "Red T-Shirt"
"store": "Zara"

"name": "Yellow T-Shirt"
"store": "Zara"

"name": "Red T-Shirt"
"store": "Bershka"

"name": "Green T-Shirt"
"store": "Benetton"

```

I'd like to rank the documents in such a manner that the documents containing frequently found fields,  
"store" in this case, are deboosted to appear lower in the results.  
This is to achieve a bit of variety, so that the search doesn't yield top results from the same store.

In the example above, if I search for "T-Shirt", I want to see one Zara T-Shirt at the top and the rest  
of Zara T-Shirts should be appearing lower, after all other unique stores.

So far I tried to research for using aggregation buckets for sorting or script sorting, but without success.  
Is it possible to achieve this inside of the search engine?

Many thanks in advance!

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood1/32/101255_2.png) [@Mark\_Harwood1](https://discuss.elastic.co/u/Mark_Harwood1)\
**Post date:** [January 20, 2023, 7:50am UTC](https://discuss.elastic.co/t/rank-based-on-rarity-of-a-field-value/323546/2 "2023-01-20T07:50:06Z")

</div>

> [@pheeria](#):
>
> This is to achieve a bit of variety, so that the search doesn't yield top results from the same store.

Showing top results sorted by natural score but with some diversity can be achieved using a ‘top\_hits’ aggregation under a [diversified sampler aggregation](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-diversified-sampler-aggregation.html)  
Pagination may be tricky using this approach though.

---

<div class="post-metadata">

**Author:** ![pheeria](https://avatars.discourse-cdn.com/v4/letter/p/7cd45c/32.png) [@pheeria](https://discuss.elastic.co/u/pheeria)\
**Post date:** [February 8, 2023, 12:19pm UTC](https://discuss.elastic.co/t/rank-based-on-rarity-of-a-field-value/323546/3 "2023-02-08T12:19:10Z")

</div>

Sorry for the late reaction and thank you very much for the help!  
Do I understand the idea correctly?

```auto
{
  "query": {}, // whatever query
  "size": 0, // since we don't use hits
  "aggs": {
    "my_unbiased_sample": {
      "diversified_sampler": {
        "shard_size": 100,
        "field": "store"
      },
      "aggs": {
        "keywords": {
          "top_hits": {
            "_source": {
              "includes": ["name", "store"]
            },
            "size": 100
          }
        }
      }
    }
  }
}

```

This works! I wanted also to ask performance implications of this approach. How much more costly is this in comparison to not doing it, or doing this kind of "diversification" on the backend?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood1/32/101255_2.png) [@Mark\_Harwood1](https://discuss.elastic.co/u/Mark_Harwood1)\
**Post date:** [February 8, 2023, 3:09pm UTC](https://discuss.elastic.co/t/rank-based-on-rarity-of-a-field-value/323546/4 "2023-02-08T15:09:34Z")

</div>

> [@pheeria](#):
>
> Do I understand the idea correctly?

Yes, maybe one thing to consider is the max items-per-store you want to see via the [max docs per value](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-diversified-sampler-aggregation.html#_max_docs_per_value) setting.

It shouldn't be too bad. For matching docs there's a cost in terms of an additional lookup to find the store and there's a small memory overhead to hold the set of best matching doc IDs for each unique store. A lot depends on your queries/data/sharding etc so benchmarking will give you the reliable answer.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 8, 2023, 3:10pm UTC](https://discuss.elastic.co/t/rank-based-on-rarity-of-a-field-value/323546/5 "2023-03-08T15:10:27Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
