# Scoring problem between 2 machine

**URL:** <https://discuss.elastic.co/t/scoring-problem-between-2-machine/8445>\
**Category:** Elasticsearch\
**Created:** [July 18, 2012, 9:26am UTC](https://discuss.elastic.co/t/scoring-problem-between-2-machine/8445 "2012-07-18T09:26:40Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sicker](https://avatars.discourse-cdn.com/v4/letter/s/a6a055/32.png) [@Sicker](https://discuss.elastic.co/u/Sicker)\
**Post date:** [July 18, 2012, 9:26am UTC](https://discuss.elastic.co/t/scoring-problem-between-2-machine/8445/1 "2012-07-18T09:26:40Z")

</div>

I create a search query that order by score but the result is not same  
order.

_This is some item from first result._

{  
\_shard: 2  
\_node: U4zV\_tkeRpyRi59NNmNWQQ  
\_index: directory  
\_type: profile  
\_id: RM631130  
\_score: 0.55288595  
},

{  
\_shard: 4  
\_node: 3KNS8qPTQEWxuCRzwgMYLw  
\_index: directory  
\_type: profile  
\_id: RM631126  
\_score: 1.7044709  
}

\*This is some item from second result. \*

{  
\_shard: 4  
\_node: U4zV\_tkeRpyRi59NNmNWQQ  
\_index: directory  
\_type: profile  
\_id: RM631126  
\_score: 0.55287325  
},  
{  
\_shard: 2  
\_node: 3KNS8qPTQEWxuCRzwgMYLw  
\_index: directory  
\_type: profile  
\_id: RM631130  
\_score: 1.6957934  
}

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [July 19, 2012, 5:32pm UTC](https://discuss.elastic.co/t/scoring-problem-between-2-machine/8445/2 "2012-07-19T17:32:36Z")

</div>

Scoring might be different due to the distributed nature of  
Elasticsearch. Try adjusting the search type:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

There is a tradeoff between performance and accuracy of scoring.

--  
Ivan

On Wed, Jul 18, 2012 at 2:26 AM, Sicker [sicker27@gmail.com](mailto:sicker27@gmail.com) wrote:

> I create a search query that order by score but the result is not same  
> order.
> 
> This is some item from first result.
> 
> {  
> \_shard: 2  
> \_node: U4zV\_tkeRpyRi59NNmNWQQ  
> \_index: directory  
> \_type: profile  
> \_id: RM631130  
> \_score: 0.55288595  
> },
> 
> {  
> \_shard: 4  
> \_node: 3KNS8qPTQEWxuCRzwgMYLw  
> \_index: directory  
> \_type: profile  
> \_id: RM631126  
> \_score: 1.7044709  
> }
> 
> This is some item from second result.
> 
> {  
> \_shard: 4  
> \_node: U4zV\_tkeRpyRi59NNmNWQQ  
> \_index: directory  
> \_type: profile  
> \_id: RM631126  
> \_score: 0.55287325  
> },  
> {  
> \_shard: 2  
> \_node: 3KNS8qPTQEWxuCRzwgMYLw  
> \_index: directory  
> \_type: profile  
> \_id: RM631130  
> \_score: 1.6957934  
> }

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [July 20, 2012, 8:20am UTC](https://discuss.elastic.co/t/scoring-problem-between-2-machine/8445/3 "2012-07-20T08:20:07Z")

</div>

On Thu, 2012-07-19 at 10:32 -0700, Ivan Brusic wrote:

> Scoring might be different due to the distributed nature of  
> Elasticsearch. Try adjusting the search type:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/api/search/search-type.html)
> 
> There is a tradeoff between performance and accuracy of scoring.

Also, as the quantity of data you have grows, these differences tend to  
even out.

>

---

<div class="post-metadata">

**Author:** ![Radim](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/radim/32/2213_2.png) [@Radim](https://discuss.elastic.co/u/Radim)\
**Post date:** [July 20, 2012, 2:31pm UTC](https://discuss.elastic.co/t/scoring-problem-between-2-machine/8445/4 "2012-07-20T14:31:36Z")

</div>

To be a little less hand-wavy (please correct me if I'm wrong): some  
stats used in the scoring, like IDF, are computed _per shard_, by  
default. These stats are effectively computed only from the document  
set present in that one shard. This means that the same document can  
be scored differently, depending on which shard it ends up in.

By changing the search-type, you can change this behaviour so that the  
stats are computed on _index-level_ (not shard-level), i.e. from the  
document set present in the entire index. This helps to score  
consistently within one index.

AFAIK there is no way to run cross-index queries accurately. You can  
rely on the "evening out" that Clinton mentions. In that case you need  
to be careful your routing doesn't skew the stats distribution too  
much -- if each shard receives very different data, then the stats  
will never even out. The default routing is fine, as it sends out  
documents to random shards evenly (using hash of the id field).

HTH,  
Radim

On Jul 20, 10:20 am, Clinton Gormley [cl...@traveljury.com](mailto:cl...@traveljury.com) wrote:

> On Thu, 2012-07-19 at 10:32 -0700, Ivan Brusic wrote:
> 
> > Scoring might be different due to the distributed nature of  
> > Elasticsearch. Try adjusting the search type:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/api/search/search-type.html)
> 
> > There is a tradeoff between performance and accuracy of scoring.
> 
> Also, as the quantity of data you have grows, these differences tend to  
> even out.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [July 20, 2012, 5:32pm UTC](https://discuss.elastic.co/t/scoring-problem-between-2-machine/8445/5 "2012-07-20T17:32:39Z")

</div>

Hiya Radim

> By changing the search-type, you can change this behaviour so that the  
> stats are computed on _index-level_ (not shard-level), i.e. from the  
> document set present in the entire index. This helps to score  
> consistently within one index.

Not just at the index-level, but for all the shards involved in your  
query. So if you're doing a multi-index search and you use

search\_type=dfs\_query\_then\_fetch

then it will fetch the term frequencies from all shards (from all  
indices in your query) before executing it.

> AFAIK there is no way to run cross-index queries accurately. You can  
> rely on the "evening out" that Clinton mentions. In that case you need  
> to be careful your routing doesn't skew the stats distribution too  
> much -- if each shard receives very different data, then the stats  
> will never even out. The default routing is fine, as it sends out  
> documents to random shards evenly (using hash of the id field).

Sure, but for typical use cases, you'll be routing on (eg) a client, and  
searching within just that client, so terms will be evenly distributed  
for that client.

clint

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:19am UTC](https://discuss.elastic.co/t/scoring-problem-between-2-machine/8445/6 "2017-07-06T03:19:38Z")

</div>


