# Inconsistent results while querying on a index

**URL:** <https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636>\
**Category:** Elasticsearch\
**Created:** [August 19, 2015, 5:29am UTC](https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636 "2015-08-19T05:29:10Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![RenuGuru1](https://avatars.discourse-cdn.com/v4/letter/r/c4cdca/32.png) [@RenuGuru1](https://discuss.elastic.co/u/RenuGuru1)\
**Post date:** [August 19, 2015, 5:29am UTC](https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636/1 "2015-08-19T05:29:10Z")

</div>

We are using elasticsearch 1.6.0 with 6 nodes.

We are getting inconsistent results while searching when the same query is requested concurrently, the order of results are changing.

Results are entirely different for search from 80 and size 2 and another search for the same query from 81 and size 2. While the 81st result should appear in both searches.

This doesn't occur when search is made only on primary shards. Replicas become inconsistent in a while.

We faced this same issue from the versions 0.90.x as well, This issue is on many versions. We are facing the same issue again on 1.6.0. On earlier versions, we solved it by making the index replica to 0 and then increasing to desired number. But since this is very risky to do it on indices of production environment, please suggest any other alternative solution to fix this inconsistency ASAP, without much risk on production cluster.

Also, Can you please fix these inconsistency issues in future release?

Thanks.

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [August 21, 2015, 3:48pm UTC](https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636/2 "2015-08-21T15:48:25Z")

</div>

What sort order are you using? When the sort value is the same on two documents, then the order in which documents are returned might different based on the queried shard. In order to make this issue less visible, you can make each user always use the same preference so that the same shard copy will be queried. [https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-preference.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-preference.html)

---

<div class="post-metadata">

**Author:** ![RenuGuru1](https://avatars.discourse-cdn.com/v4/letter/r/c4cdca/32.png) [@RenuGuru1](https://discuss.elastic.co/u/RenuGuru1)\
**Post date:** [August 24, 2015, 2:03pm UTC](https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636/3 "2015-08-24T14:03:29Z")

</div>

@jpountz, Thanks for the reply.

"What sort order are you using? " - I experienced this without using any kind of sorting.

"you can make each user always use the same preference so that the same shard copy will be queried." - I was thinking of using \_primary\_first preference for all calls, does it have any impact on performance when compared to calling without preference?

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [August 24, 2015, 2:56pm UTC](https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636/4 "2015-08-24T14:56:53Z")

</div>

> [@RenuGuru1](#):
>
> "you can make each user always use the same preference so that the same shard copy will be queried." - I was thinking of using primaryfirst preference for all calls, does it have any impact on performance when compared to calling without preference?

Using primary\_first for all calls makes little sense if you have replicas as it will prevent elasticsearch from using replicas to scale reads. It would be better to use eg. the user id as a preference for each search request.

---

<div class="post-metadata">

**Author:** ![RenuGuru1](https://avatars.discourse-cdn.com/v4/letter/r/c4cdca/32.png) [@RenuGuru1](https://discuss.elastic.co/u/RenuGuru1)\
**Post date:** [August 25, 2015, 7:20am UTC](https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636/5 "2015-08-25T07:20:57Z")

</div>

Yes! this makes sense, a question on how userId/sessionId as preference works. Say I'm using sessionId as preference, when the user searches, it hits some shards and gets results. on another search the same shards will be used for getting results for this user.

During this time, if a new document is indexed and stored in another shard different from the user's pre-allocated shards. This newly indexed documents will not be returned on search right? Yeah, this will not cause much issue as the sessionId differs and will be searched on different shards, each time a user logins'. But would like to know how this works, please enlighten me. Thanks.

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [August 25, 2015, 8:19am UTC](https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636/6 "2015-08-25T08:19:26Z")

</div>

I think you are confusing preference and routing.

Say you have an index with 2 primary shards, that have one replica shard each. This is a total of 4 shards. Let's call the primary shards 0P and 1P, and the replica shards 0R and 1R.

Routing by user id would allow you to make sure that all documents from the same user id will be on the same shard. For instance maybe all documents from user `foo` would be on shards 0P and 0R while all documents from user `bar` would end up on shards 1P and 1R. So at search time, if you know which user you are searching on, you can configure routing and only go to the shard that holds data for this user.

On the contrary, preference doesn't do anything at index time. But at search time it will make sure that you always hit the same shards if you specify the same preference. Without preference the first search request might go to shards 0P and 1R while the second search request will go to 0R and 1R. With preference configured, you are guaranteed to always hit the same shards.

---

<div class="post-metadata">

**Author:** ![RenuGuru1](https://avatars.discourse-cdn.com/v4/letter/r/c4cdca/32.png) [@RenuGuru1](https://discuss.elastic.co/u/RenuGuru1)\
**Post date:** [August 25, 2015, 10:06am UTC](https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636/7 "2015-08-25T10:06:46Z")

</div>

@Andrien, Yes I understand routing and preference are different.  
I think it would be better to understand if I give you an example of how we use elasticsearch, lets say we use it for shopping cart where user searches for all of our contents/products in our website search page. The sequence of the products displayed gets changed on concurrent requests. This is the kind of issue we are facing.

Considering this case, I'm not talking about routing since we are not storing user specific contents. We want some users to search for all of our products.

Now,

As you said, say I have an index with 2 primary shards, that have one replica shard each. This is a total of 4 shards. Let's call the primary shards 0P and 1P, and the replica shards 0R and 1R.

Now, with user's sessionId as preference, on his first search if shards 0R and 1P are searched, and  
parallely if on indexing, new products gets indexed on 0P and 1P. Before it gets replicated, if the user does a second search, would the products indexed on 0P be returned?

---

<div class="post-metadata">

**Author:** ![RenuGuru1](https://avatars.discourse-cdn.com/v4/letter/r/c4cdca/32.png) [@RenuGuru1](https://discuss.elastic.co/u/RenuGuru1)\
**Post date:** [August 25, 2015, 10:53am UTC](https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636/8 "2015-08-25T10:53:17Z")

</div>

And it is obvious that some non sync has happened while elasticsearch replicates, why and when does this non sync of replicas happen? It would be very helpful if you can avoid this non sync issue from elasticsearch itself, so that every user can see the products in same sequence.

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [August 25, 2015, 1:52pm UTC](https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636/9 "2015-08-25T13:52:15Z")

</div>

If you want every user to see products in the same sequence then you need to change your sort order to tie break on the \_uid field, so that the sort order is deterministic even in case a different replica is queried or segments are being merged.

---

<div class="post-metadata">

**Author:** ![RenuGuru1](https://avatars.discourse-cdn.com/v4/letter/r/c4cdca/32.png) [@RenuGuru1](https://discuss.elastic.co/u/RenuGuru1)\
**Post date:** [August 26, 2015, 1:33pm UTC](https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636/10 "2015-08-26T13:33:13Z")

</div>

Adding our conversation here for reference,

> [@On another topic](https://discuss.elastic.co/t/28089/1):
>
> Hi Grand,
> 
> Let me understand. Say when user searches with some keyword, we will show products ranked based on lucene's text relevance of the keyword, usage of the product (views/timespent) and metadata weight like if a product doesn't have description it will be demoted compared to documents with description. All these kind of weightage buckets are considered while ranking. So, how does this \_uid field sort order help? I don't get it.

> [@On another topic](https://discuss.elastic.co/t/28089/2):
>
> It will help by forcing elasticsearch to return documents in a deterministic order if your documents have the same score, usage and metadata so that merges or querying a different replica won't have an impact on the order in which documents are returned. It would look something like
> 
> {  
> "sort": [  
> \_score,  
> { "views": "asc" },  
> { "\_uid": "asc" } // tie-break  
> ]  
> }

> [@On another topic](https://discuss.elastic.co/t/28089/3):
>
> Nice! I can try this in future, when I face the same issue on documents with same score, but I see documents with completely different scores, so my guess is few documents are missing in replica while they are available in primary.
> 
> Lets say a document d1 is in 43th position in primary shards, while it is in 37th position in replica. We search for the doc d1 from 35 to 40 in first call, and 41 to 45 in second call. Now, if first call searches on replica and second call searches on primary shards, the doc d1 will appear on both result sets, this d1 is appearing twice as a duplicate while scrolling on the website.

> [@On another topic](https://discuss.elastic.co/t/28089/4):
>
> Right, this is why using preference is useful in addition to tie-breaking on the \_uid.
> 
> Differences in ranking may happen between the primary and the replica because they can refresh at different times and because they can have different deleted documents which are still taken into account for index statistics, but the differences should decrease as your index grows.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:53pm UTC](https://discuss.elastic.co/t/inconsistent-results-while-querying-on-a-index/27636/11 "2017-07-05T23:53:47Z")

</div>


