# Hash function for routing in ES 5.5

**URL:** <https://discuss.elastic.co/t/hash-function-for-routing-in-es-5-5/102865>\
**Category:** Elasticsearch\
**Created:** [October 5, 2017, 3:34pm UTC](https://discuss.elastic.co/t/hash-function-for-routing-in-es-5-5/102865 "2017-10-05T15:34:33Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![jmartinter](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jmartinter/32/22809_2.png) [@jmartinter](https://discuss.elastic.co/u/jmartinter)\
**Post date:** [October 5, 2017, 3:34pm UTC](https://discuss.elastic.co/t/hash-function-for-routing-in-es-5-5/102865/1 "2017-10-05T15:34:33Z")

</div>

Hi, all,

I'd like to order a long list of doc ids according to the shard each document is stored in. I paginate this list and run a query by id per page so, if most of the ids in page correspond to a same shard, system performance should improve.

To implement this I'd need to know which hash algorithm is ES 5.5 using for routing. I found some references to DJB in forums but I don't know if it is still valid.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [October 5, 2017, 3:47pm UTC](https://discuss.elastic.co/t/hash-function-for-routing-in-es-5-5/102865/2 "2017-10-05T15:47:22Z")

</div>

> [@jmartinter](#):
>
> system performance should improve.

Shards are designed to provide parallelism. Herding all requests to one shard and waiting for it to respond while all other shards stand idle is going to slow things down.

---

<div class="post-metadata">

**Author:** ![jmartinter](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jmartinter/32/22809_2.png) [@jmartinter](https://discuss.elastic.co/u/jmartinter)\
**Post date:** [October 6, 2017, 6:53am UTC](https://discuss.elastic.co/t/hash-function-for-routing-in-es-5-5/102865/3 "2017-10-06T06:53:57Z")

</div>

It seems I badly explained my use case.

My app gets a list of 5000 ids, paginates it and runs a query per ids page.  
At present, ids are in random order and every query hits most of the shards.  
If I could know the shard each id belongs to in advance, I'd group ids by shard so only a small part of the shards will be involved in each request.  
My understanding is the fewer shards you hit per request, the fewer resources (memory, threads) you require, so system performance improves.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 6, 2017, 7:05am UTC](https://discuss.elastic.co/t/hash-function-for-routing-in-es-5-5/102865/4 "2017-10-06T07:05:44Z")

</div>

If you index your documents with the id as a [routing key](https://www.elastic.co/guide/en/elasticsearch/reference/5.6/mapping-routing-field.html), you can fetch using the same routing key, which will cause only a single shard to be searched for each id.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [October 6, 2017, 8:14am UTC](https://discuss.elastic.co/t/hash-function-for-routing-in-es-5-5/102865/5 "2017-10-06T08:14:33Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> If you index your documents with the id as a routing key, you can fetch using the same routing key, which will cause only a single shard to be searched for each id.

Further, if your id **_is_** the elasticsearch document id then use GET not a search and for efficiciency's sake use [MGET](https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-multi-get.html)

---

<div class="post-metadata">

**Author:** ![jmartinter](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jmartinter/32/22809_2.png) [@jmartinter](https://discuss.elastic.co/u/jmartinter)\
**Post date:** [October 6, 2017, 10:37am UTC](https://discuss.elastic.co/t/hash-function-for-routing-in-es-5-5/102865/6 "2017-10-06T10:37:54Z")

</div>

Thanks for your replies.

Actually, I'm using MGET (a call to MGET for each ids page)

What I'm trying to do is preprocessing my 5000 ids list in order to sort it by shard (first all ids in shard1, later all ids in shard2, ...)  
Once done, I could paginate the list (100 ids/page) and and make a MGET request for each page, knowing each request will access just one or two of my shards (instead of most of them as happens now with my current unsorted list)  
All I need is to know the hash function currently used for routing in ES 5.5

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 6, 2017, 10:47am UTC](https://discuss.elastic.co/t/hash-function-for-routing-in-es-5-5/102865/7 "2017-10-06T10:47:56Z")

</div>

If you are using `mget`, it already goes to only one shard for each document, so you may not gain much by doing what you are suggesting.

---

<div class="post-metadata">

**Author:** ![jmartinter](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jmartinter/32/22809_2.png) [@jmartinter](https://discuss.elastic.co/u/jmartinter)\
**Post date:** [October 6, 2017, 12:45pm UTC](https://discuss.elastic.co/t/hash-function-for-routing-in-es-5-5/102865/8 "2017-10-06T12:45:07Z")

</div>

So, does it mean MGET just perform a bunch of GET operations?  
Would my approach make sense if I replace MGET by [ids query](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-ids-query.html)?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 6, 2017, 12:56pm UTC](https://discuss.elastic.co/t/hash-function-for-routing-in-es-5-5/102865/9 "2017-10-06T12:56:11Z")

</div>

MGET is a way to perform multiple GET requests in a single request, and is the most efficient way to return documents when you know the ids. I therefore do not see why you would need your approach.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 3, 2017, 1:07pm UTC](https://discuss.elastic.co/t/hash-function-for-routing-in-es-5-5/102865/10 "2017-11-03T13:07:30Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
