# Just Pushed: API: Allow to control document shard routing, and search shard routing

**URL:** <https://discuss.elastic.co/t/just-pushed-api-allow-to-control-document-shard-routing-and-search-shard-routing/3503>\
**Category:** Elasticsearch\
**Created:** [November 2, 2010, 6:00pm UTC](https://discuss.elastic.co/t/just-pushed-api-allow-to-control-document-shard-routing-and-search-shard-routing/3503 "2010-11-02T18:00:18Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 2, 2010, 6:00pm UTC](https://discuss.elastic.co/t/just-pushed-api-allow-to-control-document-shard-routing-and-search-shard-routing/3503/1 "2010-11-02T18:00:18Z")

</div>

Hi,

Just pushed support for explicit document routing and search routing.  
Here is the issue (content copied below):  
[http://github.com/elasticsearch/elasticsearch/issues/issue/470](http://github.com/elasticsearch/elasticsearch/issues/issue/470).

Currently, when indexing or deleting documents, they are hashed based on the  
type and id and routed to a shard based on the hash result. When searching,  
the search request is a broadcast request to all shards within an index.

Sometimes, it make sense to control this routing. For example, indexing blob  
posts for a specific user might be routed based on the user id. If routing  
is controlled, then when performing a search on posts for only that user,  
only the relevant shard that match the user id can be queried, resulting in  
much faster search and overhaul less load on the system.

The `index` and `delete` operation now allow for a `routing` parameter to be  
specified. When specified, that routing value will be used to control the  
shard placement.

The `get` operation allows to specify a `routing` parameter as well, to  
specify which shard the doc will be fetched from. Note, in our blog post  
partitioned by user example, doing a lookup just by the post id will not be  
enough and probably will not find anything, the user id will also need to be  
specified in the `routing` parameter.

The `search` and `count` operations accept a `routing` parameter as well,  
controlling which _shards_ the search will be executed on. The `routing`  
parameter accepts a comma separated list of the routing values to use, and  
the all relevant shards will execute the query.

Note, even when specifying a search for a specific user blog posts using the  
`routing` parameter set to the user id, filtering only the user posts is  
still needed by, for example, adding a `term` filter with the user id.

The `delete_by_query` operation also accepts a `routing` parameter, which is  
a comma separated list of routing values of controlling which shards the  
delete query will be executed on.

One question that might arise is why not use indices to get the same  
behavior. For example, create an index per user. The reason is that indices  
have a much lower limit on how many of them can be created. A single machine  
can easily support millions of users with millions of posts. Creating an  
index per user will mean millions of indices, which is problematic, as even  
with a single shard per index, it does mean millions of lucene indices.

-shay.banon

---

<div class="post-metadata">

**Author:** ![hartzler\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hartzler_2/32/3301_2.png) [@hartzler\_2](https://discuss.elastic.co/u/hartzler_2)\
**Post date:** [November 2, 2010, 6:15pm UTC](https://discuss.elastic.co/t/just-pushed-api-allow-to-control-document-shard-routing-and-search-shard-routing/3503/2 "2010-11-02T18:15:50Z")

</div>

Fantastic! I will definitely look into this.

I have been able to make good progress on a LRU index manager plugin for  
dealing with the "millions of lucene indexes" problem. It uses the  
open/close api you added earlier and works by extending the NodeClient  
implementation to manage the current open set of indexes. Paired with a  
reliable/available shared gateway config (I'm working on a cassandra gateway  
plugin as well, the s3 gateway would work great if on ec2) this may provide  
a workable approach to index per user strategy. Obviously this strategy  
will require you to have a small percentage of your users active over a  
reasonable time period (like 1-10% in any given hour).

However, being able to route to shards on a per user basis may well work  
better for this scenario as well, testing will tell.

On Tue, Nov 2, 2010 at 1:00 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Hi,
> 
> Just pushed support for explicit document routing and search routing.  
> Here is the issue (content copied below):  
> [API: Allow to control document shard routing, and search shard routing · Issue #470 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/470).
> 
> Currently, when indexing or deleting documents, they are hashed based on  
> the type and id and routed to a shard based on the hash result. When  
> searching, the search request is a broadcast request to all shards within an  
> index.
> 
> Sometimes, it make sense to control this routing. For example, indexing  
> blob posts for a specific user might be routed based on the user id. If  
> routing is controlled, then when performing a search on posts for only that  
> user, only the relevant shard that match the user id can be queried,  
> resulting in much faster search and overhaul less load on the system.
> 
> The `index` and `delete` operation now allow for a `routing` parameter to  
> be specified. When specified, that routing value will be used to control the  
> shard placement.
> 
> The `get` operation allows to specify a `routing` parameter as well, to  
> specify which shard the doc will be fetched from. Note, in our blog post  
> partitioned by user example, doing a lookup just by the post id will not be  
> enough and probably will not find anything, the user id will also need to be  
> specified in the `routing` parameter.
> 
> The `search` and `count` operations accept a `routing` parameter as well,  
> controlling which _shards_ the search will be executed on. The `routing`  
> parameter accepts a comma separated list of the routing values to use, and  
> the all relevant shards will execute the query.
> 
> Note, even when specifying a search for a specific user blog posts using  
> the `routing` parameter set to the user id, filtering only the user posts is  
> still needed by, for example, adding a `term` filter with the user id.
> 
> The `delete_by_query` operation also accepts a `routing` parameter, which  
> is a comma separated list of routing values of controlling which shards the  
> delete query will be executed on.
> 
> One question that might arise is why not use indices to get the same  
> behavior. For example, create an index per user. The reason is that indices  
> have a much lower limit on how many of them can be created. A single machine  
> can easily support millions of users with millions of posts. Creating an  
> index per user will mean millions of indices, which is problematic, as even  
> with a single shard per index, it does mean millions of lucene indices.
> 
> -shay.banon

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:17am UTC](https://discuss.elastic.co/t/just-pushed-api-allow-to-control-document-shard-routing-and-search-shard-routing/3503/3 "2017-07-06T04:17:17Z")

</div>


