# Single-node restarts and re-indexing

**URL:** <https://discuss.elastic.co/t/single-node-restarts-and-re-indexing/9249>\
**Category:** Elasticsearch\
**Created:** [October 5, 2012, 2:45am UTC](https://discuss.elastic.co/t/single-node-restarts-and-re-indexing/9249 "2012-10-05T02:45:39Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Carl\_C](https://avatars.discourse-cdn.com/v4/letter/c/a9a28c/32.png) [@Carl\_C](https://discuss.elastic.co/u/Carl_C)\
**Post date:** [October 5, 2012, 2:45am UTC](https://discuss.elastic.co/t/single-node-restarts-and-re-indexing/9249/1 "2012-10-05T02:45:39Z")

</div>

Hi everyone,

I'm testing a production ElasticSearch system, and I have two  
administrative questions:

1. Is there a good way to handle single node restarts? I've dug through  
past posts and can't find anything very useful. We're using the local  
gateway (and local store), so it should be possible for a single-node  
restart to trigger a minimal amount of shard reallocation. I am sure  
there's a way to delay shard reallocation with the appropriate settings,  
but I haven't figured out the right mixture yet. Any pointers would be  
great!

2. Handling index growth: I watched Shay's talk on different data flows,  
and our current setup is most like the "users flow": all documents have a  
user and most queries are restricted to a single user, so we set the  
routing parameter to be the user id. As our index grows, I am worried that  
individual shards will get to be too large. Most of our queries are over  
_all_ of a user's data, so creating a new index per-week (say) doesn't seem  
like the best solution, as it eliminates the benefits of routing.

I understand the reasons for not allowing live shard splitting. Our current  
plan in case we need to split shards is to start up a second cluster with  
more shards, duplicate live data to that cluster, and then use a backup to  
fill in the old data. Is there a better approach? Or should we really  
consider time-range indices?

Thanks,  
Carl

--

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [October 8, 2012, 9:17am UTC](https://discuss.elastic.co/t/single-node-restarts-and-re-indexing/9249/2 "2012-10-08T09:17:53Z")

</div>

Hi Carl

> 1. Is there a good way to handle single node restarts? I've dug  
> through past posts and can't find anything very useful. We're using  
> the local gateway (and local store), so it should be possible for a  
> single-node restart to trigger a minimal amount of shard reallocation.  
> I am sure there's a way to delay shard reallocation with the  
> appropriate settings, but I haven't figured out the right mixture yet.  
> Any pointers would be great!

You can use the cluster update settings API

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

to set cluster.routing.allocation.disable\_allocation before restarting,  
which will ensure that shards are not reallocated. Just remember to  
reenable it after your node has started up again.

> 1. Handling index growth: I watched Shay's talk on different data  
> flows, and our current setup is most like the "users flow": all  
> documents have a user and most queries are restricted to a single  
> user, so we set the routing parameter to be the user id. As our index  
> grows, I am worried that individual shards will get to be too large.  
> Most of our queries are over _all_ of a user's data, so creating a new  
> index per-week (say) doesn't seem like the best solution, as it  
> eliminates the benefits of routing.

If your use case fits the index-per-user model, then don't worry about  
the time-based model.

The key to this flexibility is aliases.

For example, you have a user 'foo' who starts out in your general  
all-users index. You can set up two aliases, one for writing and one for  
reading (I'll explain why later):

curl -XPOST '[http://127.0.0.1:9200/\_aliases?pretty=1](http://127.0.0.1:9200/_aliases?pretty=1)' -d '  
{  
"actions" : [  
{  
"add" : {  
"index" : "all\_clients",  
"alias" : "foo\_write",  
"routing" : "foo"  
}  
},  
{  
"add" : {  
"index" : "all\_clients",  
"filter" : {  
"term" : {  
"client\_id" : "foo"  
}  
},  
"alias" : "foo\_read",  
"routing" : "foo"  
}  
}  
]  
}  
'

The above means that:

- all documents for this client will be stored on a single shard in  
your all\_clients index
- when querying, the filter {client\_id == 'foo'} will be automatically  
applied to the results, so the 'foo' client appears to live in its  
own index

So why two aliases?

You can only write to a single index (or an alias  
which points to a single index), but you can query multiple indices. So  
to avoid making changes (like the ones explained below) to your  
application in the future, start out using separate foo\_write and  
foo\_read aliases.

Now, the client has grown large enough to warrant their own index.

You can create a new index 'foo\_v1' and then adjust your aliases so that  
foo\_write points to the new index, and foo\_read points to BOTH:

curl -XPOST '[http://127.0.0.1:9200/\_aliases?pretty=1](http://127.0.0.1:9200/_aliases?pretty=1)' -d '  
{  
"actions" : [  
{  
"remove" : {  
"index" : "all\_clients",  
"alias" : "foo\_write"  
}  
},  
{  
"add" : {  
"index" : "foo\_v1",  
"alias" : "foo\_write"  
}  
},  
{  
"add" : {  
"index" : "foo\_v1",  
"alias" : "foo\_read"  
}  
}  
]  
}  
'

So now, all writes will go the the new foo\_v1 index, but queries will go  
both to foo\_v1 and the old "all\_clients/routing:foo/client\_id==foo"  
index as well.

The only thing that you need to be careful about is getting and updating  
existing docs.

The doc-GET API can't use the 'doc\_read' alias because it points to more  
than one index. You have a few choices:

1. try the new index first, and if that fails to find the doc, try the  
old index
2. use a query instead of a GET
3. move the old data into the new index

Similarly, when you update/reindex an existing doc you must either:

1. store it in the same index that it came from, or
2. delete it from the old index and save it to the new index

clint

--

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [October 8, 2012, 9:38am UTC](https://discuss.elastic.co/t/single-node-restarts-and-re-indexing/9249/3 "2012-10-08T09:38:03Z")

</div>

> The only thing that you need to be careful about is getting and updating  
> existing docs.
> 
> The doc-GET API can't use the 'doc\_read' alias because it points to more  
> than one index. You have a few choices:
> 
> 1. try the new index first, and if that fails to find the doc, try the  
> old index
> 2. use a query instead of a GET
> 3. move the old data into the new index

I've opened an issue with some ideas about how to resolve the above more  
transparently:

> <https://github.com/elastic/elasticsearch/issues/2309>
>
> As described in this post https://groups.google.com/d/msg/elasticsearch/lhoG0gE9…Kvo/ez3XXIlsAx0J there is one bit missing from the index-per-user model which gets in the way of making things transparent to the application.
> 
> That 'bit' is that, given a document type and ID, you have to know which index the document lives in in order to be able to GET it.
> 
> So when you have documents sitting in your general all-users index and you add a user-specific index, you can't easily know which index to talk to in order to retrieve a document.
> 
> What about some mechanism for doing a GET from a list of indices? In other words, try index 1, if the document doesn't exist, try index 2, etc
> 
> Perhaps it is a \`read\_priority\` setting in an alias? Something like:
> 
> \`\`\`
> curl -XPOST 'http://127.0.0.1:9200/\_aliases?pretty=1' -d '
> {
> "actions" : \[
> {
> "add" : {
> "read\_priority" : 1,
> "index" : "foo\_v1",
> "default" : 1,
> "alias" : "foo"
> }
> },
> {
> "add" : {
> "read\_priority" : 2,
> "index" : "all\_users",
> "filter" : {
> "term" : {
> "user\_id" : "foo"
> }
> },
> "routing" : "foo",
> "alias" : "foo"
> }
> }
> \]
> }
> '
> \`\`\`
> 
> I've also added a \`default\` param which indicates that, when creating new documents, the index \`foo\_v1\` should be used (as opposed to just failing with "can't write to multiple indices")

clint

--

---

<div class="post-metadata">

**Author:** ![Carl\_C](https://avatars.discourse-cdn.com/v4/letter/c/a9a28c/32.png) [@Carl\_C](https://discuss.elastic.co/u/Carl_C)\
**Post date:** [October 8, 2012, 5:46pm UTC](https://discuss.elastic.co/t/single-node-restarts-and-re-indexing/9249/4 "2012-10-08T17:46:08Z")

</div>

Hi Clint,

Thanks for your answer --- it's really helpful. Luckily (?), all of our  
code is written in terms of queries, not GETs, as we are using  
Elasticsearch as a form of secondary indexing for our (distributed)  
database, so any documents for which we already have the id can bypass ES  
altogether.

Carl

On Monday, October 8, 2012 2:38:13 AM UTC-7, Clinton Gormley wrote:

> > The only thing that you need to be careful about is getting and updating  
> > existing docs.
> > 
> > The doc-GET API can't use the 'doc\_read' alias because it points to more  
> > than one index. You have a few choices:
> > 
> > 1. try the new index first, and if that fails to find the doc, try the  
> > old index
> > 2. use a query instead of a GET
> > 3. move the old data into the new index
> 
> I've opened an issue with some ideas about how to resolve the above more  
> transparently:  
> [Multi-index document GET · Issue #2309 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/2309)
> 
> clint

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:09am UTC](https://discuss.elastic.co/t/single-node-restarts-and-re-indexing/9249/5 "2017-07-06T03:09:42Z")

</div>


