# Stopping and Staring a big cluster : best practice?

**URL:** <https://discuss.elastic.co/t/stopping-and-staring-a-big-cluster-best-practice/15540>\
**Category:** Elasticsearch\
**Created:** [January 31, 2014, 10:25pm UTC](https://discuss.elastic.co/t/stopping-and-staring-a-big-cluster-best-practice/15540 "2014-01-31T22:25:05Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mark\_Conlin](https://avatars.discourse-cdn.com/v4/letter/m/838e76/32.png) [@Mark\_Conlin](https://discuss.elastic.co/u/Mark_Conlin)\
**Post date:** [January 31, 2014, 10:25pm UTC](https://discuss.elastic.co/t/stopping-and-staring-a-big-cluster-best-practice/15540/1 "2014-01-31T22:25:05Z")

</div>

What is the best practice for stopping and starting a running cluster?

_My setup:_  
Elasticsearch 90.6

2 Master only nodes - each on their own box  
52 Data only nodes - spread across 6 boxes with 12,12,12,2,7,7 nodes on  
each

Each node is running under supervision (supervisord) so that they will be  
restarted if they crash on their own.

_Stop/Start routine:_

- turn off rivers (no more ingest)

- turn off shard allocation  
"cluster.routing.allocation.disable\_allocation":true  
"cluster.routing.allocation.disable\_replica\_allocation":true

- stop and restart nodes on each box using supervisorctl  
$\>supervisorctl stop elasticsearch-1  
$\>supervisorctl start elasticsearch-1

- wait for "initializing\_shards" count to reach 0

- turn on shard allocation

- wait for "unassigned\_shards" count to reach 0

- turn on rivers

_Result:_  
We almost always end up with one or a combination of several isssue :

- nodes pegged on heap and un responsive (cluster cant communicate with  
them, they are not hittable via api)
- nodes stuck initializing shards forever
- nodes stuck allocating shards forever
- "ghost" nodes; a second copy of a node in the cluster state (NOT process  
actually running) with that same name, different id. This actually doesnt  
affect es performance much but it makes es-head and other tools break due  
duplicate node/key name.

Some times, repeated opening and closing and index will get its shards to  
allocate and initialize. Sometimes not.

Thanks,  
Mark

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/e8561442-69a3-4ca1-bfbc-06c45bec39e6%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/e8561442-69a3-4ca1-bfbc-06c45bec39e6%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [February 1, 2014, 12:57am UTC](https://discuss.elastic.co/t/stopping-and-staring-a-big-cluster-best-practice/15540/2 "2014-02-01T00:57:55Z")

</div>

Shutdown:

curl -XPOST node:9200/\_shutdown

In the latest versions (1.0.0.RC1) ES shutdown chooses a strategy in which  
order nodes are closed, it makes things less error prone to shut down  
current master node at last.

Startup:

with shell execute a for loop over ssh command and start your favorite  
wrapper script on remote nodes in parallel. Master-eligible nodes first,  
non-master-eligible nodes last.

Jörg

On Fri, Jan 31, 2014 at 11:25 PM, Mark Conlin [mark.conlin@gmail.com](mailto:mark.conlin@gmail.com) wrote:

> What is the best practice for stopping and starting a running cluster?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHZnUtUKKoAHmFK49rayKomv-M1f\_8brqLYX3yxKH0i8g%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHZnUtUKKoAHmFK49rayKomv-M1f_8brqLYX3yxKH0i8g%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [February 1, 2014, 3:04am UTC](https://discuss.elastic.co/t/stopping-and-staring-a-big-cluster-best-practice/15540/3 "2014-02-01T03:04:02Z")

</div>

I would add to flush the transaction log after you have indexed all your  
content.

--  
Ivan

On Fri, Jan 31, 2014 at 4:57 PM, [joergprante@gmail.com](mailto:joergprante@gmail.com) \<  
[joergprante@gmail.com](mailto:joergprante@gmail.com)\> wrote:

> Shutdown:
> 
> curl -XPOST node:9200/\_shutdown
> 
> In the latest versions (1.0.0.RC1) ES shutdown chooses a strategy in which  
> order nodes are closed, it makes things less error prone to shut down  
> current master node at last.
> 
> Startup:
> 
> with shell execute a for loop over ssh command and start your favorite  
> wrapper script on remote nodes in parallel. Master-eligible nodes first,  
> non-master-eligible nodes last.
> 
> Jörg
> 
> On Fri, Jan 31, 2014 at 11:25 PM, Mark Conlin [mark.conlin@gmail.com](mailto:mark.conlin@gmail.com)wrote:
> 
> > What is the best practice for stopping and starting a running cluster?
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHZnUtUKKoAHmFK49rayKomv-M1f\_8brqLYX3yxKH0i8g%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHZnUtUKKoAHmFK49rayKomv-M1f_8brqLYX3yxKH0i8g%40mail.gmail.com)  
> .
> 
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQCW5iP81Ki2vOJXUrRPaV\_89VnDkxUizAXMKVNi1Ok24A%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQCW5iP81Ki2vOJXUrRPaV_89VnDkxUizAXMKVNi1Ok24A%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Mark\_Conlin](https://avatars.discourse-cdn.com/v4/letter/m/838e76/32.png) [@Mark\_Conlin](https://discuss.elastic.co/u/Mark_Conlin)\
**Post date:** [February 1, 2014, 7:21pm UTC](https://discuss.elastic.co/t/stopping-and-staring-a-big-cluster-best-practice/15540/4 "2014-02-01T19:21:25Z")

</div>

I believe the heart of this issue is JVM memory usage.

So does it make sense to delete warmers before shutdown (so they dont try  
to warm during initial node recovery)?

Does it make sense to lower my (currently set at 8):  
cluster.routing.allocation.node\_initial\_primaries\_recoveries  
cluster.routing.allocation.node\_concurrent\_recoveries

to limit the amount of work any one node will do at once?

Mark

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/db2552b0-85c7-4e74-82cd-b194e43d0bf4%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/db2552b0-85c7-4e74-82cd-b194e43d0bf4%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:53am UTC](https://discuss.elastic.co/t/stopping-and-staring-a-big-cluster-best-practice/15540/5 "2017-07-06T01:53:27Z")

</div>


