# When restarting a test cluster, stop all nodes before starting any new nodes

**URL:** <https://discuss.elastic.co/t/when-restarting-a-test-cluster-stop-all-nodes-before-starting-any-new-nodes/11246>\
**Category:** Elasticsearch\
**Created:** [March 21, 2013, 4:31pm UTC](https://discuss.elastic.co/t/when-restarting-a-test-cluster-stop-all-nodes-before-starting-any-new-nodes/11246 "2013-03-21T16:31:16Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jamshid](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Jamshid](https://discuss.elastic.co/u/Jamshid)\
**Post date:** [March 21, 2013, 4:31pm UTC](https://discuss.elastic.co/t/when-restarting-a-test-cluster-stop-all-nodes-before-starting-any-new-nodes/11246/1 "2013-03-21T16:31:16Z")

</div>

Just thought I'd share the cause of a cluster startup problem I was having,  
in case anyone else makes the same mistake in their deployment script.

I wrote a simple shell script to deploy elasticsearch on three nodes, used  
in a nightly test environment.

I had a problem where (only!) sometimes the cluster would start red. Looked  
like there was a problem with a master failing. Node es1's log showed it  
detected\_master es3:

[2013-03-21 06:49:14,949][INFO][cluster.service]  
[[es1.tx.example.com](http://es1.tx.example.com)] detected\_master  
[[es3.tx.example.com](http://es3.tx.example.com)][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]],  
added  
{[[es3.tx.example.com](http://es3.tx.example.com)][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]],},  
reason: zen-disco-receive(from master  
[[[es3.tx.example.com](http://es3.tx.example.com)][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]]])

but then master\_left quickly because "[transport disconnected (with  
verified connect)]" and zen-disco-master\_failed:

[2013-03-21 06:49:14,990][INFO][discovery.zen]  
[[es1.tx.example.com](http://es1.tx.example.com)] master\_left  
[[[es3.tx.example.com](http://es3.tx.example.com)][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]]],  
reason [transport disconnected (with verified connect)]  
[2013-03-21 06:49:15,056][INFO][discovery]  
[[es1.tx.example.com](http://es1.tx.example.com)]  
[unittest-es1.tx.example.com-es2.tx.example.com-es3.tx.example.com/63m5eMNYR1-3xHgRPlYclA](http://unittest-es1.tx.example.com-es2.tx.example.com-es3.tx.example.com/63m5eMNYR1-3xHgRPlYclA)  
[2013-03-21 06:49:15,057][INFO][cluster.service]  
[[es1.tx.example.com](http://es1.tx.example.com)] master {new  
[[es1.tx.example.com](http://es1.tx.example.com)][63m5eMNYR1-3xHgRPlYclA][inet[/172.30.10.191:9300]],  
previous  
[[es3.tx.example.com](http://es3.tx.example.com)][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]]},  
removed  
{[[es3.tx.example.com](http://es3.tx.example.com)][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]],},  
reason: zen-disco-master\_failed  
([[es3.tx.example.com](http://es3.tx.example.com)][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]])

Eventually node es3 would reappear, but note it has a different unique id  
("MYc6..." instead of "4uxm...").

[2013-03-21 06:49:23,941][INFO][cluster.service]  
[[es1.tx.example.com](http://es1.tx.example.com)] added  
{[[es3.tx.example.com](http://es3.tx.example.com)][MYc6kHstS-6MbLgucUjIXg][inet[/172.30.10.193:9300]],},  
reason: zen-disco-receive(join from  
node[[[es3.tx.example.com](http://es3.tx.example.com)][MYc6kHstS-6MbLgucUjIXg][inet[/172.30.10.193:9300]]])

Of course, the problem was that I was looping through each node serially,  
killing any existing es process before starting a new one. So, sometimes a  
new node would come up before an old (master) node was killed.

for ELASTICSEARCH\_HOST in ${ELASTICSEARCH\_HOSTS[@]}  
do # ideally this would be done in parallel on each host  
ssh root@${ELASTICSEARCH\_HOST} "pkill -9 -f 'java.\*elasticsearch'"  
...  
ssh root@${ELASTICSEARCH\_HOST} "cd ${REMOTE\_DIR};  
./elasticsearch-${ELASTICSEARCH\_VERSION}/bin/elasticsearch  
-Des.max-open-files=true"  
...  
done

The fix is just to loop through all nodes and kill each es first, then  
start them up.

--Jamshid

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Jerome\_Gagnon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerome_gagnon/32/2178_2.png) [@Jerome\_Gagnon](https://discuss.elastic.co/u/Jerome_Gagnon)\
**Post date:** [March 25, 2013, 7:43pm UTC](https://discuss.elastic.co/t/when-restarting-a-test-cluster-stop-all-nodes-before-starting-any-new-nodes/11246/2 "2013-03-25T19:43:41Z")

</div>

Why not use the shutdown cluster api instead ? much much safer IMO.

On Thursday, March 21, 2013 12:31:16 PM UTC-4, Jamshid wrote:

> Just thought I'd share the cause of a cluster startup problem I was  
> having, in case anyone else makes the same mistake in their deployment  
> script.
> 
> I wrote a simple shell script to deploy elasticsearch on three nodes, used  
> in a nightly test environment.
> 
> I had a problem where (only!) sometimes the cluster would start red.  
> Looked like there was a problem with a master failing. Node es1's log  
> showed it detected\_master es3:
> 
> [2013-03-21 06:49:14,949][INFO][cluster.service] [  
> [es1.tx.example.com](http://es1.tx.example.com)] detected\_master [[es3.tx.example.com](http://es3.tx.example.com)][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]],  
> added {[[es3.tx.example.com](http://es3.tx.example.com)][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]],},  
> reason: zen-disco-receive(from master [[[es3.tx.example.com](http://es3.tx.example.com)  
> ][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]]])
> 
> but then master\_left quickly because "[transport disconnected (with  
> verified connect)]" and zen-disco-master\_failed:
> 
> [2013-03-21 06:49:14,990][INFO][discovery.zen] [  
> [es1.tx.example.com](http://es1.tx.example.com)] master\_left [[[es3.tx.example.com](http://es3.tx.example.com)][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]]],  
> reason [transport disconnected (with verified connect)]  
> [2013-03-21 06:49:15,056][INFO][discovery] [  
> [es1.tx.example.com](http://es1.tx.example.com)]  
> [unittest-es1.tx.example.com-es2.tx.example.com-es3.tx.example.com/63m5eMNYR1-3xHgRPlYclA](http://unittest-es1.tx.example.com-es2.tx.example.com-es3.tx.example.com/63m5eMNYR1-3xHgRPlYclA)  
> [2013-03-21 06:49:15,057][INFO][cluster.service] [  
> [es1.tx.example.com](http://es1.tx.example.com)] master {new [[es1.tx.example.com](http://es1.tx.example.com)][63m5eMNYR1-3xHgRPlYclA][inet[/172.30.10.191:9300]],  
> previous [[es3.tx.example.com](http://es3.tx.example.com)][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]]},  
> removed {[[es3.tx.example.com](http://es3.tx.example.com)][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]],},  
> reason: zen-disco-master\_failed ([[es3.tx.example.com](http://es3.tx.example.com)  
> ][4uxmLPLgS7qBlPo7Uv2AGA][inet[/172.30.10.193:9300]])
> 
> Eventually node es3 would reappear, but note it has a different unique id  
> ("MYc6..." instead of "4uxm...").
> 
> [2013-03-21 06:49:23,941][INFO][cluster.service] [  
> [es1.tx.example.com](http://es1.tx.example.com)] added {[[es3.tx.example.com](http://es3.tx.example.com)][MYc6kHstS-6MbLgucUjIXg][inet[/172.30.10.193:9300]],},  
> reason: zen-disco-receive(join from node[[[es3.tx.example.com](http://es3.tx.example.com)  
> ][MYc6kHstS-6MbLgucUjIXg][inet[/172.30.10.193:9300]]])
> 
> Of course, the problem was that I was looping through each node serially,  
> killing any existing es process before starting a new one. So, sometimes a  
> new node would come up before an old (master) node was killed.
> 
> for ELASTICSEARCH\_HOST in ${ELASTICSEARCH\_HOSTS[@]}  
> do # ideally this would be done in parallel on each host  
> ssh root@${ELASTICSEARCH\_HOST} "pkill -9 -f 'java.\*elasticsearch'"  
> ...  
> ssh root@${ELASTICSEARCH\_HOST} "cd ${REMOTE\_DIR};  
> ./elasticsearch-${ELASTICSEARCH\_VERSION}/bin/elasticsearch  
> -Des.max-open-files=true"  
> ...  
> done
> 
> The fix is just to loop through all nodes and kill each es first, then  
> start them up.
> 
> --Jamshid

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:44am UTC](https://discuss.elastic.co/t/when-restarting-a-test-cluster-stop-all-nodes-before-starting-any-new-nodes/11246/3 "2017-07-06T02:44:37Z")

</div>


