# Upgrade a 48-node cluster with minimal downtime?

**URL:** <https://discuss.elastic.co/t/upgrade-a-48-node-cluster-with-minimal-downtime/13101>\
**Category:** Elasticsearch\
**Created:** [August 7, 2013, 7:56am UTC](https://discuss.elastic.co/t/upgrade-a-48-node-cluster-with-minimal-downtime/13101 "2013-08-07T07:56:00Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Daniel\_Maher\_3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/daniel_maher_3/32/2014_2.png) [@Daniel\_Maher\_3](https://discuss.elastic.co/u/Daniel_Maher_3)\
**Post date:** [August 7, 2013, 7:56am UTC](https://discuss.elastic.co/t/upgrade-a-48-node-cluster-with-minimal-downtime/13101/1 "2013-08-07T07:56:00Z")

</div>

Hello,

I've got a 48-node ES cluster running 0.20.5 that I'd like to upgrade to  
0.90.3 with as little downtime as possible. I realise that this is a  
tall order, as the release notes (and past experience) make it clear  
that mixing versions is a Bad Idea, thus I can't simply roll through the  
nodes one by one.

Any hints ? 🙂

--  
dan (phrawzty).  
mozilla webops; european outpost.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Andy\_Wick](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andy_wick/32/44017_2.png) [@Andy\_Wick](https://discuss.elastic.co/u/Andy_Wick)\
**Post date:** [August 7, 2013, 6:29pm UTC](https://discuss.elastic.co/t/upgrade-a-48-node-cluster-with-minimal-downtime/13101/2 "2013-08-07T18:29:32Z")

</div>

No hints other then saying 0.90.3 was the first install/full restart I've  
done on our 30 node cluster that didn't end up with split brain! Woot!  
Usually one or two nodes don't join the cluster. So was able to shutdown,  
restart, and go yellow in under two minutes. Probably could have been  
faster if I remove some of the pauses I have that I added to "help" with  
previous full restart issues.

Andy

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [August 7, 2013, 9:47pm UTC](https://discuss.elastic.co/t/upgrade-a-48-node-cluster-with-minimal-downtime/13101/3 "2013-08-07T21:47:09Z")

</div>

If you are able to fit all your data on a subset of nodes and you're able  
to keep two clusters in sync and have smarts to know which cluster to route  
searches to (Eeesh, lots of ifs), you can run two clusters and switch nodes  
from old to new one by one. This also gives you the capability to revert to  
the old version if you run into issues with the new.

Lots of smarts needed on your side, likely quickest/easiest is to shut  
everything down, upgrade every node and start it back up.

Best Regards,  
Paul

On Wednesday, August 7, 2013 12:29:32 PM UTC-6, Andy Wick wrote:

> No hints other then saying 0.90.3 was the first install/full restart I've  
> done on our 30 node cluster that didn't end up with split brain! Woot!  
> Usually one or two nodes don't join the cluster. So was able to shutdown,  
> restart, and go yellow in under two minutes. Probably could have been  
> faster if I remove some of the pauses I have that I added to "help" with  
> previous full restart issues.
> 
> Andy

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Jerome\_Gagnon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerome_gagnon/32/2178_2.png) [@Jerome\_Gagnon](https://discuss.elastic.co/u/Jerome_Gagnon)\
**Post date:** [August 8, 2013, 1:49pm UTC](https://discuss.elastic.co/t/upgrade-a-48-node-cluster-with-minimal-downtime/13101/4 "2013-08-08T13:49:52Z")

</div>

We made the upgrade for a 100+ nodes cluster with a ~3 minutes downtime,  
wasn't that bad, you just have to be prepared and have the good tools.

On Wednesday, August 7, 2013 3:56:00 AM UTC-4, Daniel Maher wrote:

> Hello,
> 
> I've got a 48-node ES cluster running 0.20.5 that I'd like to upgrade to  
> 0.90.3 with as little downtime as possible. I realise that this is a  
> tall order, as the release notes (and past experience) make it clear  
> that mixing versions is a Bad Idea, thus I can't simply roll through the  
> nodes one by one.
> 
> Any hints ? 🙂
> 
> --  
> dan (phrawzty).  
> mozilla webops; european outpost.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [August 8, 2013, 2:00pm UTC](https://discuss.elastic.co/t/upgrade-a-48-node-cluster-with-minimal-downtime/13101/5 "2013-08-08T14:00:04Z")

</div>

On Thu, Aug 8, 2013 at 9:49 AM, Jérôme Gagnon [jerome.gagnon.1@gmail.com](mailto:jerome.gagnon.1@gmail.com)wrote:

> We made the upgrade for a 100+ nodes cluster with a ~3 minutes downtime,  
> wasn't that bad, you just have to be prepared and have the good tools.

It'd be really useful if you could explain some of your good tools!

Nik

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Jerome\_Gagnon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerome_gagnon/32/2178_2.png) [@Jerome\_Gagnon](https://discuss.elastic.co/u/Jerome_Gagnon)\
**Post date:** [August 8, 2013, 2:17pm UTC](https://discuss.elastic.co/t/upgrade-a-48-node-cluster-with-minimal-downtime/13101/6 "2013-08-08T14:17:59Z")

</div>

Tools can be as simple as parallel-ssh and (some) bash scripts.. that is  
error-prone and kind of sketchy, but this is one of the simplest possible  
solution..

You should probably more safely use chef, puppet or any other automation  
framework for more robustness and flexibility.

Jerome

On Thursday, August 8, 2013 10:00:04 AM UTC-4, Nikolas Everett wrote:

> On Thu, Aug 8, 2013 at 9:49 AM, Jérôme Gagnon \<[jerome....@gmail.com](mailto:jerome....@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > We made the upgrade for a 100+ nodes cluster with a ~3 minutes downtime,  
> > wasn't that bad, you just have to be prepared and have the good tools.
> 
> It'd be really useful if you could explain some of your good tools!
> 
> Nik

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![colinsurprenant](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colinsurprenant/32/14776_2.png) [@colinsurprenant](https://discuss.elastic.co/u/colinsurprenant)\
**Post date:** [February 14, 2014, 8:05pm UTC](https://discuss.elastic.co/t/upgrade-a-48-node-cluster-with-minimal-downtime/13101/7 "2014-02-14T20:05:39Z")

</div>

Hey! If I may chime in, you probably want to look into Ansible which offers  
very efficient and simple _automation_ facilities which other  
_provisioning_ tools like Chef & Puppet don't really have. I am not  
affiliated with Ansible, I just recently had a "ah-ah!" moment with it for  
exactly this kind of context.

Have fun,  
Colin

On Thursday, August 8, 2013 10:17:59 AM UTC-4, Jérôme Gagnon wrote:

> Tools can be as simple as parallel-ssh and (some) bash scripts.. that is  
> error-prone and kind of sketchy, but this is one of the simplest possible  
> solution..
> 
> You should probably more safely use chef, puppet or any other automation  
> framework for more robustness and flexibility.
> 
> Jerome
> 
> On Thursday, August 8, 2013 10:00:04 AM UTC-4, Nikolas Everett wrote:
> 
> > On Thu, Aug 8, 2013 at 9:49 AM, Jérôme Gagnon [jerome....@gmail.com](mailto:jerome....@gmail.com)wrote:
> > 
> > > We made the upgrade for a 100+ nodes cluster with a ~3 minutes downtime,  
> > > wasn't that bad, you just have to be prepared and have the good tools.
> > 
> > It'd be really useful if you could explain some of your good tools!
> > 
> > Nik

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/33715422-fe0c-42d3-b692-d2ed13acbb8c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/33715422-fe0c-42d3-b692-d2ed13acbb8c%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:49am UTC](https://discuss.elastic.co/t/upgrade-a-48-node-cluster-with-minimal-downtime/13101/8 "2017-07-06T01:49:57Z")

</div>


