# EC2 discovery leads to two masters

**URL:** <https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101>\
**Category:** Elasticsearch\
**Created:** [August 9, 2011, 3:19pm UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101 "2011-08-09T15:19:23Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pavel\_Penchev](https://avatars.discourse-cdn.com/v4/letter/p/ec9cab/32.png) [@Pavel\_Penchev](https://discuss.elastic.co/u/Pavel_Penchev)\
**Post date:** [August 9, 2011, 3:19pm UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/1 "2011-08-09T15:19:23Z")

</div>

Hi,

I'm having some troubles configuring ES in the cloud. Most of the time  
everything works, but sometimes the discovery fails and I endup with two  
masters using the same cluster name.  
The situation happens on roughly 1 out of 10 startups.

## I'm using 0.17.4 embedded, the configuration looks like this

cluster:  
name: default-cluster-name

index:  
number\_of\_shards: 2  
number\_of\_replicas: 1

discovery:  
type: ec2  
zen:  
minimum\_master\_nodes: 1

cloud:  
aws:  
access\_key: XXXXXXXXXX  
secret\_key: XXXXXXXXXXXXXXXXXXXXXXXXXXXXXX

Trace logs can be found here [https://gist.github.com/1134288](https://gist.github.com/1134288). Any ideas  
what am I missing?

Thanks in advance,  
Pavel

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 9, 2011, 6:22pm UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/2 "2011-08-09T18:22:06Z")

</div>

It seems like the two nodes ended up not seeing each other properly, thus  
each elected itself as the master. If you increase the ping\_timeout (it  
defaults to 3s) then it should go away. Set discovery.zen.ping.timeout to  
something like 10s or 20s.

On Tue, Aug 9, 2011 at 6:19 PM, Pavel Penchev [pavel.penchev@gmail.com](mailto:pavel.penchev@gmail.com)wrote:

> Hi,
> 
> I'm having some troubles configuring ES in the cloud. Most of the time  
> everything works, but sometimes the discovery fails and I endup with two  
> masters using the same cluster name.  
> The situation happens on roughly 1 out of 10 startups.
> 
> ## I'm using 0.17.4 embedded, the configuration looks like this
> 
> cluster:  
> name: default-cluster-name
> 
> index:  
> number\_of\_shards: 2  
> number\_of\_replicas: 1
> 
> discovery:  
> type: ec2  
> zen:  
> minimum\_master\_nodes: 1
> 
> cloud:  
> aws:  
> access\_key: XXXXXXXXXX  
> secret\_key: XXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
> 
> Trace logs can be found here [https://gist.github.com/1134288](https://gist.github.com/1134288). Any ideas  
> what am I missing?
> 
> Thanks in advance,  
> Pavel

---

<div class="post-metadata">

**Author:** ![jjasinek](https://avatars.discourse-cdn.com/v4/letter/j/b782af/32.png) [@jjasinek](https://discuss.elastic.co/u/jjasinek)\
**Post date:** [August 9, 2011, 8:06pm UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/3 "2011-08-09T20:06:46Z")

</div>

Shay,

If two nodes did participate in a network partition (even on a local  
network), and thus end up self-promoting each other to a master  
status, what happens when they see each other again?

Jason

On Aug 9, 1:22 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> It seems like the two nodes ended up not seeing each other properly, thus  
> each elected itself as the master. If you increase the ping\_timeout (it  
> defaults to 3s) then it should go away. Set discovery.zen.ping.timeout to  
> something like 10s or 20s.
> 
> On Tue, Aug 9, 2011 at 6:19 PM, Pavel Penchev [pavel.penc...@gmail.com](mailto:pavel.penc...@gmail.com)wrote:
> 
> > Hi,
> 
> > I'm having some troubles configuring ES in the cloud. Most of the time  
> > everything works, but sometimes the discovery fails and I endup with two  
> > masters using the same cluster name.  
> > The situation happens on roughly 1 out of 10 startups.
> 
> > ## I'm using 0.17.4 embedded, the configuration looks like this
> > 
> > cluster:  
> > name: default-cluster-name
> 
> > index:  
> > number\_of\_shards: 2  
> > number\_of\_replicas: 1
> 
> > discovery:  
> > type: ec2  
> > zen:  
> > minimum\_master\_nodes: 1
> 
> > cloud:  
> > aws:  
> > access\_key: XXXXXXXXXX  
> > secret\_key: XXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
> 
> > Trace logs can be found herehttps://gist.github.com/1134288. Any ideas  
> > what am I missing?
> 
> > Thanks in advance,  
> > Pavel

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 9, 2011, 8:34pm UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/4 "2011-08-09T20:34:40Z")

</div>

Nothing, they will remain partitioned, and you will need to decide which one  
to restart. The minimum\_master\_nodes is there to help reduce chances of it  
happening. (on a 2 node cluster though, this setting does not mean much).

On Tue, Aug 9, 2011 at 11:06 PM, jjasinek [jjasinek@gmail.com](mailto:jjasinek@gmail.com) wrote:

> Shay,
> 
> If two nodes did participate in a network partition (even on a local  
> network), and thus end up self-promoting each other to a master  
> status, what happens when they see each other again?
> 
> Jason
> 
> On Aug 9, 1:22 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > It seems like the two nodes ended up not seeing each other properly, thus  
> > each elected itself as the master. If you increase the ping\_timeout (it  
> > defaults to 3s) then it should go away. Set discovery.zen.ping.timeout to  
> > something like 10s or 20s.
> > 
> > On Tue, Aug 9, 2011 at 6:19 PM, Pavel Penchev \<[pavel.penc...@gmail.com](mailto:pavel.penc...@gmail.com)  
> > wrote:
> > 
> > > Hi,
> > 
> > > I'm having some troubles configuring ES in the cloud. Most of the time  
> > > everything works, but sometimes the discovery fails and I endup with  
> > > two  
> > > masters using the same cluster name.  
> > > The situation happens on roughly 1 out of 10 startups.
> > 
> > > ## I'm using 0.17.4 embedded, the configuration looks like this
> > > 
> > > cluster:  
> > > name: default-cluster-name
> > 
> > > index:  
> > > number\_of\_shards: 2  
> > > number\_of\_replicas: 1
> > 
> > > discovery:  
> > > type: ec2  
> > > zen:  
> > > minimum\_master\_nodes: 1
> > 
> > > cloud:  
> > > aws:  
> > > access\_key: XXXXXXXXXX  
> > > secret\_key: XXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
> > 
> > > Trace logs can be found herehttps://gist.github.com/1134288. Any ideas  
> > > what am I missing?
> > 
> > > Thanks in advance,  
> > > Pavel

---

<div class="post-metadata">

**Author:** ![Pavel\_Penchev](https://avatars.discourse-cdn.com/v4/letter/p/ec9cab/32.png) [@Pavel\_Penchev](https://discuss.elastic.co/u/Pavel_Penchev)\
**Post date:** [August 16, 2011, 9:28am UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/5 "2011-08-16T09:28:07Z")

</div>

Hi,

sorry to bump an old thread, for completeness I just want to confirm  
that setting discovery.zen.ping\_timeout to 15s works like a charm in my  
case.

Many thanks for the quick response,  
Pavel

On 9.08.2011 21:22, Shay Banon wrote:

> It seems like the two nodes ended up not seeing each other properly,  
> thus each elected itself as the master. If you increase the  
> ping\_timeout (it defaults to 3s) then it should go away.  
> Set discovery.zen.ping.timeout to something like 10s or 20s.
> 
> On Tue, Aug 9, 2011 at 6:19 PM, Pavel Penchev \<[pavel.penchev@gmail.com](mailto:pavel.penchev@gmail.com)  
> [mailto:pavel.penchev@gmail.com](mailto:pavel.penchev@gmail.com)\> wrote:
> 
> ```
> Hi,
> 
> I'm having some troubles configuring ES in the cloud. Most of the
> time everything works, but sometimes the discovery fails and I
> endup with two masters using the same cluster name.
> The situation happens on roughly 1 out of 10 startups.
> 
> I'm using 0.17.4 embedded, the configuration looks like this
> ------------------------------------------------------------
> cluster:
> name: default-cluster-name
> 
> index:
> number_of_shards: 2
> number_of_replicas: 1
> 
> discovery:
> type: ec2
> zen:
> minimum_master_nodes: 1
> 
> cloud:
> aws:
> access_key: XXXXXXXXXX
> secret_key: XXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
> 
> Trace logs can be found here https://gist.github.com/1134288. Any
> ideas what am I missing?
> 
> Thanks in advance,
> Pavel
> 
> ```

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 16, 2011, 7:17pm UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/6 "2011-08-16T19:17:15Z")

</div>

Hi Pavel, that timeout value will often increase based on the number of  
non-cluster nodes you have under EC2 management. At least that has been my  
experience.

A trick to keep it working well is to make sure that all the nodes that are  
in your ElasticSearch cluster are part of the same EC2 group. Then use the ES  
groups setting[http://www.elasticsearch.org/guide/reference/modules/discovery/ec2.html](http://www.elasticsearch.org/guide/reference/modules/discovery/ec2.html)to limit those nodes that ES looks for to establish membership.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 16, 2011, 11:42pm UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/7 "2011-08-16T23:42:56Z")

</div>

Two notes on that: You can use ec2 tags as well to filter down the list of  
instances needed to be pinged, and, in 0.17, the unicast discovery is  
considerably more lightweight compared to previous versions.

On Tue, Aug 16, 2011 at 10:17 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:

> Hi Pavel, that timeout value will often increase based on the number of  
> non-cluster nodes you have under EC2 management. At least that has been my  
> experience.
> 
> A trick to keep it working well is to make sure that all the nodes that are  
> in your Elasticsearch cluster are part of the same EC2 group. Then use the ES  
> groups setting[http://www.elasticsearch.org/guide/reference/modules/discovery/ec2.html](http://www.elasticsearch.org/guide/reference/modules/discovery/ec2.html)to limit those nodes that ES looks for to establish membership.

---

<div class="post-metadata">

**Author:** ![Pavel\_Penchev](https://avatars.discourse-cdn.com/v4/letter/p/ec9cab/32.png) [@Pavel\_Penchev](https://discuss.elastic.co/u/Pavel_Penchev)\
**Post date:** [August 18, 2011, 7:50am UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/8 "2011-08-18T07:50:46Z")

</div>

Thanks James, we'll make use of the setting. Indeed the production EC2  
environment is quite heterogeneous.

Pavel

On 16.08.2011 22:17, James Cook wrote:

> Hi Pavel, that timeout value will often increase based on the number  
> of non-cluster nodes you have under EC2 management. At least that has  
> been my experience.
> 
> A trick to keep it working well is to make sure that all the nodes  
> that are in your Elasticsearch cluster are part of the same EC2 group.  
> Then use the ES groups setting  
> [http://www.elasticsearch.org/guide/reference/modules/discovery/ec2.html](http://www.elasticsearch.org/guide/reference/modules/discovery/ec2.html)  
> to limit those nodes that ES looks for to establish membership.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [August 19, 2011, 8:56am UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/9 "2011-08-19T08:56:02Z")

</div>

Hi James (or anybody else with similar experience)

On Tue, 2011-08-16 at 12:17 -0700, James Cook wrote:

> Hi Pavel, that timeout value will often increase based on the number  
> of non-cluster nodes you have under EC2 management. At least that has  
> been my experience.

> A trick to keep it working well is to make sure that all the nodes  
> that are in your Elasticsearch cluster are part of the same EC2 group.  
> Then use the ES groups setting to limit those nodes that ES looks for  
> to establish membership.

Given that getting ES to work well under EC2 seems to present a bit of a  
challenge, how would you feel about writing a tutorial for  
[elasticsearch.org](http://elasticsearch.org)?

It would be an invaluable resource.

clint

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 19, 2011, 12:59pm UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/10 "2011-08-19T12:59:34Z")

</div>

I think that would be useful as well. I'll try to carve out some time to get  
something started.

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 19, 2011, 1:01pm UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/11 "2011-08-19T13:01:23Z")

</div>

And Clinton, a cookbook of search recipes would be awesome to see on a web  
page. 🙂

You have solved many gotchas for people over the past months.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [August 19, 2011, 1:04pm UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/12 "2011-08-19T13:04:40Z")

</div>

On Fri, 2011-08-19 at 06:01 -0700, James Cook wrote:

> And Clinton, a cookbook of search recipes would be awesome to see on a  
> web page. 🙂

touchÃ© 😉

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:56am UTC](https://discuss.elastic.co/t/ec2-discovery-leads-to-two-masters/5101/13 "2017-07-06T03:56:32Z")

</div>


