# 2 Nodes ES cluster becomes unavailable for 2 -3 mins if one node (master) goes down

**URL:** <https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974>\
**Category:** Elasticsearch\
**Created:** [August 6, 2015, 1:58pm UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974 "2015-08-06T13:58:58Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Gaurav\_gupta](https://avatars.discourse-cdn.com/v4/letter/g/e274bd/32.png) [@Gaurav\_gupta](https://discuss.elastic.co/u/Gaurav_gupta)\
**Post date:** [August 6, 2015, 1:58pm UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974/1 "2015-08-06T13:58:58Z")

</div>

One of the customer has only 2 nodes cluster ( and don't want to add 3rd node) which becomes inaccessible for 2-3 mins, if first node (master) goes down. Below is the error which user face :-

**SearchPhaseExecutionException: Failed to execute phase [query\_fetch], all shards failed**

And after few mins (2-3 mins) second nodes takes the charge and it start responding to incoming requests.

Since, we can't force user to add 3rd node in cluster and they need second node just for fault tolerance purpose, so can we suggest user to wait till he gets "SearchPhaseExecutionException" or any such exception. Once the another node sends master alive signal/response then he can start sending requests.

Thoughts ?

Thanks  
Gaurav

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 7, 2015, 8:27am UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974/2 "2015-08-07T08:27:57Z")

</div>

You need to find why are the shards failing.

Are you monitoring ES? Check your logs too.

---

<div class="post-metadata">

**Author:** ![Gaurav\_gupta](https://avatars.discourse-cdn.com/v4/letter/g/e274bd/32.png) [@Gaurav\_gupta](https://discuss.elastic.co/u/Gaurav_gupta)\
**Post date:** [August 7, 2015, 7:16pm UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974/3 "2015-08-07T19:16:39Z")

</div>

Below is the exception message which says that first node ( master node ) has left or not connected as user has shut down it, for testing purpose. And now, until it elects and promote Node2 as master node, it throws below exception for around 2-3 mins :-

_SearchPhaseExecutionException: Failed to execute phase [query\_fetch], all shards failed; shardFailures {[wvn04kFMTYCqNMW\_9cKd1A][qlpanoramasearchindex2][0]: SendRequestTransportException[[node1][inet[/10.2.10.185:9300]][indices:data/read/search[phase/query+fetch]]]; nested: **NodeNotConnectedException** [[node1][inet[/10.2.10.185:9300]] Node not connected]; }_

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 7, 2015, 10:58pm UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974/4 "2015-08-07T22:58:21Z")

</div>

So why did the node leave? That's what you need to answer.  
Check your network, firewalls etc. What does your config look like?

---

<div class="post-metadata">

**Author:** ![Gaurav\_gupta](https://avatars.discourse-cdn.com/v4/letter/g/e274bd/32.png) [@Gaurav\_gupta](https://discuss.elastic.co/u/Gaurav_gupta)\
**Post date:** [August 8, 2015, 12:37pm UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974/5 "2015-08-08T12:37:48Z")

</div>

Actually, user is doing User acceptance testing in which one of the scenario is to shut down or remove one of the node, manually (i.e. remove first node which is currently master). Since, he is manually, removing the master node so incoming requests fail with exception for 2-3 mins. After 2-3 mins things work fine. Also, please note that it's occurring during load testing when user manually shut down the master node. Isn't 2nd node should become master immediately. Is it an accepted behaviour as election of new master node takes 2-3 mins?

Should we try to tweak the below settings, something like below so that new master has been elected with lesser delay :-

_discovery.zen.fd.ping\_timeout: 10s_  
_discovery.zen.fd.ping\_retries: 2_

Thanks  
Gaurav

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 9, 2015, 3:03am UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974/6 "2015-08-09T03:03:30Z")

</div>

That would help.

Can I ask why you are doing this sort of testing?

---

<div class="post-metadata">

**Author:** ![Gaurav\_gupta](https://avatars.discourse-cdn.com/v4/letter/g/e274bd/32.png) [@Gaurav\_gupta](https://discuss.elastic.co/u/Gaurav_gupta)\
**Post date:** [August 10, 2015, 8:36am UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974/7 "2015-08-10T08:36:58Z")

</div>

We are doing this type of testing since any node might go down (may be a network issue or heating issue or any other reason) in production environment also. So, this test scenario is just to make sure that ES works reliably with minimum or no down time in such scenarios i.e. high availability, fault tolerant behaviour of ES cluster in even worst case scenario.

Thanks  
Gaurav

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 10, 2015, 8:39am UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974/8 "2015-08-10T08:39:24Z")

</div>

Are you running dedicated masters?

---

<div class="post-metadata">

**Author:** ![Gaurav\_gupta](https://avatars.discourse-cdn.com/v4/letter/g/e274bd/32.png) [@Gaurav\_gupta](https://discuss.elastic.co/u/Gaurav_gupta)\
**Post date:** [August 10, 2015, 8:55am UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974/9 "2015-08-10T08:55:43Z")

</div>

No, we are just IPs as :- discovery.zen.ping.unicast.hosts=152.144.226.42,152.144.226.12. Generally, we start first node i.e. 152.144.226.42 fisrt and once this node is up we start 2nd node i.e. 152.144.226.12

Note :- We are unicast instead multicast.

Thanks

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 10, 2015, 10:17pm UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974/10 "2015-08-10T22:17:01Z")

</div>

If you want the most tolerable cluster then you will want dedicated masters.

---

<div class="post-metadata">

**Author:** ![Gaurav\_gupta](https://avatars.discourse-cdn.com/v4/letter/g/e274bd/32.png) [@Gaurav\_gupta](https://discuss.elastic.co/u/Gaurav_gupta)\
**Post date:** [August 18, 2015, 12:18pm UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974/11 "2015-08-18T12:18:05Z")

</div>

After discussing with user, I come to know they using the UNICAST with nodes like - discovery.zen.ping.unicast.hosts=10. **2**.10.185,10. **9**.10.185

And when they submit 5000 requests, they observe that after processing 2000 requests all further requests fails with error "_SearchPhaseExecutionException: Failed to execute phase [query\_fetch], all shards failed; shardFailures_"

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:55pm UTC](https://discuss.elastic.co/t/2-nodes-es-cluster-becomes-unavailable-for-2-3-mins-if-one-node-master-goes-down/26974/12 "2017-07-05T23:55:19Z")

</div>


