# \[SOLVED\] Frequent node disconnects on Rackspace environment

**URL:** <https://discuss.elastic.co/t/solved-frequent-node-disconnects-on-rackspace-environment/34304>\
**Category:** Elasticsearch\
**Created:** [November 11, 2015, 8:15am UTC](https://discuss.elastic.co/t/solved-frequent-node-disconnects-on-rackspace-environment/34304 "2015-11-11T08:15:26Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![buinauskas\_evaldas](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/buinauskas_evaldas/32/31439_2.png) [@buinauskas\_evaldas](https://discuss.elastic.co/u/buinauskas_evaldas)\
**Post date:** [November 11, 2015, 8:15am UTC](https://discuss.elastic.co/t/solved-frequent-node-disconnects-on-rackspace-environment/34304/1 "2015-11-11T08:15:26Z")

</div>

We have Elasticsearch cluster deployed in Rackspace. Each machine has it's own Server created (Windows Server 2012 R2).

We have three nodes with following `elasticsearch.yml`:

```
action.disable_delete_all_indices: true

cluster.name: ClusterUK

network.publish_host: "172.24.32.10"

discovery.zen.ping.timeout: "30s"
discovery.zen.ping_timeout: "30s"
discovery.zen.minimum_master_nodes: 2
discovery.zen.ping.multicast.enabled: false
discovery.zen.ping.unicast.hosts: ["172.24.32.10", "172.24.32.5", "172.24.32.8"]

indices.fielddata.cache.size: 25%
indices.cluster.send_refresh_mapping: false

node.name: "ClusterUK Node 1" 
node.master: true
node.data: true

bootstrap.mlockall: true

```

And that's the logs it's producing:

```
[2015-11-11 07:39:37,615][INFO][http] [ClusterUK Node 1] bound_address {inet[/0:0:0:0:0:0:0:0:9200]}, publish_address {inet[/172.24.32.10:9200]}
[2015-11-11 07:39:37,615][INFO][node] [ClusterUK Node 1] started
[2015-11-11 07:39:38,896][INFO][discovery.zen] [ClusterUK Node 1] failed to send join request to master [[ClusterUK Node 1][Ar_pY4NNRBWwTbv9fV226w][elasticuk1][inet[/172.24.32.10:9300]]{master=true}], reason [RemoteTransportException[[ClusterUK Node 1][inet[/172.24.32.10:9300]][internal:discovery/zen/join]]; nested: ElasticsearchIllegalStateException[Node [[ClusterUK Node 1][z2poU5hqQT-VmBKJifD0-w][elasticuk1][inet[/172.24.32.10:9300]]{master=true}] not master for join request from [[ClusterUK Node 1][z2poU5hqQT-VmBKJifD0-w][elasticuk1][inet[/172.24.32.10:9300]]{master=true}]]; ], tried [3] times
[2015-11-11 07:40:09,974][INFO][cluster.service] [ClusterUK Node 1] detected_master [ClusterUK Node 3][m5ns1sKHTDSSdbBMWNsqwA][elasticuk3][inet[/172.24.32.8:9300]]{master=true}, added {[ClusterUK Node 3][m5ns1sKHTDSSdbBMWNsqwA][elasticuk3][inet[/172.24.32.8:9300]]{master=true},[ClusterUK Client Node STG1][Uxmn2i1iSpuxlp3IgjNNdQ][Staging1][inet[/192.168.100.248:9300]]{data=false, master=false},}, reason: zen-disco-receive(from master [[ClusterUK Node 3][m5ns1sKHTDSSdbBMWNsqwA][elasticuk3][inet[/172.24.32.8:9300]]{master=true}])
[2015-11-11 07:42:06,756][INFO][cluster.service] [ClusterUK Node 1] added {[ClusterUK Node 2][UKA81JAURsquFqvH7xiAFg][elasticuk2][inet[/172.24.32.5:9300]]{master=true},}, reason: zen-disco-receive(from master [[ClusterUK Node 3][m5ns1sKHTDSSdbBMWNsqwA][elasticuk3][inet[/172.24.32.8:9300]]{master=true}])
[2015-11-11 08:00:37,378][INFO][discovery.zen] [ClusterUK Node 1] master_left [[ClusterUK Node 3][m5ns1sKHTDSSdbBMWNsqwA][elasticuk3][inet[/172.24.32.8:9300]]{master=true}], reason [transport disconnected]
[2015-11-11 08:00:37,380][WARN][discovery.zen] [ClusterUK Node 1] master left (reason = transport disconnected), current nodes: {[ClusterUK Node 2][UKA81JAURsquFqvH7xiAFg][elasticuk2][inet[/172.24.32.5:9300]]{master=true},[ClusterUK Node 1][z2poU5hqQT-VmBKJifD0-w][elasticuk1][inet[elasticuk1/172.24.32.10:9300]]{master=true},[ClusterUK Client Node STG1][Uxmn2i1iSpuxlp3IgjNNdQ][Staging1][inet[/192.168.100.248:9300]]{data=false, master=false},}
[2015-11-11 08:00:37,380][INFO][cluster.service] [ClusterUK Node 1] removed {[ClusterUK Node 3][m5ns1sKHTDSSdbBMWNsqwA][elasticuk3][inet[/172.24.32.8:9300]]{master=true},}, reason: zen-disco-master_failed ([ClusterUK Node 3][m5ns1sKHTDSSdbBMWNsqwA][elasticuk3][inet[/172.24.32.8:9300]]{master=true})
[2015-11-11 08:00:37,985][ERROR][marvel.agent.exporter] [ClusterUK Node 1] remote target didn't respond with 200 OK response code [503 Service Unavailable]. content: [:)
��error�ClusterBlockException[blocked by: [SERVICE_UNAVAILABLE/2/no master];]��status$��]
[2015-11-11 08:00:47,996][ERROR][marvel.agent.exporter] [ClusterUK Node 1] remote target didn't respond with 200 OK response code [503 Service Unavailable]. content: [:)
��error�ClusterBlockException[blocked by: [SERVICE_UNAVAILABLE/2/no master];]��status$��]
[2015-11-11 08:01:07,407][INFO][cluster.service] [ClusterUK Node 1] detected_master [ClusterUK Node 3][m5ns1sKHTDSSdbBMWNsqwA][elasticuk3][inet[/172.24.32.8:9300]]{master=true}, added {[ClusterUK Node 3][m5ns1sKHTDSSdbBMWNsqwA][elasticuk3][inet[/172.24.32.8:9300]]{master=true},}, reason: zen-disco-receive(from master [[ClusterUK Node 3][m5ns1sKHTDSSdbBMWNsqwA][elasticuk3][inet[/172.24.32.8:9300]]{master=true}]) 

```

It seems that master node disconnects for a second and then joins the cluster back. This causes data loss if bulk-inserts are being performed and may lead to split-brain. Does anyone know what's the root cause and how this can be fixed?

Version: 1.7.3

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [November 11, 2015, 8:18am UTC](https://discuss.elastic.co/t/solved-frequent-node-disconnects-on-rackspace-environment/34304/2 "2015-11-11T08:18:52Z")

</div>

Firewall?  
Are you monitoring the network?

---

<div class="post-metadata">

**Author:** ![buinauskas\_evaldas](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/buinauskas_evaldas/32/31439_2.png) [@buinauskas\_evaldas](https://discuss.elastic.co/u/buinauskas_evaldas)\
**Post date:** [November 11, 2015, 8:23am UTC](https://discuss.elastic.co/t/solved-frequent-node-disconnects-on-rackspace-environment/34304/3 "2015-11-11T08:23:20Z")

</div>

Firewall is turned off. Network monitor is turned off too.

Looking at transport issue, i found out that it's using TCP and it might be useful to disable TCP Offload in adapter settings. Article here: [http://www.rackspace.com/knowledge\_center/article/disabling-tcp-offloading-in-windows-server-2012](http://www.rackspace.com/knowledge_center/article/disabling-tcp-offloading-in-windows-server-2012)

Trying it now. Will update

---

<div class="post-metadata">

**Author:** ![buinauskas\_evaldas](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/buinauskas_evaldas/32/31439_2.png) [@buinauskas\_evaldas](https://discuss.elastic.co/u/buinauskas_evaldas)\
**Post date:** [November 11, 2015, 10:32am UTC](https://discuss.elastic.co/t/solved-frequent-node-disconnects-on-rackspace-environment/34304/4 "2015-11-11T10:32:20Z")

</div>

It was [TCP Offloading](http://www.rackspace.com/knowledge_center/article/disabling-tcp-offloading-in-windows-server-2012).

> TCP offload engine is a function used in network interface cards (NIC)  
> to offload processing of the entire TCP/IP stack to the network  
> controller. By moving some or all of the processing to dedicated  
> hardware, a TCP offload engine frees the system's main CPU for other  
> tasks. However, TCP offloading has been known to cause some issues,  
> and disabling it can help avoid these issues.

###Disable TCP Offloading

1. In the Windows server, open the Control Panel and select **Network  
Settings **\>** Change Adapter Settings**.

[![Screenshot](https://us1.discourse-cdn.com/elastic/original/3X/a/3/a3a9e438e07917306f334ff028cf06aea57e1aec.png)](http://i.stack.imgur.com/8oFu9.png)

1. Right-click on each of the adapters ( **private** and **public** ), select  
**Configure** from the **Networking** menu, and then click the **Advanced** tab.  
The TCP offload settings are listed for the Citrix adapter.

[![Screenshot](https://us1.discourse-cdn.com/elastic/original/3X/3/4/34103a582a784951fb419d5fcee00096a6540d6a.png)](http://i.stack.imgur.com/xF6qv.png)

1. Disable each of the following TCP offload options, and then click  
**OK** :

- IPv4 Checksum Offload
- Large Receive Offload
- Large Send Offload
- TCP Checksum Offload

This solved my issue.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:39pm UTC](https://discuss.elastic.co/t/solved-frequent-node-disconnects-on-rackspace-environment/34304/5 "2017-07-05T23:39:17Z")

</div>


