# Node upgraded 2.0.0 to 2.3.2 can't communicate with other nodes in cluster

**URL:** <https://discuss.elastic.co/t/node-upgraded-2-0-0-to-2-3-2-cant-communicate-with-other-nodes-in-cluster/55925>\
**Category:** Elasticsearch\
**Created:** [July 19, 2016, 8:36pm UTC](https://discuss.elastic.co/t/node-upgraded-2-0-0-to-2-3-2-cant-communicate-with-other-nodes-in-cluster/55925 "2016-07-19T20:36:06Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![jthoni](https://avatars.discourse-cdn.com/v4/letter/j/b9e5f3/32.png) [@jthoni](https://discuss.elastic.co/u/jthoni)\
**Post date:** [July 19, 2016, 8:36pm UTC](https://discuss.elastic.co/t/node-upgraded-2-0-0-to-2-3-2-cant-communicate-with-other-nodes-in-cluster/55925/1 "2016-07-19T20:36:06Z")

</div>

We have a 4-node Elasticsearch cluster running 2.0.0. We upgraded our test environment to 2.3.2, but are running into problems upgrading our production environment. I followed the process for rolling upgrade. When the upgraded node joins the cluster, queries on the site begin to fail to return. I assume this is happening when the request is route do the upgraded node.

It seems that the newly upgraded node is unable to communicate with the other nodes (i.e. the master), and we have it configured that a single node can't become master in the absence of a quorum (i.e. the "split brain" issue).

My full description is 4x the length allowed to post here, so the full details with log data can be found here:

[http://stackoverflow.com/questions/38464357/site-queries-fail-after-bringing-upgraded-elasticsearch-cluster-online](http://stackoverflow.com/questions/38464357/site-queries-fail-after-bringing-upgraded-elasticsearch-cluster-online)

Are there any known issues with bringing a node upgraded to 2.3.2 online in a cluster with other nodes at 2.0.0? When I brought the node down and back up with 2.0.0, all works fine. The config.yml and all environment settings (such as ES\_HEAP\_SIZE, which is at 7g on my system with 14 GB ram) are identical. The only thing that changes is the version of ES.

thanks! ~john

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [July 20, 2016, 7:42am UTC](https://discuss.elastic.co/t/node-upgraded-2-0-0-to-2-3-2-cant-communicate-with-other-nodes-in-cluster/55925/2 "2016-07-20T07:42:59Z")

</div>

Hi John,

I'd concentrate on the OutOfMemoryErrors. Can you first check that the 2.3.3 Elasticsearch process uses 7 GB heap space indeed? If you have a JDK on your production server you can check the exact command with all command line parameters for all Java processes on the machine with `jps -v`.

It should show something like:

```auto
18795 Elasticsearch -Xms256m -Xmx2g -XX:+UseConcMarkSweepGC -XX:CMSInitiatingOccupancyFraction=75 -XX:+UseCMSInitiatingOccupancyOnly -XX:+DisableExplicitGC -XX:+AlwaysPreTouch -Djava.awt.headless=true -Dfile.encoding=UTF-8 -Djna.nosys=true -XX:+HeapDumpOnOutOfMemoryError -Des.path.home=/Users/dm/tests/elasticsearch-5.0.0-alpha4

```

If jps is not available, you can also use `ps`.

If you're seeing indeed 7GB for `Xms`and `Xmx` then you can look why it is using so much memory. You can produce a heap dump when the OufOfMemoryError occurs by adding `-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/path/to/your/heapdump_file"` to `ES_JAVA_OPTS`.

Daniel

---

<div class="post-metadata">

**Author:** ![jthoni](https://avatars.discourse-cdn.com/v4/letter/j/b9e5f3/32.png) [@jthoni](https://discuss.elastic.co/u/jthoni)\
**Post date:** [July 20, 2016, 4:11pm UTC](https://discuss.elastic.co/t/node-upgraded-2-0-0-to-2-3-2-cant-communicate-with-other-nodes-in-cluster/55925/3 "2016-07-20T16:11:06Z")

</div>

I think all my logs above were red herrings. I brought the upgraded node up again, and again my site started failing, but there was nothing amiss in that nodes logs. I looked at the logs on the master node, and it was filled with exceptions pointing to my upgraded node. They were all the following (pasted w/o call stacks). Manta is the master and Unthinnk the upgraded node.

> [2016-07-20 15:45:48,589][WARN][gateway] [Manta] [ml\_v7][1]: failed to list shard for shard\_store on node [Qg3ghrCkT7mjUoc7hmItiA]  
> FailedNodeException[Failed node [Qg3ghrCkT7mjUoc7hmItiA]]; nested: RemoteTransportException[[Unthinnk][10.0.0.7:9300][internal:cluster/nodes/indices/shard/store[n]]]; nested: ElasticsearchException[Failed to list store metadata for shard [[ml\_v7][1]]]; nested: IndexFormatTooNewException[Format version is not supported (resource BufferedChecksumIndexInput(SimpleFSIndexInput(path="F:\data\elasticsearch\nodes\0\indices\ml\_v7\1\index\segments\_1s1"))): 6 (needs to be between 0 and 5)];  
> Caused by: RemoteTransportException[[Unthinnk][10.0.0.7:9300][internal:cluster/nodes/indices/shard/store[n]]]; nested: ElasticsearchException[Failed to list store metadata for shard [[ml\_v7][1]]]; nested: IndexFormatTooNewException[Format version is not supported (resource BufferedChecksumIndexInput(SimpleFSIndexInput(path="F:\data\elasticsearch\nodes\0\indices\ml\_v7\1\index\segments\_1s1"))): 6 (needs to be between 0 and 5)];  
> Caused by: ElasticsearchException[Failed to list store metadata for shard [[ml\_v7][1]]]; nested: IndexFormatTooNewException[Format version is not supported (resource BufferedChecksumIndexInput(SimpleFSIndexInput(path="F:\data\elasticsearch\nodes\0\indices\ml\_v7\1\index\segments\_1s1"))): 6 (needs to be between 0 and 5)];  
> Caused by: org.apache.lucene.index.IndexFormatTooNewException: Format version is not supported (resource BufferedChecksumIndexInput(SimpleFSIndexInput(path="F:\data\elasticsearch\nodes\0\indices\ml\_v7\1\index\segments\_1s1"))): 6 (needs to be between 0 and 5)

---

<div class="post-metadata">

**Author:** ![jthoni](https://avatars.discourse-cdn.com/v4/letter/j/b9e5f3/32.png) [@jthoni](https://discuss.elastic.co/u/jthoni)\
**Post date:** [July 20, 2016, 4:13pm UTC](https://discuss.elastic.co/t/node-upgraded-2-0-0-to-2-3-2-cant-communicate-with-other-nodes-in-cluster/55925/5 "2016-07-20T16:13:17Z")

</div>

I see that we have this:

ES 2.0.0 --\> Lucene 5.2.1  
ES 2.3.2 --\> Lucene 5.5.0

Is there any compatibility issues between bringing a node up with 5.5.0 with the rest of the cluster 5.2.1?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 21, 2016, 8:49am UTC](https://discuss.elastic.co/t/node-upgraded-2-0-0-to-2-3-2-cant-communicate-with-other-nodes-in-cluster/55925/6 "2016-07-21T08:49:02Z")

</div>

Shards that have been allocated to the newer node are upgraded and can subsequently not be allocated back to nodes with lower version, so you should complete the rolling upgrade to ensure that all nodes in the cluster are running the same Elasticsearch version.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [July 21, 2016, 9:32am UTC](https://discuss.elastic.co/t/node-upgraded-2-0-0-to-2-3-2-cant-communicate-with-other-nodes-in-cluster/55925/7 "2016-07-21T09:32:43Z")

</div>

@jthoni When doing a rolling upgrade, Elasticsearch should not relocate indices from new nodes to old nodes, so you shouldn't get the format version exceptions that you're seeing. Did you start a 2.3.2 node then stop it and go back to 2.0.0 on the same box? If so, that would account for the "format version too new" exceptions.

---

<div class="post-metadata">

**Author:** ![jthoni](https://avatars.discourse-cdn.com/v4/letter/j/b9e5f3/32.png) [@jthoni](https://discuss.elastic.co/u/jthoni)\
**Post date:** [July 21, 2016, 7:44pm UTC](https://discuss.elastic.co/t/node-upgraded-2-0-0-to-2-3-2-cant-communicate-with-other-nodes-in-cluster/55925/8 "2016-07-21T19:44:35Z")

</div>

Well, that would explain the exceptions. I took the node down, then brought it up with 2.3.2. All queries to the production cluster started failing, so I immediately took it down and brought it back up with 2.0.0. I guess the index got flagged as 2.3.2, so that accounts for the exception.

---

<div class="post-metadata">

**Author:** ![jthoni](https://avatars.discourse-cdn.com/v4/letter/j/b9e5f3/32.png) [@jthoni](https://discuss.elastic.co/u/jthoni)\
**Post date:** [July 21, 2016, 7:45pm UTC](https://discuss.elastic.co/t/node-upgraded-2-0-0-to-2-3-2-cant-communicate-with-other-nodes-in-cluster/55925/9 "2016-07-21T19:45:47Z")

</div>

I would assume, therefore, that data in the now 2.3.2 flagged shard will not be getting replicated to another node (as all the others are 2.0.0), right?

---

<div class="post-metadata">

**Author:** ![jthoni](https://avatars.discourse-cdn.com/v4/letter/j/b9e5f3/32.png) [@jthoni](https://discuss.elastic.co/u/jthoni)\
**Post date:** [July 21, 2016, 7:48pm UTC](https://discuss.elastic.co/t/node-upgraded-2-0-0-to-2-3-2-cant-communicate-with-other-nodes-in-cluster/55925/10 "2016-07-21T19:48:19Z")

</div>

To verify, I should be able to use [Rolling Upgrade](https://www.elastic.co/guide/en/elasticsearch/reference/current/rolling-upgrades.html) to go from 2.0.0 to 2.3.2 without taking down the entire cluster, right? I am now hesitant to experiment as it is having adverse effects on production (note that the upgrade worked fine in our test environment).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:33pm UTC](https://discuss.elastic.co/t/node-upgraded-2-0-0-to-2-3-2-cant-communicate-with-other-nodes-in-cluster/55925/11 "2017-07-05T22:33:35Z")

</div>


