# Cluster Redundancy

**URL:** <https://discuss.elastic.co/t/cluster-redundancy/329415>\
**Category:** Elasticsearch\
**Created:** [April 5, 2023, 10:42am UTC](https://discuss.elastic.co/t/cluster-redundancy/329415 "2023-04-05T10:42:00Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![George\_Smith](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/george_smith/32/119458_2.png) [@George\_Smith](https://discuss.elastic.co/u/George_Smith)\
**Post date:** [April 5, 2023, 10:42am UTC](https://discuss.elastic.co/t/cluster-redundancy/329415/1 "2023-04-05T10:42:01Z")

</div>

Hi all,

I have a question regarding network redundancy with an Elasticsearch Cluster.  
I am attempting to add redundancy to my cluster so that if a network adapter a node is using fails, it can use another network adapter to continue communicating with the cluster.

The way I am attempting to do this is by using the special value `0.0.0.0` for `network.host` to bind all nodes in the cluster to all available network adapters on the machine the node is running on, then explicitly specifying the IP addresses for `network.bind_host` and `network.publish_host`.

For example, say I have a 3 node cluster with each node running on a different machine.  
The addresses/adapters available to the machines are both 10.0.2.xx and 10.0.3.xx .

Therefore for the machine where xx is 30, it's node YAML looks as follows:

```auto
network.host: 0.0.0.0
network.bind_host:
  - 10.0.2.30
  - 10.0.3.30
network.publish_host:
  - 10.0.2.30
  - 10.0.3.30

```

I would then expect that if I disabled the adapter on the machine with the address `10.0.2.30`, the node should still be able to communicate with the cluster using the `10.0.3.30`.  
Unfortunately this is not the case, and when I disable the adapter the node is unable to communicate with the rest of the cluster.

I have noticed that using `/_nodes/?pretty` to check if the node is configured correctly, this information looks correct:

```auto
"network" : {
          "host" : "0.0.0.0",
          "bind_host" : [
            "10.0.2.30",
            "10.0.3.30"
          ],
          "publish_host" : [
            "10.0.2.30",
            "10.0.3.30"
          ]
        }

```

However the publish address only shows one value:

```auto
"transport" : {
        "bound_address" : [
          "10.0.2.30:9202",
          "10.0.3.30:9202"
        ],
        "publish_address" : "10.0.3.30:9202",
        "profiles" : { }
      },
      "http" : {
        "bound_address" : [
          "10.0.2.30:9201",
          "10.0.3.30:9201"
        ],
        "publish_address" : "10.0.3.30:9201",
        "max_content_length_in_bytes" : 104857600
      },

```

Could this be why my node fails to communicate with the cluster when 1 of the two network adapters available fails or is there something else I am missing?

Thanks in advance 🙂

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 5, 2023, 12:40pm UTC](https://discuss.elastic.co/t/cluster-redundancy/329415/2 "2023-04-05T12:40:16Z")

</div>

Hi there @George_Smith. Thanks for the questions. I think [these docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/modules-network.html#advanced-network-settings) will help to answer them:

> You can specify a list of addresses for `network.host` and `network.publish_host`. [...] If you do this then Elasticsearch chooses one of the addresses for its publish address. This choice uses heuristics based on IPv4/IPv6 stack preference and reachability and may change when the node restarts. Ensure each node is accessible at all possible publish addresses.

Elasticsearch is based on redundancy at the _node_ level. If a NIC (or disk, or PSU, or anything else within a node) fails then ES treats that as a whole-node failure and acts accordingly.

You can pull some system-level tricks to handle this kind of failure more gracefully (e.g. using RAID for disks or bonding for NICs) but that sort of thing is transparent to ES.

---

<div class="post-metadata">

**Author:** ![George\_Smith](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/george_smith/32/119458_2.png) [@George\_Smith](https://discuss.elastic.co/u/George_Smith)\
**Post date:** [April 5, 2023, 1:09pm UTC](https://discuss.elastic.co/t/cluster-redundancy/329415/3 "2023-04-05T13:09:36Z")

</div>

Hi @DavidTurner, many thanks for your timely response!

I had a feeling this was how Elasticsearch treated failures, looks like we'll need to investigate NIC bonding in this case.

Alternatively if NIC bonding is not appropriate for our usecase, I think we can use a workaround whereby we create n elasticsearch instances/nodes where n is the number of adapters on the machine, with each instance/node bound to 1 of the n adapters which should give us the desired redundancy characteristics, albeit at the cost of increased resource usage.

Thanks!

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 5, 2023, 2:11pm UTC](https://discuss.elastic.co/t/cluster-redundancy/329415/4 "2023-04-05T14:11:23Z")

</div>

I'm curious whether there's anything special about your environment that makes it worth putting so much effort into mitigating the risk of NIC failure. NICs do fail occasionally for sure, but do they fail unusually often for you?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 5, 2023, 3:00pm UTC](https://discuss.elastic.co/t/cluster-redundancy/329415/5 "2023-04-05T15:00:46Z")

</div>

> [@George\_Smith](#):
>
> Alternatively if NIC bonding is not appropriate for our usecase, I think we can use a workaround whereby we create n elasticsearch instances/nodes where n is the number of adapters on the machine, with each instance/node bound to 1 of the n adapters which should give us the desired redundancy characteristics, albeit at the cost of increased resource usage.

This would be 3 separate nodes, each with their own data, and you would still lose one of them on NIC failure. I do not see how this is any better.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 3, 2023, 3:00pm UTC](https://discuss.elastic.co/t/cluster-redundancy/329415/6 "2023-05-03T15:00:56Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
