# During a rolling restart sometimes all replicas of a single shard go into PRIMARY\_FAILED

**URL:** https://discuss.elastic.co/t/during-a-rolling-restart-sometimes-all-replicas-of-a-single-shard-go-into-primary-failed/299269
**Category:** Elasticsearch
**Created:** [March 9, 2022, 11:09pm UTC](https://discuss.elastic.co/t/during-a-rolling-restart-sometimes-all-replicas-of-a-single-shard-go-into-primary-failed/299269 "2022-03-09T23:09:33Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![rmb938](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rmb938/32/102838_2.png) [@rmb938](https://discuss.elastic.co/u/rmb938)
#### Post date: [March 9, 2022, 11:09pm UTC](https://discuss.elastic.co/t/during-a-rolling-restart-sometimes-all-replicas-of-a-single-shard-go-into-primary-failed/299269/1 "2022-03-09T23:09:33Z")

</div>

I have an issue where sometimes during a rolling restart when it gets to a node that has a primary replica once the node is offline the replica shards go into a `PRIMARY_FAILED` state.

i.e

```auto
my-index 11 p UNASSIGNED NODE_LEFT 
my-index 11 r UNASSIGNED PRIMARY_FAILED
my-index 11 r UNASSIGNED PRIMARY_FAILED

```

This doesn't seem to happen all the time and I can't really find a way to make it consistently happen.

According to the documentation this means `The shard was initializing as a replica, but the primary shard failed before the initialization completed.`

How do I prevent this? I am restarting one node at a time and wiating for all shards to be allocated and have the cluster in a green state before moving onto the next node. Shard allocation is turned off before each node is taken down and turned back on when brought online.

I can't really find any documentation that says how to prevent this and I am following all the steps here [Full-cluster restart and rolling restart | Elasticsearch Guide [8.1] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/restart-cluster.html#restart-cluster-rolling)

So I am not really sure what is causing this. Any ideas or insights would be great!

---

<div class="post-metadata">

### Author: ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)
#### Post date: [March 9, 2022, 11:23pm UTC](https://discuss.elastic.co/t/during-a-rolling-restart-sometimes-all-replicas-of-a-single-shard-go-into-primary-failed/299269/2 "2022-03-09T23:23:37Z")

</div>

Welcome to our community! 😃

What do the logs on your master node show at this time for that index?

---

<div class="post-metadata">

### Author: ![rmb938](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rmb938/32/102838_2.png) [@rmb938](https://discuss.elastic.co/u/rmb938)
#### Post date: [March 10, 2022, 12:45am UTC](https://discuss.elastic.co/t/during-a-rolling-restart-sometimes-all-replicas-of-a-single-shard-go-into-primary-failed/299269/3 "2022-03-10T00:45:18Z")

</div>

Unfortunately not much.

I see shard allocation being turned off, the node leaving the cluster, logs about marking unavailable shards as stale (posted bellow), then a bit later the node joining the cluster and allocation being turned back on.

```auto
{"type": "server", "timestamp": "2022-03-09T22:40:52,226Z", "level": "WARN", "component": "o.e.c.r.a.AllocationService", "cluster.name": "my-cluster", "node.name": "my-cluster-es-master-4", "message": "[my-index][7] marking unavailable shards as stale: [cyzHRssCRd-PJ8FYF9zAGQ]", "cluster.uuid": "nVZb27XkRkmc5vsbGxqfng", "node.id": "jZz5usD4SSKR4hwrJRJCtw" }
{"type": "server", "timestamp": "2022-03-09T22:40:53,107Z", "level": "WARN", "component": "o.e.c.r.a.AllocationService", "cluster.name": "my-cluster", "node.name": "my-cluster-es-master-4", "message": "[my-index][2] marking unavailable shards as stale: [iRQsE2dkQ2qgoKJup7PmFw]", "cluster.uuid": "nVZb27XkRkmc5vsbGxqfng", "node.id": "jZz5usD4SSKR4hwrJRJCtw" }
{"type": "server", "timestamp": "2022-03-09T22:40:53,456Z", "level": "WARN", "component": "o.e.c.r.a.AllocationService", "cluster.name": "my-cluster", "node.name": "my-cluster-es-master-4", "message": "[my-index][6] marking unavailable shards as stale: [IABho5n9TJqOdYFy_K90Yw]", "cluster.uuid": "nVZb27XkRkmc5vsbGxqfng", "node.id": "jZz5usD4SSKR4hwrJRJCtw" }
{"type": "server", "timestamp": "2022-03-09T22:40:53,491Z", "level": "WARN", "component": "o.e.c.r.a.AllocationService", "cluster.name": "my-cluster", "node.name": "my-cluster-es-master-4", "message": "[my-index][11] marking unavailable shards as stale: [ChD9fRVhSyeh7EuGT_N_Fg]", "cluster.uuid": "nVZb27XkRkmc5vsbGxqfng", "node.id": "jZz5usD4SSKR4hwrJRJCtw" }

```

No logs in any of my master nodes about anything else. I'm using the default out of the box logging setup.

---

<div class="post-metadata">

### Author: ![rmb938](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rmb938/32/102838_2.png) [@rmb938](https://discuss.elastic.co/u/rmb938)
#### Post date: [March 16, 2022, 11:43pm UTC](https://discuss.elastic.co/t/during-a-rolling-restart-sometimes-all-replicas-of-a-single-shard-go-into-primary-failed/299269/4 "2022-03-16T23:43:28Z")

</div>

@warkolm Any ideas on what to check next?

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [April 13, 2022, 11:43pm UTC](https://discuss.elastic.co/t/during-a-rolling-restart-sometimes-all-replicas-of-a-single-shard-go-into-primary-failed/299269/5 "2022-04-13T23:43:48Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
