# Quickly restarting a node

**URL:** <https://discuss.elastic.co/t/quickly-restarting-a-node/172367>\
**Category:** Elasticsearch\
**Created:** [March 14, 2019, 3:16pm UTC](https://discuss.elastic.co/t/quickly-restarting-a-node/172367 "2019-03-14T15:16:10Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![yann-soubeyrand](https://avatars.discourse-cdn.com/v4/letter/y/ed8c4c/32.png) [@yann-soubeyrand](https://discuss.elastic.co/u/yann-soubeyrand)\
**Post date:** [March 14, 2019, 3:16pm UTC](https://discuss.elastic.co/t/quickly-restarting-a-node/172367/1 "2019-03-14T15:16:11Z")

</div>

Hi,

Trying to restart a node in our cluster as quickly as possible, I use the following procedure:

1. I disable shard allocation except for new primaries:

```auto
curl -s -u "$username:$password" -X 'PUT' -H 'Content-Type: application/json' -d '{ "transient": { "cluster.routing.allocation.enable": "new_primaries" } }' "$cluster_url/_cluster/settings"

```

1. I perform a synced flush:

```auto
curl -s -u "$username:$password" -X 'POST' "$cluster_url/_flush/synced"

```

1. I restart the node.
2. When the node joins the cluster, I re-enable shard allocation:

```auto
curl -s -u "$username:$password" -X 'PUT' -H 'Content-Type: application/json' -d '{ "transient": { "cluster.routing.allocation.enable": null } }' "$cluster_url/_cluster/settings"

```

My understanding is that cluster state should rapidly transition from yellow to green thanks to the synced flush. However, shard allocation hits throttling and is therefore slow as hell.

Did I miss something?

Our cluster currently contains too many shards and we are working toward reducing it. Will it solve our problem or is there other factors influencing node restart duration?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [March 14, 2019, 3:41pm UTC](https://discuss.elastic.co/t/quickly-restarting-a-node/172367/2 "2019-03-14T15:41:33Z")

</div>

Did the response to the synced flush indicate that it was completely successful? Did you stop indexing while the node was offline? If the answer to either question is no then it's possible that the synced flush marker isn't there on every shard (either it wasn't put in place, or it was put there and then removed) and this results in a slower recovery.

Which version are you using?

---

<div class="post-metadata">

**Author:** ![yann-soubeyrand](https://avatars.discourse-cdn.com/v4/letter/y/ed8c4c/32.png) [@yann-soubeyrand](https://discuss.elastic.co/u/yann-soubeyrand)\
**Post date:** [March 14, 2019, 4:19pm UTC](https://discuss.elastic.co/t/quickly-restarting-a-node/172367/3 "2019-03-14T16:19:47Z")

</div>

The synced flush indicates that almost every shards are successful: only 22 out of 22472 failed (we really have too many shards). Indexing wasn't stopped during node restart but only a small number of shards should be touched (I estimate the maximum number to be 642).

Having 6 data nodes, 3745 (22472 / 6) shards are unassigned after a node restart and I expect maximum 107 (642 / 6) shards to be slow recovering and the remaining shards to recover very quickly (as their flush marker shouldn't have changed).

For a shard which has been touched during node restart (resulting in its flush marker changing), is its recovery duration function of its size?

---

<div class="post-metadata">

**Author:** ![yann-soubeyrand](https://avatars.discourse-cdn.com/v4/letter/y/ed8c4c/32.png) [@yann-soubeyrand](https://discuss.elastic.co/u/yann-soubeyrand)\
**Post date:** [March 14, 2019, 4:33pm UTC](https://discuss.elastic.co/t/quickly-restarting-a-node/172367/4 "2019-03-14T16:33:42Z")

</div>

I forgot to mention that we are using version 6.6.1.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [March 14, 2019, 4:38pm UTC](https://discuss.elastic.co/t/quickly-restarting-a-node/172367/5 "2019-03-14T16:38:33Z")

</div>

> [@yann-soubeyrand](#):
>
> is its recovery duration function of its size?

It depends. In some recoveries Elasticsearch has to make a brand-new copy of the shard. It will re-use any segments that it can, but often there aren't many of these. This was the case for all recoveries in versions before 6.0, and is still the case in more recent versions if there's been too many changes (\>512MB of translog), or the node has been offline for too long (\>12h), or the new copy is assigned to a different node from the node that holds the previous, stale, copy of the shard.

> [@yann-soubeyrand](#):
>
> I expect maximum 107 (642 / 6) shards to be slow recovering and the remaining shards to recover very quickly (as their flush marker shouldn't have changed).

Is that different from what you're seeing? Are you seeing shards recover that you weren't expecting to need recovery?

---

<div class="post-metadata">

**Author:** ![yann-soubeyrand](https://avatars.discourse-cdn.com/v4/letter/y/ed8c4c/32.png) [@yann-soubeyrand](https://discuss.elastic.co/u/yann-soubeyrand)\
**Post date:** [March 14, 2019, 4:56pm UTC](https://discuss.elastic.co/t/quickly-restarting-a-node/172367/6 "2019-03-14T16:56:53Z")

</div>

I'm not sure of my interpretation of the /\_recovery informations here, but I see almost all our indices there.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 11, 2019, 4:56pm UTC](https://discuss.elastic.co/t/quickly-restarting-a-node/172367/7 "2019-04-11T16:56:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
