# Cluster unstable when recovering a node

**URL:** <https://discuss.elastic.co/t/cluster-unstable-when-recovering-a-node/8270>\
**Category:** Elasticsearch\
**Created:** [June 30, 2012, 7:48pm UTC](https://discuss.elastic.co/t/cluster-unstable-when-recovering-a-node/8270 "2012-06-30T19:48:25Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![rore](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rore/32/399_2.png) [@rore](https://discuss.elastic.co/u/rore)\
**Post date:** [June 30, 2012, 7:48pm UTC](https://discuss.elastic.co/t/cluster-unstable-when-recovering-a-node/8270/1 "2012-06-30T19:48:25Z")

</div>

I'm encountering again and again stability issues on the cluster when  
rebuilding one of the nodes, this happens while the shards are re-balancing.

In more details: I have a cluster with 3 nodes. Data is by time so indexes  
are created 1 per month, with 1 shard and 1 replica per index. Meaning each  
node contains about 2/3 of the indexes.  
This setup works great on general. Problems starts when one of the nodes  
goes down.

For instance, yesterday Amazon had issues in one AZ, which brought one of  
the nodes down for several hours, and I had to rebuild a new node from  
backup. So the new node was added to the cluster, and the cluster began  
rebalancing the shards.

Now while rebalancing occurred the cluster performance degraded severely.  
Indexing times went up (and even came to a halt occasionally), search times  
went app, meaning bad impact on our application. At some point another node  
stopped responding and was removed from the cluster, and I needed to  
restart that node also, which meant more rebalancing. So bringing the  
cluster back to a smooth running state takes a long couple of hours in  
which performance is between bad and terrible.

Is there any way to handle this issue? Am I missing something in the  
cluster configuration that can prevent these problems?

---

<div class="post-metadata">

**Author:** ![rore](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rore/32/399_2.png) [@rore](https://discuss.elastic.co/u/rore)\
**Post date:** [June 30, 2012, 7:48pm UTC](https://discuss.elastic.co/t/cluster-unstable-when-recovering-a-node/8270/2 "2012-06-30T19:48:54Z")

</div>

Forgot to mention, I'm using 0.19.2

On Saturday, June 30, 2012 10:48:25 PM UTC+3, Rotem wrote:

> I'm encountering again and again stability issues on the cluster when  
> rebuilding one of the nodes, this happens while the shards are re-balancing.
> 
> In more details: I have a cluster with 3 nodes. Data is by time so indexes  
> are created 1 per month, with 1 shard and 1 replica per index. Meaning each  
> node contains about 2/3 of the indexes.  
> This setup works great on general. Problems starts when one of the nodes  
> goes down.
> 
> For instance, yesterday Amazon had issues in one AZ, which brought one of  
> the nodes down for several hours, and I had to rebuild a new node from  
> backup. So the new node was added to the cluster, and the cluster began  
> rebalancing the shards.
> 
> Now while rebalancing occurred the cluster performance degraded severely.  
> Indexing times went up (and even came to a halt occasionally), search times  
> went app, meaning bad impact on our application. At some point another node  
> stopped responding and was removed from the cluster, and I needed to  
> restart that node also, which meant more rebalancing. So bringing the  
> cluster back to a smooth running state takes a long couple of hours in  
> which performance is between bad and terrible.
> 
> Is there any way to handle this issue? Am I missing something in the  
> cluster configuration that can prevent these problems?

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [July 2, 2012, 8:29am UTC](https://discuss.elastic.co/t/cluster-unstable-when-recovering-a-node/8270/3 "2012-07-02T08:29:51Z")

</div>

On Sat, 2012-06-30 at 12:48 -0700, Rotem wrote:

> Forgot to mention, I'm using 0.19.2

Yeah, copying large amounts of data causes a lot of I/O, which can  
degrade performance to the point of not being usable.

Version 0.19.5 comes with throttling, which gives you more control over  
how fast data is copied over:

> <https://github.com/elastic/elasticsearch/issues/2041>
>
> Allow to configure store throttling (only applied on file system based storage),… which allows to control the maximum bytes per sec written to the file system. It can be configured to only apply while merging, or on all output operations. The setting can eb set on the node level (in which case the throttling is done \_across\_ all shards allocated on the node), or index level, in which case it only applied to that index.
> 
> The node level settings are \`indices.store.throttle.type\` to set the type, with values of \`none\`, \`merge\` and \`all\` (defaults to \`none\`). And, also, \`indices.store.throttle.max\_bytes\_per\_sec\` (defaults to \`0\`), which can be set to something like \`1mb\`.
> 
> The index level settings is \`index.store.throttle.type\` for the type, with values of \`node\`, \`none\`, \`merge\`, and \`all\`. Defaults to \`node\` which will use the "shared" throttling on the node level. And, \`index.store.throttle.max\_bytes\_per\_sec\` (defaults to \`0\`).

Obviously, if you throttle, it'll take longer to recover, but if you  
don't, your cluster may become unusable while you're recovering 🙂

clint

> On Saturday, June 30, 2012 10:48:25 PM UTC+3, Rotem wrote:  
> I'm encountering again and again stability issues on the  
> cluster when rebuilding one of the nodes, this happens while  
> the shards are re-balancing.
> 
> ```
> In more details: I have a cluster with 3 nodes. Data is by
> time so indexes are created 1 per month, with 1 shard and 1
> replica per index. Meaning each node contains about 2/3 of the
> indexes.
> This setup works great on general. Problems starts when one of
> the nodes goes down.
>     
> For instance, yesterday Amazon had issues in one AZ, which
> brought one of the nodes down for several hours, and I had to
> rebuild a new node from backup. So the new node was added to
> the cluster, and the cluster began rebalancing the shards. 
>     
> Now while rebalancing occurred the cluster performance
> degraded severely. Indexing times went up (and even came to a
> halt occasionally), search times went app, meaning bad impact
> on our application. At some point another node stopped
> responding and was removed from the cluster, and I needed to
> restart that node also, which meant more rebalancing. So
> bringing the cluster back to a smooth running state takes a
> long couple of hours in which performance is between bad and
> terrible. 
>     
> Is there any way to handle this issue? Am I missing something
> in the cluster configuration that can prevent these problems?
> 
> ```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:21am UTC](https://discuss.elastic.co/t/cluster-unstable-when-recovering-a-node/8270/4 "2017-07-06T03:21:51Z")

</div>


