# Restarting of node taking much time

**URL:** <https://discuss.elastic.co/t/restarting-of-node-taking-much-time/13884>\
**Category:** Elasticsearch\
**Created:** [October 9, 2013, 12:06pm UTC](https://discuss.elastic.co/t/restarting-of-node-taking-much-time/13884 "2013-10-09T12:06:35Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ankit\_Jain](https://avatars.discourse-cdn.com/v4/letter/a/96bed5/32.png) [@Ankit\_Jain](https://discuss.elastic.co/u/Ankit_Jain)\
**Post date:** [October 9, 2013, 12:06pm UTC](https://discuss.elastic.co/t/restarting-of-node-taking-much-time/13884/1 "2013-10-09T12:06:35Z")

</div>

Hi All,

We have deployed 5 nodes cluster and each node is serving around 200 shards  
(total number of indices are 200 and each index has 200 shards).

While restarting each node is taking around 20 to 30 minutes to move all  
shards from unassigned state to assigned state.

How we can quickly move all the shards from unassigned state to assigned  
state?

Is the recovery time dependent on data size?

Regards,  
Ankit Jain

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [October 9, 2013, 2:45pm UTC](https://discuss.elastic.co/t/restarting-of-node-taking-much-time/13884/2 "2013-10-09T14:45:53Z")

</div>

Are the restarts planned or are they server crashes? If they are planned,  
you should disabling indexing (if possible), flush the index and  
temporarily disable allocation.

Here is a repost of something I wrote two days ago:

Elasticsearch has been throttling I/O recovery since version 0.90. The  
defaults are fairly low. Trying increasing the  
indices.recovery.max\_bytes\_per\_sec  
setting.

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

You can also increase the number of shards that are recovered at the same  
time. The default is 2. Increase either value too much and you will have  
long IO waits.

Cheers,

Ivan

On Wed, Oct 9, 2013 at 5:06 AM, Ankit Jain [ankitjaincs06@gmail.com](mailto:ankitjaincs06@gmail.com) wrote:

> Hi All,
> 
> We have deployed 5 nodes cluster and each node is serving around 200  
> shards (total number of indices are 200 and each index has 200 shards).
> 
> While restarting each node is taking around 20 to 30 minutes to move all  
> shards from unassigned state to assigned state.
> 
> How we can quickly move all the shards from unassigned state to assigned  
> state?
> 
> Is the recovery time dependent on data size?
> 
> Regards,  
> Ankit Jain
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ankit\_Jain](https://avatars.discourse-cdn.com/v4/letter/a/96bed5/32.png) [@Ankit\_Jain](https://discuss.elastic.co/u/Ankit_Jain)\
**Post date:** [October 9, 2013, 4:54pm UTC](https://discuss.elastic.co/t/restarting-of-node-taking-much-time/13884/3 "2013-10-09T16:54:58Z")

</div>

Thanks Ivan.

We are planning to store TB's of data into Elasticsearch cluster. I don't  
think \*indices.recovery.max\_bytes\_\*\*per\_sec \*would going to help me much.

Is any specific files which are required ES node during restart or any  
specific metadata information required during restart?

Also, can you suggest some alternative ways to improve node recovery time?

Regards,  
Ankit Jain

On Wednesday, 9 October 2013 20:15:53 UTC+5:30, Ivan Brusic wrote:

> Are the restarts planned or are they server crashes? If they are planned,  
> you should disabling indexing (if possible), flush the index and  
> temporarily disable allocation.
> 
> Here is a repost of something I wrote two days ago:
> 
> Elasticsearch has been throttling I/O recovery since version 0.90. The  
> defaults are fairly low. Trying increasing the indices.recovery.max\_bytes\_per\_sec  
> setting.
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/cluster-update-settings.html)
> 
> You can also increase the number of shards that are recovered at the same  
> time. The default is 2. Increase either value too much and you will have  
> long IO waits.
> 
> Cheers,
> 
> Ivan
> 
> On Wed, Oct 9, 2013 at 5:06 AM, Ankit Jain \<[ankitj...@gmail.com](mailto:ankitj...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > Hi All,
> > 
> > We have deployed 5 nodes cluster and each node is serving around 200  
> > shards (total number of indices are 200 and each index has 200 shards).
> > 
> > While restarting each node is taking around 20 to 30 minutes to move all  
> > shards from unassigned state to assigned state.
> > 
> > How we can quickly move all the shards from unassigned state to assigned  
> > state?
> > 
> > Is the recovery time dependent on data size?
> > 
> > Regards,  
> > Ankit Jain
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [October 9, 2013, 5:59pm UTC](https://discuss.elastic.co/t/restarting-of-node-taking-much-time/13884/4 "2013-10-09T17:59:29Z")

</div>

You have decided to put 200 shards on a single node. Such a high number  
combined with a significant shard size can not be recovered very quick.  
Although you can put many thousands shards on a single node and shards can  
grow into many hundreds of GB, you always have a price to pay when the  
shards start moving around between nodes at recovery time.

The improvement depends on how much you want to stress a node while  
recovering. The default settings are chosen wisely so when a node comes up,  
it can always respond to search and index requests immediately while  
recovery. It does not confuse clients with timeouts, and it does not  
confuse sysadmins with iowaits. But there is a long down time, and this  
down time is directly configured by the number of shards per node and the  
current shard size.

So if you want to stress your nodes, you can select a higher value in  
cluster.routing.allocation.node\_concurrent\_recoveries.

[http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/modules-cluster.html](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/modules-cluster.html)

Pros:

- recovery may be faster

Cons

- recovery takes many network resources
- queries and indexing may not respond in time
- higher iowaits
- network bandwidth must be available

The best method to achieve quick recovery is to select a wise shards per  
node ratio and a sane shard size.

Here is my approach: on my 32-core machines, I plan to never run more than  
32 shards, and no shard shall grow beyond 5-10g, so transporting on a  
10GBit/s takes reasonable time.

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![shadyabhi](https://avatars.discourse-cdn.com/v4/letter/s/edb3f5/32.png) [@shadyabhi](https://discuss.elastic.co/u/shadyabhi)\
**Post date:** [October 10, 2013, 7:50am UTC](https://discuss.elastic.co/t/restarting-of-node-taking-much-time/13884/5 "2013-10-10T07:50:27Z")

</div>

How are you restarting your nodes? Are you using via init scripts or

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

? Using the API seems to do pretty fast recovery for me.

On Wed, Oct 9, 2013 at 5:36 PM, Ankit Jain [ankitjaincs06@gmail.com](mailto:ankitjaincs06@gmail.com) wrote:

> Hi All,
> 
> We have deployed 5 nodes cluster and each node is serving around 200 shards  
> (total number of indices are 200 and each index has 200 shards).
> 
> While restarting each node is taking around 20 to 30 minutes to move all  
> shards from unassigned state to assigned state.
> 
> How we can quickly move all the shards from unassigned state to assigned  
> state?
> 
> Is the recovery time dependent on data size?
> 
> Regards,  
> Ankit Jain
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
Regards,  
Abhijeet Rastogi (shadyabhi)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ankit\_Jain](https://avatars.discourse-cdn.com/v4/letter/a/96bed5/32.png) [@Ankit\_Jain](https://discuss.elastic.co/u/Ankit_Jain)\
**Post date:** [October 10, 2013, 10:38am UTC](https://discuss.elastic.co/t/restarting-of-node-taking-much-time/13884/6 "2013-10-10T10:38:56Z")

</div>

Thanks Jorg and Abhijeet

@Abhijeet we are taking scenario of server crashes.  
Also, the amount data server by each shard is around 100 GB.

Regards,  
Ankit Jain

On Thursday, 10 October 2013 13:20:27 UTC+5:30, Abhijeet Rastogi wrote:

> How are you restarting your nodes? Are you using via init scripts or
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/cluster-nodes-shutdown.html)  
> ? Using the API seems to do pretty fast recovery for me.
> 
> On Wed, Oct 9, 2013 at 5:36 PM, Ankit Jain \<[ankitj...@gmail.com](mailto:ankitj...@gmail.com)\<javascript:\>\>  
> wrote:
> 
> > Hi All,
> > 
> > We have deployed 5 nodes cluster and each node is serving around 200  
> > shards  
> > (total number of indices are 200 and each index has 200 shards).
> > 
> > While restarting each node is taking around 20 to 30 minutes to move all  
> > shards from unassigned state to assigned state.
> > 
> > How we can quickly move all the shards from unassigned state to assigned  
> > state?
> > 
> > Is the recovery time dependent on data size?
> > 
> > Regards,  
> > Ankit Jain
> > 
> > --  
> > You received this message because you are subscribed to the Google  
> > Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send  
> > an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> Regards,  
> Abhijeet Rastogi (shadyabhi)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:13am UTC](https://discuss.elastic.co/t/restarting-of-node-taking-much-time/13884/7 "2017-07-06T02:13:03Z")

</div>


