# Help for removing a crashed node?

**URL:** <https://discuss.elastic.co/t/help-for-removing-a-crashed-node/43456>\
**Category:** Elasticsearch\
**Created:** [March 3, 2016, 11:46pm UTC](https://discuss.elastic.co/t/help-for-removing-a-crashed-node/43456 "2016-03-03T23:46:28Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![boreal](https://avatars.discourse-cdn.com/v4/letter/b/cab0a1/32.png) [@boreal](https://discuss.elastic.co/u/boreal)\
**Post date:** [March 3, 2016, 11:46pm UTC](https://discuss.elastic.co/t/help-for-removing-a-crashed-node/43456/1 "2016-03-03T23:46:28Z")

</div>

Hi,

I have a two node cluster, and master node crashed while running update on an index.  
Number of replicas is set to 1. I want to remove sick node, then add new node. I saw a page for how to shutdown a node. I didn't see complete step-by-step guide for removing a sick node, so I wanted to ask here:

1. After I ran "shutdown" on master node, I can just halt elasticsearch process on that node?  
2.If I ran "shutdown" on master node, will other node start throwing error because there is a number of replicas that it is set to?
2. Can I delete problematic index on a healthy node as nothing have happened?

Thanks a lot!  
b

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 4, 2016, 1:27am UTC](https://discuss.elastic.co/t/help-for-removing-a-crashed-node/43456/2 "2016-03-04T01:27:53Z")

</div>

If you shutdown the bad node you can just replace it. As you have replicas your data will be safe and it will copy part of it over to the new node when it joins.

---

<div class="post-metadata">

**Author:** ![boreal](https://avatars.discourse-cdn.com/v4/letter/b/cab0a1/32.png) [@boreal](https://discuss.elastic.co/u/boreal)\
**Post date:** [March 4, 2016, 7:27am UTC](https://discuss.elastic.co/t/help-for-removing-a-crashed-node/43456/3 "2016-03-04T07:27:09Z")

</div>

Thanks for your advice.

Well, the sick node seemed to have recovered itself. However, now the healthy node cannot connect to the previously sick node anymore. I'm getting an error below.

**Node 1(previously had OOM error)**

> [2016-03-04 01:18:46,067][WARN][shield.transport.netty] [node-1] exception caught on transport layer [[id: 0x5d126fb9, /MYIPForNode2:38674 :\> /MyIPForNode1:9300]], closing connection  
> java.io.StreamCorruptedException: invalid internal transport message format, got (ff,f4,ff,fd)  
> at org.elasticsearch.transport.netty.SizeHeaderFrameDecoder.decode(SizeHeaderFrameDecoder.java:64)  
> at org.jboss.netty.handler.codec.frame.FrameDecoder.callDecode(FrameDecoder.java:425)

**Node 2 (healthy node)**

> [2016-03-04 01:18:26,842][INFO][discovery.zen] [node-2] failed to send join request to master [{node-1}{MUm8WoRyS9Gi5IdIncCe-w}{MyIPForNode1}{MyIPForNode1:9300}{master=true}], reason [RemoteTransportException[[node-1][MyIPForNode1:9300][internal:discovery/zen/join]]; nested: ConnectTransportException[[node-2][MyIPForNode2:9300] connect\_timeout[30s]]; ]

Also, on the previously sick node I deleted the index that was problematic. However, it seems that update to that deleted index is still happening, as I see these exceptions in the log.

> [2016-03-04 00:38:03,813][INFO][rest.suppressed] /.marvel-es-data/cluster\_info/\_search Params: {index=.marvel-es-data, type=cluster\_info}  
> [.marvel-es-data] IndexNotFoundException[no such index]

Any idea as to how to stop this? I am thinking maybe I should still shut down the previously bad node anyway.

Thanks in advance!

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 4, 2016, 7:45am UTC](https://discuss.elastic.co/t/help-for-removing-a-crashed-node/43456/4 "2016-03-04T07:45:42Z")

</div>

> [@boreal](#):
>
> 169.53.138.84

Are you on Windows? That looks like an IP Windows assigns to itself when it cannot get a valid one.

---

<div class="post-metadata">

**Author:** ![boreal](https://avatars.discourse-cdn.com/v4/letter/b/cab0a1/32.png) [@boreal](https://discuss.elastic.co/u/boreal)\
**Post date:** [March 4, 2016, 7:51am UTC](https://discuss.elastic.co/t/help-for-removing-a-crashed-node/43456/5 "2016-03-04T07:51:49Z")

</div>

No, this is Ubuntu launched through Softlayer. Thanks!

I at least found why updating index caused OOM -- it was "Mapping Explosion" problem mentioned here:

> **[Six Ways to Crash Elasticsearch
	  	 | Elastic](https://www.elastic.co/blog/found-crash-elasticsearch)**
>
> As much as we love Elasticsearch, at Found we've seen customers crash their clusters in numerous ways. Mostly due to simple misunderstandings and usually the fixes are fairly straightforward. In our quest to enlighten new adopters and entertain the...

But at least, I deleted the entire index which had too many mapping though.. but the effect still seems to be continuing.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:11pm UTC](https://discuss.elastic.co/t/help-for-removing-a-crashed-node/43456/6 "2017-07-05T23:11:20Z")

</div>


