# Node.close() gets stuck in the 'stopping' state

**URL:** <https://discuss.elastic.co/t/node-close-gets-stuck-in-the-stopping-state/3992>\
**Category:** Elasticsearch\
**Created:** [February 27, 2011, 3:16pm UTC](https://discuss.elastic.co/t/node-close-gets-stuck-in-the-stopping-state/3992 "2011-02-27T15:16:51Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jacob\_Perkins](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jacob_perkins/32/3370_2.png) [@Jacob\_Perkins](https://discuss.elastic.co/u/Jacob_Perkins)\
**Post date:** [February 27, 2011, 3:16pm UTC](https://discuss.elastic.co/t/node-close-gets-stuck-in-the-stopping-state/3992/1 "2011-02-27T15:16:51Z")

</div>

We're using hadoop + the java bulk api to index our data using a one-  
hop strategy. This has worked out great so far, see  
[http://github.com/infochimps/wonderdog](http://github.com/infochimps/wonderdog). There's one major issue I'm  
trying to understand:

After using an elasticsearch 'node' object (embedded into each hadoop  
map task) the node needs to call the 'close' method. While this seems  
like an easy fix, here's the problem:

1. About 40% of the time this works great. The shutdown procedure is a  
finite state machine that looks like the following:

(stopping -\> stopped -\> closing -\> closed).

Only when the node object goes through that entire procedure (thereby  
severing all ties to elasticsearch) does the hadoop task commit and  
complete successfully.

1. The other 60% of the time the node gets stuck in the (stopping)  
state which ultimately results in the task (which completed  
successfully mind you) to timeout and fail.

Now, in the current hackety version, the 'node' object itself does not  
call close. Instead the node's client (the thing actually using the  
open connection afaik) is closed. However, this is essentially a  
meaningless operation since the 'node' maintains a persistent  
connection. What this results in are 'rogue' hadoop processes that  
trick elasticsearch into thinking there are many more 'nodes' than  
there actually are. When enough rogue processes accumulate this causes  
a 'too many open files' issue.

What is the node doing during the stopping phase and how can I tell  
what's causing it to hang?

--jacob  
@thedatachef

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 27, 2011, 7:32pm UTC](https://discuss.elastic.co/t/node-close-gets-stuck-in-the-stopping-state/3992/2 "2011-02-27T19:32:14Z")

</div>

Is there a chance you can gist a thread dump of the task when its stuck? It will help seeing where exactly its stuck. Which version are you using?  
On Sunday, February 27, 2011 at 5:16 PM, Jacob Perkins wrote:

> We're using hadoop + the java bulk api to index our data using a one-  
> hop strategy. This has worked out great so far, see  
> [GitHub - infochimps/wonderdog: Wonderdog is now at https://github.com/infochimps-labs/wonderdog) ElasticSearch and Hadoop and beautiful bouncy elephant love.](http://github.com/infochimps/wonderdog). There's one major issue I'm  
> trying to understand:
> 
> After using an elasticsearch 'node' object (embedded into each hadoop  
> map task) the node needs to call the 'close' method. While this seems  
> like an easy fix, here's the problem:
> 
> 1. About 40% of the time this works great. The shutdown procedure is a  
> finite state machine that looks like the following:
> 
> (stopping -\> stopped -\> closing -\> closed).
> 
> Only when the node object goes through that entire procedure (thereby  
> severing all ties to elasticsearch) does the hadoop task commit and  
> complete successfully.
> 
> 1. The other 60% of the time the node gets stuck in the (stopping)  
> state which ultimately results in the task (which completed  
> successfully mind you) to timeout and fail.
> 
> Now, in the current hackety version, the 'node' object itself does not  
> call close. Instead the node's client (the thing actually using the  
> open connection afaik) is closed. However, this is essentially a  
> meaningless operation since the 'node' maintains a persistent  
> connection. What this results in are 'rogue' hadoop processes that  
> trick elasticsearch into thinking there are many more 'nodes' than  
> there actually are. When enough rogue processes accumulate this causes  
> a 'too many open files' issue.
> 
> What is the node doing during the stopping phase and how can I tell  
> what's causing it to hang?
> 
> --jacob  
> @thedatachef

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:11am UTC](https://discuss.elastic.co/t/node-close-gets-stuck-in-the-stopping-state/3992/3 "2017-07-06T04:11:18Z")

</div>


