# Increasing CLOSE\_WAIT connections and HTTP current\_open metric

**URL:** <https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223>\
**Category:** Elasticsearch\
**Created:** [June 26, 2012, 3:26am UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223 "2012-06-26T03:26:42Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [June 26, 2012, 3:26am UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/1 "2012-06-26T03:26:42Z")

</div>

Hello,

I've been fighting with ES that stops working after it gets hit for several  
minutes by about 1200 QPS. This is happening with ES 0.19.3 and 0.19.4 on  
big boxes (24 cores, 96 GB RAM).

What seems to be happening is that after a while we start seeing more and  
more and more CLOSE\_WAIT connections between search clients (~500 frontend  
PHP apps) and ES, like this:

$ netstat -T | head  
Active Internet connections (w/o servers)  
Proto Recv-Q Send-Q Local Address Foreign Address  
State  
tcp 325 0 [11.11.11.11-static.reverse.softlayer.com](http://11.11.11.11-static.reverse.softlayer.com):wap-wsp  
[184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035) CLOSE\_WAIT  
...  
...

The number of CLOSE\_WAIT connections goes from being 0 for a while to going  
up into thousands. And then at some point ES stops responding on port 9200  
(but still responds on port 9300).

The number of these CLOSE\_WAIT connections seems to _roughly_ correspond to  
the "current\_open" HTTP metric:

$ curl --silent  
11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true'  
| egrep 'address|current|curr|server\_open'  
"transport\_address" : "inet[/ 11.11.11.11 :9300]",  
"curr\_estab" : 87,  
"server\_open" : 36,  
"current\_open" : 7,  
\<== healthy, just restarted  
"transport\_address" : "inet[/ 22.22.22.22 :9300]",  
"curr\_estab" : 1245,  
"server\_open" : 612,  
"current\_open" : 14,  
\<== healthy, just restarted  
"transport\_address" : "inet[/ 33.33.33.33:9300]",  
"curr\_estab" : 93,  
"server\_open" : 36,  
"current\_open" : 14,  
\<== healthy, just restarted  
"transport\_address" : "inet[/ 44.44.44.44 :9300]",  
"curr\_estab" : 171,  
"server\_open" : 36,  
"current\_open" : 15776,  
\<== baaad, not restarted

I tried using using threadpool (both fixed and blocking with both abort and  
client rejection policies) to try stopping this "current\_open" from  
growing, e.g.:

threadpool:  
search:  
type: fixed  
size: 120  
queue\_size: 100  
reject\_policy: abort

But that didn't help.

I should say that the search apps hitting this ES cluster are not using  
persistent/keep-alive connections. And while this is clearly not ideal and  
not efficient, I think it still shouldn't cause this "leak" that ends up  
accumulating connections in CLOSE\_WAIT state and eventually getting ES to  
stop being responsive.

Is there anything one can do on the ES side to more aggressively close  
connections?

## Thanks, Otis

Search Analytics - [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
Scalable Performance Monitoring - [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)

---

<div class="post-metadata">

**Author:** ![Paul\_Brown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul_brown/32/1449_2.png) [@Paul\_Brown](https://discuss.elastic.co/u/Paul_Brown)\
**Post date:** [June 26, 2012, 6:50am UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/2 "2012-06-26T06:50:19Z")

</div>

Hi, Otis --

The wikipedia article on TCP has a state chart that may be helpful:

[Transmission Control Protocol - Wikipedia](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)

CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I assume) hasn't done a full job of closing the connection, e.g., closing the TCP connection but not the underlying socket, so that's where I'd look. For example, the option CURLOPT\_FORBID\_REUSE might be useful.

-- Paul

On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:

> Hello,
> 
> I've been fighting with ES that stops working after it gets hit for several minutes by about 1200 QPS. This is happening with ES 0.19.3 and 0.19.4 on big boxes (24 cores, 96 GB RAM).
> 
> What seems to be happening is that after a while we start seeing more and more and more CLOSE\_WAIT connections between search clients (~500 frontend PHP apps) and ES, like this:
> 
> $ netstat -T | head  
> Active Internet connections (w/o servers)  
> Proto Recv-Q Send-Q Local Address Foreign Address State  
> tcp 325 0 [11.11.11.11-static.reverse.softlayer.com](http://11.11.11.11-static.reverse.softlayer.com):wap-wsp [184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035) CLOSE\_WAIT  
> ...  
> ...
> 
> The number of CLOSE\_WAIT connections goes from being 0 for a while to going up into thousands. And then at some point ES stops responding on port 9200 (but still responds on port 9300).
> 
> The number of these CLOSE\_WAIT connections seems to _roughly_ correspond to the "current\_open" HTTP metric:
> 
> $ curl --silent 11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true' | egrep 'address|current|curr|server\_open'  
> "transport\_address" : "inet[/ 11.11.11.11 :9300]",  
> "curr\_estab" : 87,  
> "server\_open" : 36,  
> "current\_open" : 7, \<== healthy, just restarted  
> "transport\_address" : "inet[/ 22.22.22.22 :9300]",  
> "curr\_estab" : 1245,  
> "server\_open" : 612,  
> "current\_open" : 14, \<== healthy, just restarted  
> "transport\_address" : "inet[/ 33.33.33.33:9300]",  
> "curr\_estab" : 93,  
> "server\_open" : 36,  
> "current\_open" : 14, \<== healthy, just restarted  
> "transport\_address" : "inet[/ 44.44.44.44 :9300]",  
> "curr\_estab" : 171,  
> "server\_open" : 36,  
> "current\_open" : 15776, \<== baaad, not restarted
> 
> I tried using using threadpool (both fixed and blocking with both abort and client rejection policies) to try stopping this "current\_open" from growing, e.g.:
> 
> threadpool:  
> search:  
> type: fixed  
> size: 120  
> queue\_size: 100  
> reject\_policy: abort
> 
> But that didn't help.
> 
> I should say that the search apps hitting this ES cluster are not using persistent/keep-alive connections. And while this is clearly not ideal and not efficient, I think it still shouldn't cause this "leak" that ends up accumulating connections in CLOSE\_WAIT state and eventually getting ES to stop being responsive.
> 
> Is there anything one can do on the ES side to more aggressively close connections?
> 
> ## Thanks, Otis
> 
> Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [June 27, 2012, 9:12pm UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/3 "2012-06-27T21:12:30Z")

</div>

Hi Paul,

On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:

> Hi, Otis --
> 
> The wikipedia article on TCP has a state chart that may be helpful:
> 
> [Transmission Control Protocol - Wikipedia](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)
> 
> CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I  
> assume) hasn't done a full job of closing the connection, e.g., closing the  
> TCP connection but not the underlying socket, so that's where I'd look.  
> For example, the option CURLOPT\_FORBID\_REUSE might be useful.

Is that really so?  
I looked at this diagram: [File:TCP CLOSE.svg - Wikipedia](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
If I read that correctly, it looks like CLOSE\_WAIT happens when the client  
(left side) which initiated the connection issues a FIN, which I understand  
as the client saying "I want to close this connection". After that FIN is  
received by the server/receiver, that server/receiver goes into the  
CLOSE\_WAIT state and at that time it is supposed to answer by sending the  
ACK and then (after some time?) by sending FIN going going into LAST\_ACK  
state and then, after client responds with ACK, into CLOSE state.

So if the server side is in CLOSE\_WAIT, doesn't that mean that the server  
received a FIN, but did not send ACK back to client?

## Thanks, Otis

Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

> On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:
> 
> > Hello,
> > 
> > I've been fighting with ES that stops working after it gets hit for  
> > several minutes by about 1200 QPS. This is happening with ES 0.19.3 and  
> > 0.19.4 on big boxes (24 cores, 96 GB RAM).
> > 
> > What seems to be happening is that after a while we start seeing more  
> > and more and more CLOSE\_WAIT connections between search clients (~500  
> > frontend PHP apps) and ES, like this:
> > 
> > $ netstat -T | head  
> > Active Internet connections (w/o servers)  
> > Proto Recv-Q Send-Q Local Address Foreign Address  
> > State  
> > tcp 325 0 [11.11.11.11-static.reverse.softlayer.com](http://11.11.11.11-static.reverse.softlayer.com):wap-wsp  
> > [184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035) CLOSE\_WAIT  
> > ...  
> > ...
> > 
> > The number of CLOSE\_WAIT connections goes from being 0 for a while to  
> > going up into thousands. And then at some point ES stops responding on  
> > port 9200 (but still responds on port 9300).
> > 
> > The number of these CLOSE\_WAIT connections seems to _roughly_ correspond  
> > to the "current\_open" HTTP metric:
> > 
> > $ curl --silent  
> > 11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true'  
> > | egrep 'address|current|curr|server\_open'  
> > "transport\_address" : "inet[/ 11.11.11.11 :9300]",  
> > "curr\_estab" : 87,  
> > "server\_open" : 36,  
> > "current\_open" : 7,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 22.22.22.22 :9300]",  
> > "curr\_estab" : 1245,  
> > "server\_open" : 612,  
> > "current\_open" : 14,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 33.33.33.33:9300]",  
> > "curr\_estab" : 93,  
> > "server\_open" : 36,  
> > "current\_open" : 14,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 44.44.44.44 :9300]",  
> > "curr\_estab" : 171,  
> > "server\_open" : 36,  
> > "current\_open" : 15776,  
> > \<== baaad, not restarted
> > 
> > I tried using using threadpool (both fixed and blocking with both abort  
> > and client rejection policies) to try stopping this "current\_open" from  
> > growing, e.g.:
> > 
> > threadpool:  
> > search:  
> > type: fixed  
> > size: 120  
> > queue\_size: 100  
> > reject\_policy: abort
> > 
> > But that didn't help.
> > 
> > I should say that the search apps hitting this ES cluster are not using  
> > persistent/keep-alive connections. And while this is clearly not ideal and  
> > not efficient, I think it still shouldn't cause this "leak" that ends up  
> > accumulating connections in CLOSE\_WAIT state and eventually getting ES to  
> > stop being responsive.
> > 
> > Is there anything one can do on the ES side to more aggressively close  
> > connections?
> > 
> > ## Thanks, Otis
> > 
> > Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> > Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [June 27, 2012, 10:00pm UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/4 "2012-06-27T22:00:33Z")

</div>

First, can you make sure you ran your test with reject policy of abort for  
the thread pool?  
Second, Can you try two things:

1. After you stop the load test, do you still have CLOSE\_WAIT?
2. If you run a single "client" load test, do you see CLOSE\_WAIT?

-shay.banon

On Wed, Jun 27, 2012 at 11:12 PM, Otis Gospodnetic \<  
[otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:

> Hi Paul,
> 
> On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:
> 
> > Hi, Otis --
> > 
> > The wikipedia article on TCP has a state chart that may be helpful:
> > 
> > [http://en.wikipedia.org/wiki/\*\*Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/**Transmission_Control_Protocol)[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)
> > 
> > CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I  
> > assume) hasn't done a full job of closing the connection, e.g., closing the  
> > TCP connection but not the underlying socket, so that's where I'd look.  
> > For example, the option CURLOPT\_FORBID\_REUSE might be useful.
> 
> Is that really so?  
> I looked at this diagram: [File:TCP CLOSE.svg - Wikipedia](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> If I read that correctly, it looks like CLOSE\_WAIT happens when the client  
> (left side) which initiated the connection issues a FIN, which I understand  
> as the client saying "I want to close this connection". After that FIN is  
> received by the server/receiver, that server/receiver goes into the  
> CLOSE\_WAIT state and at that time it is supposed to answer by sending the  
> ACK and then (after some time?) by sending FIN going going into LAST\_ACK  
> state and then, after client responds with ACK, into CLOSE state.
> 
> So if the server side is in CLOSE\_WAIT, doesn't that mean that the server  
> received a FIN, but did not send ACK back to client?
> 
> ## Thanks, Otis
> 
> Search Analytics - [http://sematext.com/search-\*\*analytics/index.html](http://sematext.com/search-**analytics/index.html)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> Scalable Performance Monitoring - [http://sematext.com/spm/index.\*\*html](http://sematext.com/spm/index.**html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> 
> > On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:
> > 
> > > Hello,
> > > 
> > > I've been fighting with ES that stops working after it gets hit for  
> > > several minutes by about 1200 QPS. This is happening with ES 0.19.3 and  
> > > 0.19.4 on big boxes (24 cores, 96 GB RAM).
> > > 
> > > What seems to be happening is that after a while we start seeing more  
> > > and more and more CLOSE\_WAIT connections between search clients (~500  
> > > frontend PHP apps) and ES, like this:
> > > 
> > > $ netstat -T | head  
> > > Active Internet connections (w/o servers)  
> > > Proto Recv-Q Send-Q Local Address Foreign Address  
> > > State  
> > > tcp 325 0 11.11.11.11-static.reverse.\*\*[softlayer.com](http://softlayer.com):wap-wsp  
> > > 184.184.184.184-static.\*\*[reverse.softlayer.com:32035](http://reverse.softlayer.com:32035)[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)CLOSE\_WAIT  
> > > ...  
> > > ...
> > > 
> > > The number of CLOSE\_WAIT connections goes from being 0 for a while to  
> > > going up into thousands. And then at some point ES stops responding on  
> > > port 9200 (but still responds on port 9300).
> > > 
> > > The number of these CLOSE\_WAIT connections seems to _roughly_  
> > > correspond to the "current\_open" HTTP metric:
> > > 
> > > $ curl --silent 11.11.11.11:9200/_cluster/ **nodes/stats?network=true&**  
> > > transport=true&http=true&\*\*thread\_pool=true&indices=\*\*false&pretty=true[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)'  
> > > | egrep 'address|current|curr|server_\*\*open'  
> > > "transport\_address" : "inet[/ 11.11.11.11 :9300]",  
> > > "curr\_estab" : 87,  
> > > "server\_open" : 36,  
> > > "current\_open" : 7,  
> > > \<== healthy, just restarted  
> > > "transport\_address" : "inet[/ 22.22.22.22 :9300]",  
> > > "curr\_estab" : 1245,  
> > > "server\_open" : 612,  
> > > "current\_open" : 14,  
> > > \<== healthy, just restarted  
> > > "transport\_address" : "inet[/ 33.33.33.33:9300]",  
> > > "curr\_estab" : 93,  
> > > "server\_open" : 36,  
> > > "current\_open" : 14,  
> > > \<== healthy, just restarted  
> > > "transport\_address" : "inet[/ 44.44.44.44 :9300]",  
> > > "curr\_estab" : 171,  
> > > "server\_open" : 36,  
> > > "current\_open" : 15776,  
> > > \<== baaad, not restarted
> > > 
> > > I tried using using threadpool (both fixed and blocking with both abort  
> > > and client rejection policies) to try stopping this "current\_open" from  
> > > growing, e.g.:
> > > 
> > > threadpool:  
> > > search:  
> > > type: fixed  
> > > size: 120  
> > > queue\_size: 100  
> > > reject\_policy: abort
> > > 
> > > But that didn't help.
> > > 
> > > I should say that the search apps hitting this ES cluster are not using  
> > > persistent/keep-alive connections. And while this is clearly not ideal and  
> > > not efficient, I think it still shouldn't cause this "leak" that ends up  
> > > accumulating connections in CLOSE\_WAIT state and eventually getting ES to  
> > > stop being responsive.
> > > 
> > > Is there anything one can do on the ES side to more aggressively close  
> > > connections?
> > > 
> > > ## Thanks, Otis
> > > 
> > > Search Analytics - [http://sematext.com/search-\*\*analytics/index.html](http://sematext.com/search-**analytics/index.html)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > > Scalable Performance Monitoring - [http://sematext.com/spm/index.\*\*html](http://sematext.com/spm/index.**html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)

---

<div class="post-metadata">

**Author:** ![Rafal\_Kuc\_3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rafal_kuc_3/32/799_2.png) [@Rafal\_Kuc\_3](https://discuss.elastic.co/u/Rafal_Kuc_3)\
**Post date:** [June 27, 2012, 10:11pm UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/5 "2012-06-27T22:11:38Z")

</div>

Hello!

We didn't see any CLOSE\_WAIT's while we were doing performance testing using keep alive. I'll change reject policy to abort and will see how that goes.

_--_

Regards,

Rafał Kuć

Sematext :: [http://sematext.com/](http://sematext.com/) _:: Solr - Lucene - Nutch - ElasticSearch_

First, can you make sure you ran your test with reject policy of abort for the thread pool?

Second, Can you try two things:

1. After you stop the load test, do you still have CLOSE\_WAIT?

2. If you run a single "client" load test, do you see CLOSE\_WAIT?

-shay.banon

On Wed, Jun 27, 2012 at 11:12 PM, Otis Gospodnetic \<[otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:

Hi Paul,

On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:

Hi, Otis --

The wikipedia article on TCP has a state chart that may be helpful:

[http://en.wikipedia.org/wiki/](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)[Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)

CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I assume) hasn't done a full job of closing the connection, e.g., closing the TCP connection but not the underlying socket, so that's where I'd look. For example, the option CURLOPT\_FORBID\_REUSE might be useful.

Is that really so?

I looked at this diagram: [http://en.wikipedia.org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)

If I read that correctly, it looks like CLOSE\_WAIT happens when the client (left side) which initiated the connection issues a FIN, which I understand as the client saying "I want to close this connection". After that FIN is received by the server/receiver, that server/receiver goes into the CLOSE\_WAIT state and at that time it is supposed to answer by sending the ACK and then (after some time?) by sending FIN going going into LAST\_ACK state and then, after client responds with ACK, into CLOSE state.

So if the server side is in CLOSE\_WAIT, doesn't that mean that the server received a FIN, but did not send ACK back to client?

Thanks,

Otis

--

Search Analytics - [http://sematext.com/search-](http://sematext.com/search-analytics/index.html)[analytics/index.html](http://sematext.com/search-analytics/index.html)

Scalable Performance Monitoring - [http://sematext.com/spm/index.](http://sematext.com/spm/index.html)[html](http://sematext.com/spm/index.html)

On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:

\> Hello,

\>

\> I've been fighting with ES that stops working after it gets hit for several minutes by about 1200 QPS. This is happening with ES 0.19.3 and 0.19.4 on big boxes (24 cores, 96 GB RAM).

\>

\> What seems to be happening is that after a while we start seeing more and more and more CLOSE\_WAIT connections between search clients (~500 frontend PHP apps) and ES, like this:

\>

\> $ netstat -T | head

\> Active Internet connections (w/o servers)

\> Proto Recv-Q Send-Q Local Address Foreign Address State

\> tcp 325 0 [11.11.11.11-static.reverse.softlayer.com](http://11.11.11.11-static.reverse.softlayer.com):wap-wsp [184.184.184.184-static.](http://184.184.184.184-static.reverse.softlayer.com:32035)[reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035) CLOSE\_WAIT

\> ...

\> ...

\>

\> The number of CLOSE\_WAIT connections goes from being 0 for a while to going up into thousands. And then at some point ES stops responding on port 9200 (but still responds on port 9300).

\>

\> The number of these CLOSE\_WAIT connections seems to _roughly_ correspond to the "current\_open" HTTP metric:

\>

\> $ curl --silent [11.11.11.11:9200/\_cluster/](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[nodes/stats?network=true&](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[transport=true&http=true&](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[thread\_pool=true&indices=](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)' | egrep 'address|current|curr|server\_open'

\> "transport\_address" : "inet[/ 11.11.11.11 :9300]",

\> "curr\_estab" : 87,

\> "server\_open" : 36,

\> "current\_open" : 7, \<== healthy, just restarted

\> "transport\_address" : "inet[/ 22.22.22.22 :9300]",

\> "curr\_estab" : 1245,

\> "server\_open" : 612,

\> "current\_open" : 14, \<== healthy, just restarted

\> "transport\_address" : "inet[/ [33.33.33.33:9300](http://33.33.33.33:9300)]",

\> "curr\_estab" : 93,

\> "server\_open" : 36,

\> "current\_open" : 14, \<== healthy, just restarted

\> "transport\_address" : "inet[/ 44.44.44.44 :9300]",

\> "curr\_estab" : 171,

\> "server\_open" : 36,

\> "current\_open" : 15776, \<== baaad, not restarted

\>

\> I tried using using threadpool (both fixed and blocking with both abort and client rejection policies) to try stopping this "current\_open" from growing, e.g.:

\>

\> threadpool:

\> search:

\> type: fixed

\> size: 120

\> queue\_size: 100

\> reject\_policy: abort

\>

\> But that didn't help.

\>

\> I should say that the search apps hitting this ES cluster are not using persistent/keep-alive connections. And while this is clearly not ideal and not efficient, I think it still shouldn't cause this "leak" that ends up accumulating connections in CLOSE\_WAIT state and eventually getting ES to stop being responsive.

\>

\> Is there anything one can do on the ES side to more aggressively close connections?

\>

\> Thanks,

\> Otis

\> --

\> Search Analytics - [http://sematext.com/search-](http://sematext.com/search-analytics/index.html)[analytics/index.html](http://sematext.com/search-analytics/index.html)

\> Scalable Performance Monitoring - [http://sematext.com/spm/index.](http://sematext.com/spm/index.html)[html](http://sematext.com/spm/index.html)

\>

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [June 27, 2012, 10:33pm UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/6 "2012-06-27T22:33:09Z")

</div>

- Based on otis first post, he was using reject\_policy of abort, can you  
clarify what reject\_policy was used with the test with no keep alive that  
resulted in many CLOSE\_WAIT?

- If the CLOSE\_WAIT is still a problem with reject policy of abort, can you  
run the test without keep alive and the two options I asked for?

On Thu, Jun 28, 2012 at 12:11 AM, Rafał Kuć [r.kuc@solr.pl](mailto:r.kuc@solr.pl) wrote:

> Hello!
> 
> We didn't see any CLOSE\_WAIT's while we were doing performance testing  
> using keep alive. I'll change reject policy to abort and will see how that  
> goes.
> 
> \*--  
> Regards,  
> Rafał Kuć  
> Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> Elasticsearch
> 
> - 
> 
> First, can you make sure you ran your test with reject policy of abort  
> for the thread pool?  
> Second, Can you try two things:
> 
> 1. After you stop the load test, do you still have CLOSE\_WAIT?
> 2. If you run a single "client" load test, do you see CLOSE\_WAIT?
> 
> -shay.banon
> 
> On Wed, Jun 27, 2012 at 11:12 PM, Otis Gospodnetic \<  
> [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:  
> Hi Paul,
> 
> On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:
> 
> Hi, Otis --
> 
> The wikipedia article on TCP has a state chart that may be helpful:
> 
> [Wikipedia, the free encyclopedia](http://en.wikipedia.org/wiki/)[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> Transmission\_Control\_Protocol[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)
> 
> CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I  
> assume) hasn't done a full job of closing the connection, e.g., closing the  
> TCP connection but not the underlying socket, so that's where I'd look.  
> For example, the option CURLOPT\_FORBID\_REUSE might be useful.
> 
> Is that really so?  
> I looked at this diagram: [File:TCP CLOSE.svg - Wikipedia](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> If I read that correctly, it looks like CLOSE\_WAIT happens when the client  
> (left side) which initiated the connection issues a FIN, which I understand  
> as the client saying "I want to close this connection". After that FIN is  
> received by the server/receiver, that server/receiver goes into the  
> CLOSE\_WAIT state and at that time it is supposed to answer by sending the  
> ACK and then (after some time?) by sending FIN going going into LAST\_ACK  
> state and then, after client responds with ACK, into CLOSE state.
> 
> So if the server side is in CLOSE\_WAIT, doesn't that mean that the server  
> received a FIN, but did not send ACK back to client?
> 
> ## Thanks, Otis
> 
> Search Analytics - [Sematext | IT System Monitoring Tools for DevOps](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> analytics/index.html [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> Scalable Performance Monitoring - [http://sematext.com/spm/index.](http://sematext.com/spm/index.)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> html [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> 
> On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:
> 
> > Hello,
> > 
> > I've been fighting with ES that stops working after it gets hit for  
> > several minutes by about 1200 QPS. This is happening with ES 0.19.3 and  
> > 0.19.4 on big boxes (24 cores, 96 GB RAM).
> > 
> > What seems to be happening is that after a while we start seeing more  
> > and more and more CLOSE\_WAIT connections between search clients (~500  
> > frontend PHP apps) and ES, like this:
> > 
> > $ netstat -T | head  
> > Active Internet connections (w/o servers)  
> > Proto Recv-Q Send-Q Local Address Foreign Address  
> > State  
> > tcp 325 0 [11.11.11.11-static.reverse.softlayer.com](http://11.11.11.11-static.reverse.softlayer.com):wap-wsp  
> > 184.184.184.184-static.[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > [reverse.softlayer.com:32035](http://reverse.softlayer.com:32035)[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > CLOSE\_WAIT  
> > ...  
> > ...
> > 
> > The number of CLOSE\_WAIT connections goes from being 0 for a while to  
> > going up into thousands. And then at some point ES stops responding on  
> > port 9200 (but still responds on port 9300).
> > 
> > The number of these CLOSE\_WAIT connections seems to _roughly_ correspond  
> > to the "current\_open" HTTP metric:
> > 
> > $ curl --silent 11.11.11.11:9200/\_cluster/[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > nodes/stats?network=true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > transport=true&http=true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > thread\_pool=true&indices=[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > false&pretty=true[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)'  
> > | egrep 'address|current|curr|server\_open'  
> > "transport\_address" : "inet[/ 11.11.11.11 :9300]",  
> > "curr\_estab" : 87,  
> > "server\_open" : 36,  
> > "current\_open" : 7,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 22.22.22.22 :9300]",  
> > "curr\_estab" : 1245,  
> > "server\_open" : 612,  
> > "current\_open" : 14,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 33.33.33.33:9300]",  
> > "curr\_estab" : 93,  
> > "server\_open" : 36,  
> > "current\_open" : 14,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 44.44.44.44 :9300]",  
> > "curr\_estab" : 171,  
> > "server\_open" : 36,  
> > "current\_open" : 15776,  
> > \<== baaad, not restarted
> > 
> > I tried using using threadpool (both fixed and blocking with both abort  
> > and client rejection policies) to try stopping this "current\_open" from  
> > growing, e.g.:
> > 
> > threadpool:  
> > search:  
> > type: fixed  
> > size: 120  
> > queue\_size: 100  
> > reject\_policy: abort
> > 
> > But that didn't help.
> > 
> > I should say that the search apps hitting this ES cluster are not using  
> > persistent/keep-alive connections. And while this is clearly not ideal and  
> > not efficient, I think it still shouldn't cause this "leak" that ends up  
> > accumulating connections in CLOSE\_WAIT state and eventually getting ES to  
> > stop being responsive.
> > 
> > Is there anything one can do on the ES side to more aggressively close  
> > connections?
> > 
> > ## Thanks, Otis
> > 
> > Search Analytics - [Sematext | IT System Monitoring Tools for DevOps](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > analytics/index.html [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > Scalable Performance Monitoring - [http://sematext.com/spm/index.](http://sematext.com/spm/index.)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> > html [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)

---

<div class="post-metadata">

**Author:** ![Rafal\_Kuc\_3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rafal_kuc_3/32/799_2.png) [@Rafal\_Kuc\_3](https://discuss.elastic.co/u/Rafal_Kuc_3)\
**Post date:** [June 27, 2012, 10:37pm UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/7 "2012-06-27T22:37:49Z")

</div>

Hello!

We tried different settings on nodes - one of it was reject policy abort with size 500 and 120, which resulted in CLOSE\_WAIT's under a certain load. Right now I've configured all the nodes in the cluster to 'reject\_policy: abort'.

As for the tests, I think it won't be a problem, but first, lets see how ElasticSearch will behave with the current reject policy.

Ah, just to clarify things, we did try 0.19.3, 0.19.4 and now we are running 0.19.7 🙂

_--_

Thanks,

Rafał Kuć

Sematext :: [http://sematext.com/](http://sematext.com/) _:: Solr - Lucene - Nutch - ElasticSearch_

- Based on otis first post, he was using reject\_policy of abort, can you clarify what reject\_policy was used with the test with no keep alive that resulted in many CLOSE\_WAIT?

- If the CLOSE\_WAIT is still a problem with reject policy of abort, can you run the test without keep alive and the two options I asked for?

On Thu, Jun 28, 2012 at 12:11 AM, Rafał Kuć \<[r.kuc@solr.pl](mailto:r.kuc@solr.pl)\> wrote:

Hello!

We didn't see any CLOSE\_WAIT's while we were doing performance testing using keep alive. I'll change reject policy to abort and will see how that goes.

_--_

Regards,

Rafał Kuć

Sematext :: [http://sematext.com/](http://sematext.com/) _:: Solr - Lucene - Nutch - ElasticSearch_

First, can you make sure you ran your test with reject policy of abort for the thread pool?

Second, Can you try two things:

1. After you stop the load test, do you still have CLOSE\_WAIT?

2. If you run a single "client" load test, do you see CLOSE\_WAIT?

-shay.banon

On Wed, Jun 27, 2012 at 11:12 PM, Otis Gospodnetic \<[otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:

Hi Paul,

On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:

Hi, Otis --

The wikipedia article on TCP has a state chart that may be helpful:

[http://en.wikipedia.org/wiki/](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)[Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)

CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I assume) hasn't done a full job of closing the connection, e.g., closing the TCP connection but not the underlying socket, so that's where I'd look. For example, the option CURLOPT\_FORBID\_REUSE might be useful.

Is that really so?

I looked at this diagram: [http://en.wikipedia.org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)

If I read that correctly, it looks like CLOSE\_WAIT happens when the client (left side) which initiated the connection issues a FIN, which I understand as the client saying "I want to close this connection". After that FIN is received by the server/receiver, that server/receiver goes into the CLOSE\_WAIT state and at that time it is supposed to answer by sending the ACK and then (after some time?) by sending FIN going going into LAST\_ACK state and then, after client responds with ACK, into CLOSE state.

So if the server side is in CLOSE\_WAIT, doesn't that mean that the server received a FIN, but did not send ACK back to client?

Thanks,

Otis

--

Search Analytics - [http://sematext.com/search-](http://sematext.com/search-analytics/index.html)[analytics/index.html](http://sematext.com/search-analytics/index.html)

Scalable Performance Monitoring - [http://sematext.com/spm/index.](http://sematext.com/spm/index.html)[html](http://sematext.com/spm/index.html)

On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:

\> Hello,

\>

\> I've been fighting with ES that stops working after it gets hit for several minutes by about 1200 QPS. This is happening with ES 0.19.3 and 0.19.4 on big boxes (24 cores, 96 GB RAM).

\>

\> What seems to be happening is that after a while we start seeing more and more and more CLOSE\_WAIT connections between search clients (~500 frontend PHP apps) and ES, like this:

\>

\> $ netstat -T | head

\> Active Internet connections (w/o servers)

\> Proto Recv-Q Send-Q Local Address Foreign Address State

\> tcp 325 0 [11.11.11.11-static.reverse.softlayer.com](http://11.11.11.11-static.reverse.softlayer.com):wap-wsp [184.184.184.184-static.](http://184.184.184.184-static.reverse.softlayer.com:32035)[reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035) CLOSE\_WAIT

\> ...

\> ...

\>

\> The number of CLOSE\_WAIT connections goes from being 0 for a while to going up into thousands. And then at some point ES stops responding on port 9200 (but still responds on port 9300).

\>

\> The number of these CLOSE\_WAIT connections seems to _roughly_ correspond to the "current\_open" HTTP metric:

\>

\> $ curl --silent [11.11.11.11:9200/\_cluster/](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[nodes/stats?network=true&](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[transport=true&http=true&](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[thread\_pool=true&indices=](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)' | egrep 'address|current|curr|server\_open'

\> "transport\_address" : "inet[/ 11.11.11.11 :9300]",

\> "curr\_estab" : 87,

\> "server\_open" : 36,

\> "current\_open" : 7, \<== healthy, just restarted

\> "transport\_address" : "inet[/ 22.22.22.22 :9300]",

\> "curr\_estab" : 1245,

\> "server\_open" : 612,

\> "current\_open" : 14, \<== healthy, just restarted

\> "transport\_address" : "inet[/ [33.33.33.33:9300](http://33.33.33.33:9300)]",

\> "curr\_estab" : 93,

\> "server\_open" : 36,

\> "current\_open" : 14, \<== healthy, just restarted

\> "transport\_address" : "inet[/ 44.44.44.44 :9300]",

\> "curr\_estab" : 171,

\> "server\_open" : 36,

\> "current\_open" : 15776, \<== baaad, not restarted

\>

\> I tried using using threadpool (both fixed and blocking with both abort and client rejection policies) to try stopping this "current\_open" from growing, e.g.:

\>

\> threadpool:

\> search:

\> type: fixed

\> size: 120

\> queue\_size: 100

\> reject\_policy: abort

\>

\> But that didn't help.

\>

\> I should say that the search apps hitting this ES cluster are not using persistent/keep-alive connections. And while this is clearly not ideal and not efficient, I think it still shouldn't cause this "leak" that ends up accumulating connections in CLOSE\_WAIT state and eventually getting ES to stop being responsive.

\>

\> Is there anything one can do on the ES side to more aggressively close connections?

\>

\> Thanks,

\> Otis

\> --

\> Search Analytics - [http://sematext.com/search-](http://sematext.com/search-analytics/index.html)[analytics/index.html](http://sematext.com/search-analytics/index.html)

\> Scalable Performance Monitoring - [http://sematext.com/spm/index.](http://sematext.com/spm/index.html)[html](http://sematext.com/spm/index.html)

\>

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [June 28, 2012, 1:35am UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/8 "2012-06-28T01:35:59Z")

</div>

Hi Shay,

This is what we are using now:

network:  
tcp:  
keep\_alive: true # also tried setting to false

# timeout: 5s # this didn't help

threadpool:  
search:  
type: fixed # also tried blocking  
size: 40  
queue\_size: 10  
reject\_policy: abort # also tried client

And ES sees this:

$ curl --silent  
'11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true'  
| egrep 'address|current|curr|server\_open'  
"transport\_address" : "inet[/11.11.11.11:9300]",  
"curr\_estab" : 764,  
"server\_open" : 360,  
"current\_open" : 1053, \<== this is approximately the number  
of CLOSE\_WAITs as shown by netstat -T

And this is where we see the threadpool numbers:

$ curl --silent  
'11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true'  
| egrep -C 4 search"

```
    "search" : {
      "threads" : 40,
      "queue" : 0,
      "active" : 6

```

The above is actually from a production environment we are trying to  
stabilize.

So because of only 40 threads, we do see rejections:

[2012-06-27 21:23:51,541][WARN][search.action] [search 2]  
Failed to send release search context  
org.elasticsearch.transport.RemoteTransportException: [search  
4][inet[/11.11.11.11:9300]][search/freeContext]  
Caused by:  
org.elasticsearch.common.util.concurrent.EsRejectedExecutionException  
at  
org.elasticsearch.common.util.concurrent.EsAbortPolicy.rejectedExecution(EsAbortPolicy.java:33)  
at  
java.util.concurrent.ThreadPoolExecutor.reject(ThreadPoolExecutor.java:767)

But CLOSE\_WAITs are still piling up.

The tricky thing is that under a certain load (\< 1200 QPS) CLOSE\_WAITs are  
nowhere to be found. The number of ESTABLISHED connections is more or less  
constant and the number of CLOSE\_WAITs is 0.  
But a few minutes after we increase the load to \> 1200 QPS we start seeing  
CLOSE\_WAITs and they just keep increasing.  
Actually, now that I look at things, I see that even the number of  
ESTABLISHED connections start to grow at some point, too. Not as fast as  
CLOSE\_WAITs, but growing.

## Otis

Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

On Wednesday, June 27, 2012 6:37:49 PM UTC-4, Rafał Kuć wrote:

> Hello!
> 
> We tried different settings on nodes - one of it was reject policy abort  
> with size 500 and 120, which resulted in CLOSE\_WAIT's under a certain load.  
> Right now I've configured all the nodes in the cluster to 'reject\_policy:  
> abort'.
> 
> As for the tests, I think it won't be a problem, but first, lets see how  
> Elasticsearch will behave with the current reject policy.
> 
> Ah, just to clarify things, we did try 0.19.3, 0.19.4 and now we are  
> running 0.19.7 🙂
> 
> \*--  
> Thanks,  
> Rafał Kuć  
> Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> Elasticsearch
> 
> - 
> 
> - Based on otis first post, he was using reject\_policy of abort, can you  
> clarify what reject\_policy was used with the test with no keep alive that  
> resulted in many CLOSE\_WAIT?
> 
> - If the CLOSE\_WAIT is still a problem with reject policy of abort, can  
> you run the test without keep alive and the two options I asked for?
> 
> On Thu, Jun 28, 2012 at 12:11 AM, Rafał Kuć [r.kuc@solr.pl](mailto:r.kuc@solr.pl) wrote:  
> Hello!
> 
> We didn't see any CLOSE\_WAIT's while we were doing performance testing  
> using keep alive. I'll change reject policy to abort and will see how that  
> goes.
> 
> \*--  
> Regards,  
> Rafał Kuć  
> Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> Elasticsearch
> 
> - 
> 
> First, can you make sure you ran your test with reject policy of abort  
> for the thread pool?  
> Second, Can you try two things:
> 
> 1. After you stop the load test, do you still have CLOSE\_WAIT?
> 2. If you run a single "client" load test, do you see CLOSE\_WAIT?
> 
> -shay.banon
> 
> On Wed, Jun 27, 2012 at 11:12 PM, Otis Gospodnetic \<  
> [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:  
> Hi Paul,
> 
> On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:
> 
> Hi, Otis --
> 
> The wikipedia article on TCP has a state chart that may be helpful:
> 
> [Wikipedia, the free encyclopedia](http://en.wikipedia.org/wiki/)[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> Transmission\_Control\_Protocol[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)
> 
> CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I  
> assume) hasn't done a full job of closing the connection, e.g., closing the  
> TCP connection but not the underlying socket, so that's where I'd look.  
> For example, the option CURLOPT\_FORBID\_REUSE might be useful.
> 
> Is that really so?  
> I looked at this diagram: [File:TCP CLOSE.svg - Wikipedia](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> If I read that correctly, it looks like CLOSE\_WAIT happens when the client  
> (left side) which initiated the connection issues a FIN, which I understand  
> as the client saying "I want to close this connection". After that FIN is  
> received by the server/receiver, that server/receiver goes into the  
> CLOSE\_WAIT state and at that time it is supposed to answer by sending the  
> ACK and then (after some time?) by sending FIN going going into LAST\_ACK  
> state and then, after client responds with ACK, into CLOSE state.
> 
> So if the server side is in CLOSE\_WAIT, doesn't that mean that the server  
> received a FIN, but did not send ACK back to client?
> 
> ## Thanks, Otis
> 
> Search Analytics - [http://sematext.com/search-](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> analytics/index.html [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> Scalable Performance Monitoring - [http://sematext.com/spm/index.](http://sematext.com/spm/index.)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> html [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> 
> On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:
> 
> > Hello,
> > 
> > I've been fighting with ES that stops working after it gets hit for  
> > several minutes by about 1200 QPS. This is happening with ES 0.19.3 and  
> > 0.19.4 on big boxes (24 cores, 96 GB RAM).
> > 
> > What seems to be happening is that after a while we start seeing more  
> > and more and more CLOSE\_WAIT connections between search clients (~500  
> > frontend PHP apps) and ES, like this:
> > 
> > $ netstat -T | head  
> > Active Internet connections (w/o servers)  
> > Proto Recv-Q Send-Q Local Address Foreign Address  
> > State  
> > tcp 325 0 [11.11.11.11-static.reverse.softlayer.com](http://11.11.11.11-static.reverse.softlayer.com):wap-wsp  
> > 184.184.184.184-static.[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > [reverse.softlayer.com:32035](http://reverse.softlayer.com:32035)[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > CLOSE\_WAIT  
> > ...  
> > ...
> > 
> > The number of CLOSE\_WAIT connections goes from being 0 for a while to  
> > going up into thousands. And then at some point ES stops responding on  
> > port 9200 (but still responds on port 9300).
> > 
> > The number of these CLOSE\_WAIT connections seems to _roughly_ correspond  
> > to the "current\_open" HTTP metric:
> > 
> > $ curl --silent 11.11.11.11:9200/\_cluster/[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > nodes/stats?network=true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > transport=true&http=true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > thread\_pool=true&indices=[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > false&pretty=true[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)'  
> > | egrep 'address|current|curr|server\_open'  
> > "transport\_address" : "inet[/ 11.11.11.11 :9300]",  
> > "curr\_estab" : 87,  
> > "server\_open" : 36,  
> > "current\_open" : 7,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 22.22.22.22 :9300]",  
> > "curr\_estab" : 1245,  
> > "server\_open" : 612,  
> > "current\_open" : 14,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 33.33.33.33:9300]",  
> > "curr\_estab" : 93,  
> > "server\_open" : 36,  
> > "current\_open" : 14,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 44.44.44.44 :9300]",  
> > "curr\_estab" : 171,  
> > "server\_open" : 36,  
> > "current\_open" : 15776,  
> > \<== baaad, not restarted
> > 
> > I tried using using threadpool (both fixed and blocking with both abort  
> > and client rejection policies) to try stopping this "current\_open" from  
> > growing, e.g.:
> > 
> > threadpool:  
> > search:  
> > type: fixed  
> > size: 120  
> > queue\_size: 100  
> > reject\_policy: abort
> > 
> > But that didn't help.
> > 
> > I should say that the search apps hitting this ES cluster are not using  
> > persistent/keep-alive connections. And while this is clearly not ideal and  
> > not efficient, I think it still shouldn't cause this "leak" that ends up  
> > accumulating connections in CLOSE\_WAIT state and eventually getting ES to  
> > stop being responsive.
> > 
> > Is there anything one can do on the ES side to more aggressively close  
> > connections?
> > 
> > ## Thanks, Otis
> > 
> > Search Analytics - [http://sematext.com/search-](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > analytics/index.html [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > Scalable Performance Monitoring - [http://sematext.com/spm/index.](http://sematext.com/spm/index.)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> > html [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)

---

<div class="post-metadata">

**Author:** ![Rafal\_Kuc\_3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rafal_kuc_3/32/799_2.png) [@Rafal\_Kuc\_3](https://discuss.elastic.co/u/Rafal_Kuc_3)\
**Post date:** [June 28, 2012, 7:38am UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/9 "2012-06-28T07:38:49Z")

</div>

Hello!

Right now, after a few hours the number of CLOSE\_WAIT connections are \> 20k on all nodes in the cluster, which makes ElasticSearch unresponsive to the API calls on 9200 port. However, clients using Java API are still working without a problem.

_--_

Regards,

Rafał Kuć

Sematext :: [http://sematext.com/](http://sematext.com/) _:: Solr - Lucene - Nutch - ElasticSearch_

Hi Shay,

This is what we are using now:

network:

tcp:

```
keep_alive: true # also tried setting to false

```

# timeout: 5s # this didn't help

threadpool:

search:

```
type: fixed # also tried blocking

size: 40

queue_size: 10

reject_policy: abort # also tried client

```

And ES sees this:

$ curl --silent '11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true' | egrep 'address|current|curr|server\_open'

```
  "transport_address" : "inet[/11.11.11.11:9300]",

      "curr_estab" : 764,

    "server_open" : 360,

    "current_open" : 1053, &lt;== this is approximately the number of CLOSE_WAITs as shown by netstat -T

```

And this is where we see the threadpool numbers:

$ curl --silent '11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true' | egrep -C 4 search"

```
    "search" : {

      "threads" : 40,

      "queue" : 0,

      "active" : 6

```

The above is actually from a production environment we are trying to stabilize.

So because of only 40 threads, we do see rejections:

[2012-06-27 21:23:51,541][WARN][search.action] [search 2] Failed to send release search context

org.elasticsearch.transport.RemoteTransportException: [search 4][inet[/11.11.11.11:9300]][search/freeContext]

Caused by: org.elasticsearch.common.util.concurrent.EsRejectedExecutionException

```
    at org.elasticsearch.common.util.concurrent.EsAbortPolicy.rejectedExecution(EsAbortPolicy.java:33)

    at java.util.concurrent.ThreadPoolExecutor.reject(ThreadPoolExecutor.java:767)

```

But CLOSE\_WAITs are still piling up.

The tricky thing is that under a certain load (\< 1200 QPS) CLOSE\_WAITs are nowhere to be found. The number of ESTABLISHED connections is more or less constant and the number of CLOSE\_WAITs is 0.

But a few minutes after we increase the load to \> 1200 QPS we start seeing CLOSE\_WAITs and they just keep increasing.

Actually, now that I look at things, I see that even the number of ESTABLISHED connections start to grow at some point, too. Not as fast as CLOSE\_WAITs, but growing.

Otis

--

Search Analytics - [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)

Scalable Performance Monitoring - [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)

On Wednesday, June 27, 2012 6:37:49 PM UTC-4, Rafał Kuć wrote:

Hello!

We tried different settings on nodes - one of it was reject policy abort with size 500 and 120, which resulted in CLOSE\_WAIT's under a certain load. Right now I've configured all the nodes in the cluster to 'reject\_policy: abort'.

As for the tests, I think it won't be a problem, but first, lets see how ElasticSearch will behave with the current reject policy.

Ah, just to clarify things, we did try 0.19.3, 0.19.4 and now we are running 0.19.7 🙂

_--_

Thanks,

Rafał Kuć

Sematext :: [http://sematext.com/](http://sematext.com/) _:: Solr - Lucene - Nutch - ElasticSearch_

- Based on otis first post, he was using reject\_policy of abort, can you clarify what reject\_policy was used with the test with no keep alive that resulted in many CLOSE\_WAIT?

- If the CLOSE\_WAIT is still a problem with reject policy of abort, can you run the test without keep alive and the two options I asked for?

On Thu, Jun 28, 2012 at 12:11 AM, Rafał Kuć \<[r.kuc@solr.pl](mailto:r.kuc@solr.pl)\> wrote:

Hello!

We didn't see any CLOSE\_WAIT's while we were doing performance testing using keep alive. I'll change reject policy to abort and will see how that goes.

_--_

Regards,

Rafał Kuć

Sematext :: [http://sematext.com/](http://sematext.com/) _:: Solr - Lucene - Nutch - ElasticSearch_

First, can you make sure you ran your test with reject policy of abort for the thread pool?

Second, Can you try two things:

1. After you stop the load test, do you still have CLOSE\_WAIT?

2. If you run a single "client" load test, do you see CLOSE\_WAIT?

-shay.banon

On Wed, Jun 27, 2012 at 11:12 PM, Otis Gospodnetic \<[otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:

Hi Paul,

On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:

Hi, Otis --

The wikipedia article on TCP has a state chart that may be helpful:

[http://en.wikipedia.org/](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)[wiki/](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)[Transmission\_Control\_](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)[Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)

CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I assume) hasn't done a full job of closing the connection, e.g., closing the TCP connection but not the underlying socket, so that's where I'd look. For example, the option CURLOPT\_FORBID\_REUSE might be useful.

Is that really so?

I looked at this diagram: [http://en.wikipedia.](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)[org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)

If I read that correctly, it looks like CLOSE\_WAIT happens when the client (left side) which initiated the connection issues a FIN, which I understand as the client saying "I want to close this connection". After that FIN is received by the server/receiver, that server/receiver goes into the CLOSE\_WAIT state and at that time it is supposed to answer by sending the ACK and then (after some time?) by sending FIN going going into LAST\_ACK state and then, after client responds with ACK, into CLOSE state.

So if the server side is in CLOSE\_WAIT, doesn't that mean that the server received a FIN, but did not send ACK back to client?

Thanks,

Otis

--

Search Analytics - [http://sematext.com/search-](http://sematext.com/search-analytics/index.html)[a](http://sematext.com/search-analytics/index.html)[nalytics/index.html](http://sematext.com/search-analytics/index.html)

Scalable Performance Monitoring - [http://sematext.com/spm/](http://sematext.com/spm/index.html)[index.](http://sematext.com/spm/index.html)[html](http://sematext.com/spm/index.html)

On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:

\> Hello,

\>

\> I've been fighting with ES that stops working after it gets hit for several minutes by about 1200 QPS. This is happening with ES 0.19.3 and 0.19.4 on big boxes (24 cores, 96 GB RAM).

\>

\> What seems to be happening is that after a while we start seeing more and more and more CLOSE\_WAIT connections between search clients (~500 frontend PHP apps) and ES, like this:

\>

\> $ netstat -T | head

\> Active Internet connections (w/o servers)

\> Proto Recv-Q Send-Q Local Address Foreign Address State

\> tcp 325 0 [11.11.11.11-static.reverse.softlayer.com](http://11.11.11.11-static.reverse.softlayer.com):wap-wsp [184.184.](http://184.184.184.184-static.reverse.softlayer.com:32035)[184.184-static.](http://184.184.184.184-static.reverse.softlayer.com:32035)[reverse.](http://184.184.184.184-static.reverse.softlayer.com:32035)[softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035) CLOSE\_WAIT

\> ...

\> ...

\>

\> The number of CLOSE\_WAIT connections goes from being 0 for a while to going up into thousands. And then at some point ES stops responding on port 9200 (but still responds on port 9300).

\>

\> The number of these CLOSE\_WAIT connections seems to _roughly_ correspond to the "current\_open" HTTP metric:

\>

\> $ curl --silent [11.11.11.11:9200/\_](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[cluster/](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[nodes/stats?network=](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[true&](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[transport=true&http=true&](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[thread\_pool=true&indices=](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[false](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)' | egrep 'address|current|curr|server\_open'

\> "transport\_address" : "inet[/ 11.11.11.11 :9300]",

\> "curr\_estab" : 87,

\> "server\_open" : 36,

\> "current\_open" : 7, \<== healthy, just restarted

\> "transport\_address" : "inet[/ 22.22.22.22 :9300]",

\> "curr\_estab" : 1245,

\> "server\_open" : 612,

\> "current\_open" : 14, \<== healthy, just restarted

\> "transport\_address" : "inet[/ [33.33.33.33:9300](http://33.33.33.33:9300)]",

\> "curr\_estab" : 93,

\> "server\_open" : 36,

\> "current\_open" : 14, \<== healthy, just restarted

\> "transport\_address" : "inet[/ 44.44.44.44 :9300]",

\> "curr\_estab" : 171,

\> "server\_open" : 36,

\> "current\_open" : 15776, \<== baaad, not restarted

\>

\> I tried using using threadpool (both fixed and blocking with both abort and client rejection policies) to try stopping this "current\_open" from growing, e.g.:

\>

\> threadpool:

\> search:

\> type: fixed

\> size: 120

\> queue\_size: 100

\> reject\_policy: abort

\>

\> But that didn't help.

\>

\> I should say that the search apps hitting this ES cluster are not using persistent/keep-alive connections. And while this is clearly not ideal and not efficient, I think it still shouldn't cause this "leak" that ends up accumulating connections in CLOSE\_WAIT state and eventually getting ES to stop being responsive.

\>

\> Is there anything one can do on the ES side to more aggressively close connections?

\>

\> Thanks,

\> Otis

\> --

\> Search Analytics - [http://sematext.com/search-](http://sematext.com/search-analytics/index.html)[a](http://sematext.com/search-analytics/index.html)[nalytics/index.html](http://sematext.com/search-analytics/index.html)

\> Scalable Performance Monitoring - [http://sematext.com/spm/](http://sematext.com/spm/index.html)[index.](http://sematext.com/spm/index.html)[html](http://sematext.com/spm/index.html)

\>

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [June 28, 2012, 6:42pm UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/10 "2012-06-28T18:42:36Z")

</div>

FYI, in a non-ES context, I had to deal with many low-level socket  
server/client error conditions, and encountered also a lot of CLOSE\_WAIT  
states that pile up and could make systems at worst case completely stall  
for minutes. As described, the client aborts the close prematurely and is  
not behaving correctly. As a consequence, the operating system has a  
system-wide tcp timeout before removing the socket. One (not so good)  
method is to tweak this operating system value (I was surprised how  
differently Linux and Solaris managed such cases).

The best method I found to work around this issue was both server and  
client using keepalive connections and handle the timeout states.

It would also be interesting to examine the curl / libcurl version,  
probably it is an old one and there is some issue there.

Best,

Jörg

>

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [June 28, 2012, 11:15pm UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/11 "2012-06-28T23:15:54Z")

</div>

Another question, are the PHP clients running on the same nodes as  
elasticsearch, or on different nodes (I assume on different nodes, just  
want to make sure).

I have a suspicion that maybe netty, the networking library we use, does  
not get around to actually close the connections because of the load on the  
system. Trying to chase that one down with Trustin. When this happens, can  
you gist a thread dump (jstack) just so we have it?

But, as suggested, the best thing to do is to use keep alive. nginx by the  
way can be used to abstract that nicely as Karmi found out, see here:

> <https://gist.github.com/karmi/0a2b0e0df83813a4045f/d237b1f2425353edc3b13c593ee2f960dfae0fca>

.

On Thu, Jun 28, 2012 at 9:38 AM, Rafał Kuć [r.kuc@solr.pl](mailto:r.kuc@solr.pl) wrote:

> Hello!
> 
> Right now, after a few hours the number of CLOSE\_WAIT connections are \>  
> 20k on all nodes in the cluster, which makes Elasticsearch unresponsive to  
> the API calls on 9200 port. However, clients using Java API are still  
> working without a problem.
> 
> _--  
> Regards,  
> Rafał Kuć  
> Sematext :: \*  
> [http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch - Elasticsearch
> 
> - 
> 
> Hi Shay,
> 
> This is what we are using now:
> 
> network:  
> tcp:  
> keep\_alive: true # also tried setting to false
> 
> # timeout: 5s # this didn't help
> 
> threadpool:  
> search:  
> type: fixed # also tried blocking  
> size: 40  
> queue\_size: 10  
> reject\_policy: abort # also tried client
> 
> And ES sees this:
> 
> $ curl --silent '  
> 11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true'  
> | egrep 'address|current|curr|server\_open'  
> "transport\_address" : "inet[/11.11.11.11:9300]",  
> "curr\_estab" : 764,  
> "server\_open" : 360,  
> "current\_open" : 1053, \<== this is approximately the  
> number of CLOSE\_WAITs as shown by netstat -T
> 
> And this is where we see the threadpool numbers:
> 
> $ curl --silent '  
> 11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true'  
> | egrep -C 4 search"
> 
> ```
> "search" : {
> "threads" : 40,
> "queue" : 0,
> "active" : 6
> 
> ```
> 
> The above is actually from a production environment we are trying to  
> stabilize.
> 
> So because of only 40 threads, we do see rejections:
> 
> [2012-06-27 21:23:51,541][WARN][search.action] [search 2]  
> Failed to send release search context  
> org.elasticsearch.transport.RemoteTransportException: [search  
> 4][inet[/11.11.11.11:9300]][search/freeContext]  
> Caused by:  
> org.elasticsearch.common.util.concurrent.EsRejectedExecutionException  
> at  
> org.elasticsearch.common.util.concurrent.EsAbortPolicy.rejectedExecution(EsAbortPolicy.java:33)  
> at  
> java.util.concurrent.ThreadPoolExecutor.reject(ThreadPoolExecutor.java:767)
> 
> But CLOSE\_WAITs are still piling up.
> 
> The tricky thing is that under a certain load (\< 1200 QPS) CLOSE\_WAITs are  
> nowhere to be found. The number of ESTABLISHED connections is more or less  
> constant and the number of CLOSE\_WAITs is 0.  
> But a few minutes after we increase the load to \> 1200 QPS we start seeing  
> CLOSE\_WAITs and they just keep increasing.  
> Actually, now that I look at things, I see that even the number of  
> ESTABLISHED connections start to grow at some point, too. Not as fast as  
> CLOSE\_WAITs, but growing.
> 
> ## Otis
> 
> Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Wednesday, June 27, 2012 6:37:49 PM UTC-4, Rafał Kuć wrote:  
> Hello!
> 
> We tried different settings on nodes - one of it was reject policy abort  
> with size 500 and 120, which resulted in CLOSE\_WAIT's under a certain load.  
> Right now I've configured all the nodes in the cluster to 'reject\_policy:  
> abort'.
> 
> As for the tests, I think it won't be a problem, but first, lets see how  
> Elasticsearch will behave with the current reject policy.
> 
> Ah, just to clarify things, we did try 0.19.3, 0.19.4 and now we are  
> running 0.19.7 🙂
> 
> \*--  
> Thanks,  
> Rafał Kuć  
> Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> Elasticsearch
> 
> - 
> 
> - Based on otis first post, he was using reject\_policy of abort, can you  
> clarify what reject\_policy was used with the test with no keep alive that  
> resulted in many CLOSE\_WAIT?
> 
> - If the CLOSE\_WAIT is still a problem with reject policy of abort, can  
> you run the test without keep alive and the two options I asked for?
> 
> On Thu, Jun 28, 2012 at 12:11 AM, Rafał Kuć [r.kuc@solr.pl](mailto:r.kuc@solr.pl) wrote:  
> Hello!
> 
> We didn't see any CLOSE\_WAIT's while we were doing performance testing  
> using keep alive. I'll change reject policy to abort and will see how that  
> goes.
> 
> \*--  
> Regards,  
> Rafał Kuć  
> Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> Elasticsearch
> 
> - 
> 
> First, can you make sure you ran your test with reject policy of abort  
> for the thread pool?  
> Second, Can you try two things:
> 
> 1. After you stop the load test, do you still have CLOSE\_WAIT?
> 2. If you run a single "client" load test, do you see CLOSE\_WAIT?
> 
> -shay.banon
> 
> On Wed, Jun 27, 2012 at 11:12 PM, Otis Gospodnetic \<  
> [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:  
> Hi Paul,
> 
> On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:
> 
> Hi, Otis --
> 
> The wikipedia article on TCP has a state chart that may be helpful:
> 
> [http://en.wikipedia.org/](http://en.wikipedia.org/)[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> wiki/ [http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> Transmission\_Control\_[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> Protocol [http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)
> 
> CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I  
> assume) hasn't done a full job of closing the connection, e.g., closing the  
> TCP connection but not the underlying socket, so that's where I'd look.  
> For example, the option CURLOPT\_FORBID\_REUSE might be useful.
> 
> Is that really so?  
> I looked at this diagram: [http://en.wikipedia](http://en.wikipedia).[http://en.wikipedia.org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> org/wiki/File:TCP\_CLOSE.svg[http://en.wikipedia.org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> If I read that correctly, it looks like CLOSE\_WAIT happens when the client  
> (left side) which initiated the connection issues a FIN, which I understand  
> as the client saying "I want to close this connection". After that FIN is  
> received by the server/receiver, that server/receiver goes into the  
> CLOSE\_WAIT state and at that time it is supposed to answer by sending the  
> ACK and then (after some time?) by sending FIN going going into LAST\_ACK  
> state and then, after client responds with ACK, into CLOSE state.
> 
> So if the server side is in CLOSE\_WAIT, doesn't that mean that the server  
> received a FIN, but did not send ACK back to client?
> 
> ## Thanks, Otis
> 
> Search Analytics - [Sematext | IT System Monitoring Tools for DevOps](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> a [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)nalytics/index.html[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> index. [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)html[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> 
> On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:
> 
> > Hello,
> > 
> > I've been fighting with ES that stops working after it gets hit for  
> > several minutes by about 1200 QPS. This is happening with ES 0.19.3 and  
> > 0.19.4 on big boxes (24 cores, 96 GB RAM).
> > 
> > What seems to be happening is that after a while we start seeing more  
> > and more and more CLOSE\_WAIT connections between search clients (~500  
> > frontend PHP apps) and ES, like this:
> > 
> > $ netstat -T | head  
> > Active Internet connections (w/o servers)  
> > Proto Recv-Q Send-Q Local Address Foreign Address  
> > State  
> > tcp 325 0 [11.11.11.11-static.reverse.softlayer.com](http://11.11.11.11-static.reverse.softlayer.com):wap-wsp  
> > 184.184. [http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > 184.184-static.[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > reverse. [http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > [softlayer.com:32035](http://softlayer.com:32035)[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > CLOSE\_WAIT  
> > ...  
> > ...
> > 
> > The number of CLOSE\_WAIT connections goes from being 0 for a while to  
> > going up into thousands. And then at some point ES stops responding on  
> > port 9200 (but still responds on port 9300).
> > 
> > The number of these CLOSE\_WAIT connections seems to _roughly_ correspond  
> > to the "current\_open" HTTP metric:
> > 
> > $ curl --silent 11.11.11.11:9200/\_[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > cluster/[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > nodes/stats?network=[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > transport=true&http=true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > thread\_pool=true&indices=[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > false[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > &pretty=true[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)'  
> > | egrep 'address|current|curr|server\_open'  
> > "transport\_address" : "inet[/ 11.11.11.11 :9300]",  
> > "curr\_estab" : 87,  
> > "server\_open" : 36,  
> > "current\_open" : 7,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 22.22.22.22 :9300]",  
> > "curr\_estab" : 1245,  
> > "server\_open" : 612,  
> > "current\_open" : 14,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 33.33.33.33:9300]",  
> > "curr\_estab" : 93,  
> > "server\_open" : 36,  
> > "current\_open" : 14,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 44.44.44.44 :9300]",  
> > "curr\_estab" : 171,  
> > "server\_open" : 36,  
> > "current\_open" : 15776,  
> > \<== baaad, not restarted
> > 
> > I tried using using threadpool (both fixed and blocking with both abort  
> > and client rejection policies) to try stopping this "current\_open" from  
> > growing, e.g.:
> > 
> > threadpool:  
> > search:  
> > type: fixed  
> > size: 120  
> > queue\_size: 100  
> > reject\_policy: abort
> > 
> > But that didn't help.
> > 
> > I should say that the search apps hitting this ES cluster are not using  
> > persistent/keep-alive connections. And while this is clearly not ideal and  
> > not efficient, I think it still shouldn't cause this "leak" that ends up  
> > accumulating connections in CLOSE\_WAIT state and eventually getting ES to  
> > stop being responsive.
> > 
> > Is there anything one can do on the ES side to more aggressively close  
> > connections?
> > 
> > ## Thanks, Otis
> > 
> > Search Analytics - [Sematext | IT System Monitoring Tools for DevOps](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > a [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)nalytics/index.html[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)
> 
> > Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> > index. [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)html[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> 
> >

---

<div class="post-metadata">

**Author:** ![Rafal\_Kuc\_3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rafal_kuc_3/32/799_2.png) [@Rafal\_Kuc\_3](https://discuss.elastic.co/u/Rafal_Kuc_3)\
**Post date:** [June 29, 2012, 7:26am UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/12 "2012-06-29T07:26:39Z")

</div>

Hello!

PHP application is running on different machines. We will gist the jstack as soon as CLOSE\_WAIT situation will happen again.

_--_

Regards,

Rafał Kuć

Sematext :: [http://sematext.com/](http://sematext.com/) _:: Solr - Lucene - Nutch - ElasticSearch_

Another question, are the PHP clients running on the same nodes as elasticsearch, or on different nodes (I assume on different nodes, just want to make sure).

I have a suspicion that maybe netty, the networking library we use, does not get around to actually close the connections because of the load on the system. Trying to chase that one down with Trustin. When this happens, can you gist a thread dump (jstack) just so we have it?

But, as suggested, the best thing to do is to use keep alive. nginx by the way can be used to abstract that nicely as Karmi found out, see here: [https://gist.github.com/0a2b0e0df83813a4045f/d237b1f2425353edc3b13c593ee2f960dfae0fca](https://gist.github.com/0a2b0e0df83813a4045f/d237b1f2425353edc3b13c593ee2f960dfae0fca).

On Thu, Jun 28, 2012 at 9:38 AM, Rafał Kuć \<[r.kuc@solr.pl](mailto:r.kuc@solr.pl)\> wrote:

Hello!

Right now, after a few hours the number of CLOSE\_WAIT connections are \> 20k on all nodes in the cluster, which makes ElasticSearch unresponsive to the API calls on 9200 port. However, clients using Java API are still working without a problem.

_--_

Regards,

Rafał Kuć

Sematext ::

[http://sematext.com/](http://sematext.com/) _:: Solr - Lucene - Nutch - ElasticSearch_

Hi Shay,

This is what we are using now:

network:

tcp:

```
keep_alive: true # also tried setting to false

```

# timeout: 5s # this didn't help

threadpool:

search:

```
type: fixed # also tried blocking

size: 40

queue_size: 10

reject_policy: abort # also tried client

```

And ES sees this:

$ curl --silent '[11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)' | egrep 'address|current|curr|server\_open'

```
  "transport_address" : "inet[/<a style=" font-family:'courier new'; font-size: 9pt;" href="http://11.11.11.11:9300">11.11.11.11:9300</a>]",

      "curr_estab" : 764,

    "server_open" : 360,

    "current_open" : 1053, &lt;== this is approximately the number of CLOSE_WAITs as shown by netstat -T

```

And this is where we see the threadpool numbers:

$ curl --silent '[11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)' | egrep -C 4 search"

```
    "search" : {

      "threads" : 40,

      "queue" : 0,

      "active" : 6

```

The above is actually from a production environment we are trying to stabilize.

So because of only 40 threads, we do see rejections:

[2012-06-27 21:23:51,541][WARN][search.action] [search 2] Failed to send release search context

org.elasticsearch.transport.RemoteTransportException: [search 4][inet[/11.11.11.11:9300]][search/freeContext]

Caused by: org.elasticsearch.common.util.concurrent.EsRejectedExecutionException

```
    at org.elasticsearch.common.util.concurrent.EsAbortPolicy.rejectedExecution(EsAbortPolicy.java:33)

    at java.util.concurrent.ThreadPoolExecutor.reject(ThreadPoolExecutor.java:767)

```

But CLOSE\_WAITs are still piling up.

The tricky thing is that under a certain load (\< 1200 QPS) CLOSE\_WAITs are nowhere to be found. The number of ESTABLISHED connections is more or less constant and the number of CLOSE\_WAITs is 0.

But a few minutes after we increase the load to \> 1200 QPS we start seeing CLOSE\_WAITs and they just keep increasing.

Actually, now that I look at things, I see that even the number of ESTABLISHED connections start to grow at some point, too. Not as fast as CLOSE\_WAITs, but growing.

Otis

--

Search Analytics - [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)

Scalable Performance Monitoring - [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)

On Wednesday, June 27, 2012 6:37:49 PM UTC-4, Rafał Kuć wrote:

Hello!

We tried different settings on nodes - one of it was reject policy abort with size 500 and 120, which resulted in CLOSE\_WAIT's under a certain load. Right now I've configured all the nodes in the cluster to 'reject\_policy: abort'.

As for the tests, I think it won't be a problem, but first, lets see how ElasticSearch will behave with the current reject policy.

Ah, just to clarify things, we did try 0.19.3, 0.19.4 and now we are running 0.19.7 🙂

_--_

Thanks,

Rafał Kuć

Sematext :: [http://sematext.com/](http://sematext.com/) _:: Solr - Lucene - Nutch - ElasticSearch_

- Based on otis first post, he was using reject\_policy of abort, can you clarify what reject\_policy was used with the test with no keep alive that resulted in many CLOSE\_WAIT?

- If the CLOSE\_WAIT is still a problem with reject policy of abort, can you run the test without keep alive and the two options I asked for?

On Thu, Jun 28, 2012 at 12:11 AM, Rafał Kuć \<[r.kuc@solr.pl](mailto:r.kuc@solr.pl)\> wrote:

Hello!

We didn't see any CLOSE\_WAIT's while we were doing performance testing using keep alive. I'll change reject policy to abort and will see how that goes.

_--_

Regards,

Rafał Kuć

Sematext :: [http://sematext.com/](http://sematext.com/) _:: Solr - Lucene - Nutch - ElasticSearch_

First, can you make sure you ran your test with reject policy of abort for the thread pool?

Second, Can you try two things:

1. After you stop the load test, do you still have CLOSE\_WAIT?

2. If you run a single "client" load test, do you see CLOSE\_WAIT?

-shay.banon

On Wed, Jun 27, 2012 at 11:12 PM, Otis Gospodnetic \<[otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:

Hi Paul,

On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:

Hi, Otis --

The wikipedia article on TCP has a state chart that may be helpful:

[http://en.wikipedia.org/](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)[wiki/](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)[Transmission\_Control\_](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)[Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)

CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I assume) hasn't done a full job of closing the connection, e.g., closing the TCP connection but not the underlying socket, so that's where I'd look. For example, the option CURLOPT\_FORBID\_REUSE might be useful.

Is that really so?

I looked at this diagram: [http://en.wikipedia.](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)[org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)

If I read that correctly, it looks like CLOSE\_WAIT happens when the client (left side) which initiated the connection issues a FIN, which I understand as the client saying "I want to close this connection". After that FIN is received by the server/receiver, that server/receiver goes into the CLOSE\_WAIT state and at that time it is supposed to answer by sending the ACK and then (after some time?) by sending FIN going going into LAST\_ACK state and then, after client responds with ACK, into CLOSE state.

So if the server side is in CLOSE\_WAIT, doesn't that mean that the server received a FIN, but did not send ACK back to client?

Thanks,

Otis

--

Search Analytics - [http://sematext.com/search-](http://sematext.com/search-analytics/index.html)[a](http://sematext.com/search-analytics/index.html)[nalytics/index.html](http://sematext.com/search-analytics/index.html)

Scalable Performance Monitoring - [http://sematext.com/spm/](http://sematext.com/spm/index.html)[index.](http://sematext.com/spm/index.html)[html](http://sematext.com/spm/index.html)

On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:

\> Hello,

\>

\> I've been fighting with ES that stops working after it gets hit for several minutes by about 1200 QPS. This is happening with ES 0.19.3 and 0.19.4 on big boxes (24 cores, 96 GB RAM).

\>

\> What seems to be happening is that after a while we start seeing more and more and more CLOSE\_WAIT connections between search clients (~500 frontend PHP apps) and ES, like this:

\>

\> $ netstat -T | head

\> Active Internet connections (w/o servers)

\> Proto Recv-Q Send-Q Local Address Foreign Address State

\> tcp 325 0 [11.11.11.11-static.reverse.softlayer.com](http://11.11.11.11-static.reverse.softlayer.com):wap-wsp [184.184.](http://184.184.184.184-static.reverse.softlayer.com:32035)[184.184-static.](http://184.184.184.184-static.reverse.softlayer.com:32035)[reverse.](http://184.184.184.184-static.reverse.softlayer.com:32035)[softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035) CLOSE\_WAIT

\> ...

\> ...

\>

\> The number of CLOSE\_WAIT connections goes from being 0 for a while to going up into thousands. And then at some point ES stops responding on port 9200 (but still responds on port 9300).

\>

\> The number of these CLOSE\_WAIT connections seems to _roughly_ correspond to the "current\_open" HTTP metric:

\>

\> $ curl --silent [11.11.11.11:9200/\_](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[cluster/](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[nodes/stats?network=](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[true&](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[transport=true&http=true&](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[thread\_pool=true&indices=](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[false](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)[&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)' | egrep 'address|current|curr|server\_open'

\> "transport\_address" : "inet[/ 11.11.11.11 :9300]",

\> "curr\_estab" : 87,

\> "server\_open" : 36,

\> "current\_open" : 7, \<== healthy, just restarted

\> "transport\_address" : "inet[/ 22.22.22.22 :9300]",

\> "curr\_estab" : 1245,

\> "server\_open" : 612,

\> "current\_open" : 14, \<== healthy, just restarted

\> "transport\_address" : "inet[/ [33.33.33.33:9300](http://33.33.33.33:9300)]",

\> "curr\_estab" : 93,

\> "server\_open" : 36,

\> "current\_open" : 14, \<== healthy, just restarted

\> "transport\_address" : "inet[/ 44.44.44.44 :9300]",

\> "curr\_estab" : 171,

\> "server\_open" : 36,

\> "current\_open" : 15776, \<== baaad, not restarted

\>

\> I tried using using threadpool (both fixed and blocking with both abort and client rejection policies) to try stopping this "current\_open" from growing, e.g.:

\>

\> threadpool:

\> search:

\> type: fixed

\> size: 120

\> queue\_size: 100

\> reject\_policy: abort

\>

\> But that didn't help.

\>

\> I should say that the search apps hitting this ES cluster are not using persistent/keep-alive connections. And while this is clearly not ideal and not efficient, I think it still shouldn't cause this "leak" that ends up accumulating connections in CLOSE\_WAIT state and eventually getting ES to stop being responsive.

\>

\> Is there anything one can do on the ES side to more aggressively close connections?

\>

\> Thanks,

\> Otis

\> --

\> Search Analytics - [http://sematext.com/search-](http://sematext.com/search-analytics/index.html)[a](http://sematext.com/search-analytics/index.html)[nalytics/index.html](http://sematext.com/search-analytics/index.html)

\> Scalable Performance Monitoring - [http://sematext.com/spm/](http://sematext.com/spm/index.html)[index.](http://sematext.com/spm/index.html)[html](http://sematext.com/spm/index.html)

\>

---

<div class="post-metadata">

**Author:** ![Rafal\_Kuc\_3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rafal_kuc_3/32/799_2.png) [@Rafal\_Kuc\_3](https://discuss.elastic.co/u/Rafal_Kuc_3)\
**Post date:** [June 29, 2012, 2:03pm UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/13 "2012-06-29T14:03:51Z")

</div>

Hello,

Shay, I've created a gist from one of the nodes, when CLOSE\_WAIT were  
happening again - [Dump when ElasticSearch CLOSE\_WAIT are increasing · GitHub](https://gist.github.com/3018085). Basically, the node just  
got unresponsive, not responding to searches and the CLOSE\_WAIT's are  
increasing. Ah one more thing - there was more than 250 CLOSE\_WAIT's while  
I was creating that dump.

Thanks,  
Rafał Kuć

W dniu piątek, 29 czerwca 2012 09:26:39 UTC+2 użytkownik Rafał Kuć napisał:

> Hello!
> 
> PHP application is running on different machines. We will gist the jstack  
> as soon as CLOSE\_WAIT situation will happen again.
> 
> \*--  
> Regards,  
> Rafał Kuć  
> Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> Elasticsearch
> 
> - 
> 
> Another question, are the PHP clients running on the same nodes as  
> elasticsearch, or on different nodes (I assume on different nodes, just  
> want to make sure).
> 
> I have a suspicion that maybe netty, the networking library we use, does  
> not get around to actually close the connections because of the load on the  
> system. Trying to chase that one down with Trustin. When this happens, can  
> you gist a thread dump (jstack) just so we have it?
> 
> But, as suggested, the best thing to do is to use keep alive. nginx by the  
> way can be used to abstract that nicely as Karmi found out, see here:  
> [Nginx as a keep-alive proxy for elasticsearch. · GitHub](https://gist.github.com/0a2b0e0df83813a4045f/d237b1f2425353edc3b13c593ee2f960dfae0fca)  
> .
> 
> On Thu, Jun 28, 2012 at 9:38 AM, Rafał Kuć [r.kuc@solr.pl](mailto:r.kuc@solr.pl) wrote:  
> Hello!
> 
> Right now, after a few hours the number of CLOSE\_WAIT connections are \>  
> 20k on all nodes in the cluster, which makes Elasticsearch unresponsive to  
> the API calls on 9200 port. However, clients using Java API are still  
> working without a problem.
> 
> \*--  
> Regards,  
> Rafał Kuć  
> Sematext ::  
> _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch - Elasticsearch
> 
> - 
> 
> Hi Shay,
> 
> This is what we are using now:
> 
> network:  
> tcp:  
> keep\_alive: true # also tried setting to false
> 
> # timeout: 5s # this didn't help
> 
> threadpool:  
> search:  
> type: fixed # also tried blocking  
> size: 40  
> queue\_size: 10  
> reject\_policy: abort # also tried client
> 
> And ES sees this:
> 
> $ curl --silent '  
> 11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true'  
> | egrep 'address|current|curr|server\_open'  
> "transport\_address" : "inet[/11.11.11.11:9300]",  
> "curr\_estab" : 764,  
> "server\_open" : 360,  
> "current\_open" : 1053, \<== this is approximately the  
> number of CLOSE\_WAITs as shown by netstat -T
> 
> And this is where we see the threadpool numbers:
> 
> $ curl --silent '  
> 11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true'  
> | egrep -C 4 search"
> 
> ```
> "search" : {
> "threads" : 40,
> "queue" : 0,
> "active" : 6
> 
> ```
> 
> The above is actually from a production environment we are trying to  
> stabilize.
> 
> So because of only 40 threads, we do see rejections:
> 
> [2012-06-27 21:23:51,541][WARN][search.action] [search 2]  
> Failed to send release search context  
> org.elasticsearch.transport.RemoteTransportException: [search  
> 4][inet[/11.11.11.11:9300]][search/freeContext]  
> Caused by:  
> org.elasticsearch.common.util.concurrent.EsRejectedExecutionException  
> at  
> org.elasticsearch.common.util.concurrent.EsAbortPolicy.rejectedExecution(EsAbortPolicy.java:33)  
> at  
> java.util.concurrent.ThreadPoolExecutor.reject(ThreadPoolExecutor.java:767)
> 
> But CLOSE\_WAITs are still piling up.
> 
> The tricky thing is that under a certain load (\< 1200 QPS) CLOSE\_WAITs are  
> nowhere to be found. The number of ESTABLISHED connections is more or less  
> constant and the number of CLOSE\_WAITs is 0.  
> But a few minutes after we increase the load to \> 1200 QPS we start seeing  
> CLOSE\_WAITs and they just keep increasing.  
> Actually, now that I look at things, I see that even the number of  
> ESTABLISHED connections start to grow at some point, too. Not as fast as  
> CLOSE\_WAITs, but growing.
> 
> ## Otis
> 
> Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Wednesday, June 27, 2012 6:37:49 PM UTC-4, Rafał Kuć wrote:  
> Hello!
> 
> We tried different settings on nodes - one of it was reject policy abort  
> with size 500 and 120, which resulted in CLOSE\_WAIT's under a certain load.  
> Right now I've configured all the nodes in the cluster to 'reject\_policy:  
> abort'.
> 
> As for the tests, I think it won't be a problem, but first, lets see how  
> Elasticsearch will behave with the current reject policy.
> 
> Ah, just to clarify things, we did try 0.19.3, 0.19.4 and now we are  
> running 0.19.7 🙂
> 
> \*--  
> Thanks,  
> Rafał Kuć  
> Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> Elasticsearch
> 
> - 
> 
> - Based on otis first post, he was using reject\_policy of abort, can you  
> clarify what reject\_policy was used with the test with no keep alive that  
> resulted in many CLOSE\_WAIT?
> 
> - If the CLOSE\_WAIT is still a problem with reject policy of abort, can  
> you run the test without keep alive and the two options I asked for?
> 
> On Thu, Jun 28, 2012 at 12:11 AM, Rafał Kuć [r.kuc@solr.pl](mailto:r.kuc@solr.pl) wrote:  
> Hello!
> 
> We didn't see any CLOSE\_WAIT's while we were doing performance testing  
> using keep alive. I'll change reject policy to abort and will see how that  
> goes.
> 
> \*--  
> Regards,  
> Rafał Kuć  
> Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> Elasticsearch
> 
> - 
> 
> First, can you make sure you ran your test with reject policy of abort  
> for the thread pool?  
> Second, Can you try two things:
> 
> 1. After you stop the load test, do you still have CLOSE\_WAIT?
> 2. If you run a single "client" load test, do you see CLOSE\_WAIT?
> 
> -shay.banon
> 
> On Wed, Jun 27, 2012 at 11:12 PM, Otis Gospodnetic \<  
> [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:  
> Hi Paul,
> 
> On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:
> 
> Hi, Otis --
> 
> The wikipedia article on TCP has a state chart that may be helpful:
> 
> [http://en.wikipedia.org/](http://en.wikipedia.org/)[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> wiki/ [http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> Transmission\_Control\_[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> Protocol [http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)
> 
> CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I  
> assume) hasn't done a full job of closing the connection, e.g., closing the  
> TCP connection but not the underlying socket, so that's where I'd look.  
> For example, the option CURLOPT\_FORBID\_REUSE might be useful.
> 
> Is that really so?  
> I looked at this diagram: [http://en.wikipedia](http://en.wikipedia).[http://en.wikipedia.org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> org/wiki/File:TCP\_CLOSE.svg[http://en.wikipedia.org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> If I read that correctly, it looks like CLOSE\_WAIT happens when the client  
> (left side) which initiated the connection issues a FIN, which I understand  
> as the client saying "I want to close this connection". After that FIN is  
> received by the server/receiver, that server/receiver goes into the  
> CLOSE\_WAIT state and at that time it is supposed to answer by sending the  
> ACK and then (after some time?) by sending FIN going going into LAST\_ACK  
> state and then, after client responds with ACK, into CLOSE state.
> 
> So if the server side is in CLOSE\_WAIT, doesn't that mean that the server  
> received a FIN, but did not send ACK back to client?
> 
> ## Thanks, Otis
> 
> Search Analytics - [Sematext | IT System Monitoring Tools for DevOps](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> a [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)nalytics/index.html[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> index. [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)html[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> 
> On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:
> 
> > Hello,
> > 
> > I've been fighting with ES that stops working after it gets hit for  
> > several minutes by about 1200 QPS. This is happening with ES 0.19.3 and  
> > 0.19.4 on big boxes (24 cores, 96 GB RAM).
> > 
> > What seems to be happening is that after a while we start seeing more  
> > and more and more CLOSE\_WAIT connections between search clients (~500  
> > frontend PHP apps) and ES, like this:
> > 
> > $ netstat -T | head  
> > Active Internet connections (w/o servers)  
> > Proto Recv-Q Send-Q Local Address Foreign Address  
> > State  
> > tcp 325 0 [11.11.11.11-static.reverse.softlayer.com](http://11.11.11.11-static.reverse.softlayer.com):wap-wsp  
> > 184.184. [http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > 184.184-static.[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > reverse. [http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > [softlayer.com:32035](http://softlayer.com:32035)[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > CLOSE\_WAIT  
> > ...  
> > ...
> > 
> > The number of CLOSE\_WAIT connections goes from being 0 for a while to  
> > going up into thousands. And then at some point ES stops responding on  
> > port 9200 (but still responds on port 9300).
> > 
> > The number of these CLOSE\_WAIT connections seems to _roughly_ correspond  
> > to the "current\_open" HTTP metric:
> > 
> > $ curl --silent 11.11.11.11:9200/\_[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > cluster/[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > nodes/stats?network=[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > transport=true&http=true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > thread\_pool=true&indices=[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > false[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > &pretty=true[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)'  
> > | egrep 'address|current|curr|server\_open'  
> > "transport\_address" : "inet[/ 11.11.11.11 :9300]",  
> > "curr\_estab" : 87,  
> > "server\_open" : 36,  
> > "current\_open" : 7,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 22.22.22.22 :9300]",  
> > "curr\_estab" : 1245,  
> > "server\_open" : 612,  
> > "current\_open" : 14,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 33.33.33.33:9300]",  
> > "curr\_estab" : 93,  
> > "server\_open" : 36,  
> > "current\_open" : 14,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 44.44.44.44 :9300]",  
> > "curr\_estab" : 171,  
> > "server\_open" : 36,  
> > "current\_open" : 15776,  
> > \<== baaad, not restarted
> > 
> > I tried using using threadpool (both fixed and blocking with both abort  
> > and client rejection policies) to try stopping this "current\_open" from  
> > growing, e.g.:
> > 
> > threadpool:  
> > search:  
> > type: fixed  
> > size: 120  
> > queue\_size: 100  
> > reject\_policy: abort
> > 
> > But that didn't help.
> > 
> > I should say that the search apps hitting this ES cluster are not using  
> > persistent/keep-alive connections. And while this is clearly not ideal and  
> > not efficient, I think it still shouldn't cause this "leak" that ends up  
> > accumulating connections in CLOSE\_WAIT state and eventually getting ES to  
> > stop being responsive.
> > 
> > Is there anything one can do on the ES side to more aggressively close  
> > connections?
> > 
> > ## Thanks, Otis
> > 
> > Search Analytics - [Sematext | IT System Monitoring Tools for DevOps](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > a [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)nalytics/index.html[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)
> 
> > Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> > index. [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)html[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> 
> >

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [June 29, 2012, 6:37pm UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/14 "2012-06-29T18:37:55Z")

</div>

Hi,

If it matters - this PHP application is really running on 500 front-end  
servers. Each of them is connecting to ES independently. And on top of  
that the PHP process is not persistent, so a new connection is created  
(with PHP curl lib) for each query sent to ES. (not our design! :))

## Otis

Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

On Friday, June 29, 2012 3:26:39 AM UTC-4, Rafał Kuć wrote:

> Hello!
> 
> PHP application is running on different machines. We will gist the jstack  
> as soon as CLOSE\_WAIT situation will happen again.
> 
> \*--  
> Regards,  
> Rafał Kuć  
> Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> Elasticsearch
> 
> - 
> 
> Another question, are the PHP clients running on the same nodes as  
> elasticsearch, or on different nodes (I assume on different nodes, just  
> want to make sure).
> 
> I have a suspicion that maybe netty, the networking library we use, does  
> not get around to actually close the connections because of the load on the  
> system. Trying to chase that one down with Trustin. When this happens, can  
> you gist a thread dump (jstack) just so we have it?
> 
> But, as suggested, the best thing to do is to use keep alive. nginx by the  
> way can be used to abstract that nicely as Karmi found out, see here:  
> [Nginx as a keep-alive proxy for elasticsearch. · GitHub](https://gist.github.com/0a2b0e0df83813a4045f/d237b1f2425353edc3b13c593ee2f960dfae0fca)  
> .
> 
> On Thu, Jun 28, 2012 at 9:38 AM, Rafał Kuć [r.kuc@solr.pl](mailto:r.kuc@solr.pl) wrote:  
> Hello!
> 
> Right now, after a few hours the number of CLOSE\_WAIT connections are \>  
> 20k on all nodes in the cluster, which makes Elasticsearch unresponsive to  
> the API calls on 9200 port. However, clients using Java API are still  
> working without a problem.
> 
> \*--  
> Regards,  
> Rafał Kuć  
> Sematext ::  
> _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch - Elasticsearch
> 
> - 
> 
> Hi Shay,
> 
> This is what we are using now:
> 
> network:  
> tcp:  
> keep\_alive: true # also tried setting to false
> 
> # timeout: 5s # this didn't help
> 
> threadpool:  
> search:  
> type: fixed # also tried blocking  
> size: 40  
> queue\_size: 10  
> reject\_policy: abort # also tried client
> 
> And ES sees this:
> 
> $ curl --silent '  
> 11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true'  
> | egrep 'address|current|curr|server\_open'  
> "transport\_address" : "inet[/11.11.11.11:9300]",  
> "curr\_estab" : 764,  
> "server\_open" : 360,  
> "current\_open" : 1053, \<== this is approximately the  
> number of CLOSE\_WAITs as shown by netstat -T
> 
> And this is where we see the threadpool numbers:
> 
> $ curl --silent '  
> 11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true'  
> | egrep -C 4 search"
> 
> ```
> "search" : {
> "threads" : 40,
> "queue" : 0,
> "active" : 6
> 
> ```
> 
> The above is actually from a production environment we are trying to  
> stabilize.
> 
> So because of only 40 threads, we do see rejections:
> 
> [2012-06-27 21:23:51,541][WARN][search.action] [search 2]  
> Failed to send release search context  
> org.elasticsearch.transport.RemoteTransportException: [search  
> 4][inet[/11.11.11.11:9300]][search/freeContext]  
> Caused by:  
> org.elasticsearch.common.util.concurrent.EsRejectedExecutionException  
> at  
> org.elasticsearch.common.util.concurrent.EsAbortPolicy.rejectedExecution(EsAbortPolicy.java:33)  
> at  
> java.util.concurrent.ThreadPoolExecutor.reject(ThreadPoolExecutor.java:767)
> 
> But CLOSE\_WAITs are still piling up.
> 
> The tricky thing is that under a certain load (\< 1200 QPS) CLOSE\_WAITs are  
> nowhere to be found. The number of ESTABLISHED connections is more or less  
> constant and the number of CLOSE\_WAITs is 0.  
> But a few minutes after we increase the load to \> 1200 QPS we start seeing  
> CLOSE\_WAITs and they just keep increasing.  
> Actually, now that I look at things, I see that even the number of  
> ESTABLISHED connections start to grow at some point, too. Not as fast as  
> CLOSE\_WAITs, but growing.
> 
> ## Otis
> 
> Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Wednesday, June 27, 2012 6:37:49 PM UTC-4, Rafał Kuć wrote:  
> Hello!
> 
> We tried different settings on nodes - one of it was reject policy abort  
> with size 500 and 120, which resulted in CLOSE\_WAIT's under a certain load.  
> Right now I've configured all the nodes in the cluster to 'reject\_policy:  
> abort'.
> 
> As for the tests, I think it won't be a problem, but first, lets see how  
> Elasticsearch will behave with the current reject policy.
> 
> Ah, just to clarify things, we did try 0.19.3, 0.19.4 and now we are  
> running 0.19.7 🙂
> 
> \*--  
> Thanks,  
> Rafał Kuć  
> Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> Elasticsearch
> 
> - 
> 
> - Based on otis first post, he was using reject\_policy of abort, can you  
> clarify what reject\_policy was used with the test with no keep alive that  
> resulted in many CLOSE\_WAIT?
> 
> - If the CLOSE\_WAIT is still a problem with reject policy of abort, can  
> you run the test without keep alive and the two options I asked for?
> 
> On Thu, Jun 28, 2012 at 12:11 AM, Rafał Kuć [r.kuc@solr.pl](mailto:r.kuc@solr.pl) wrote:  
> Hello!
> 
> We didn't see any CLOSE\_WAIT's while we were doing performance testing  
> using keep alive. I'll change reject policy to abort and will see how that  
> goes.
> 
> \*--  
> Regards,  
> Rafał Kuć  
> Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> Elasticsearch
> 
> - 
> 
> First, can you make sure you ran your test with reject policy of abort  
> for the thread pool?  
> Second, Can you try two things:
> 
> 1. After you stop the load test, do you still have CLOSE\_WAIT?
> 2. If you run a single "client" load test, do you see CLOSE\_WAIT?
> 
> -shay.banon
> 
> On Wed, Jun 27, 2012 at 11:12 PM, Otis Gospodnetic \<  
> [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:  
> Hi Paul,
> 
> On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:
> 
> Hi, Otis --
> 
> The wikipedia article on TCP has a state chart that may be helpful:
> 
> [http://en.wikipedia.org/](http://en.wikipedia.org/)[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> wiki/ [http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> Transmission\_Control\_[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> Protocol [http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)
> 
> CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I  
> assume) hasn't done a full job of closing the connection, e.g., closing the  
> TCP connection but not the underlying socket, so that's where I'd look.  
> For example, the option CURLOPT\_FORBID\_REUSE might be useful.
> 
> Is that really so?  
> I looked at this diagram: [http://en.wikipedia](http://en.wikipedia).[http://en.wikipedia.org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> org/wiki/File:TCP\_CLOSE.svg[http://en.wikipedia.org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> If I read that correctly, it looks like CLOSE\_WAIT happens when the client  
> (left side) which initiated the connection issues a FIN, which I understand  
> as the client saying "I want to close this connection". After that FIN is  
> received by the server/receiver, that server/receiver goes into the  
> CLOSE\_WAIT state and at that time it is supposed to answer by sending the  
> ACK and then (after some time?) by sending FIN going going into LAST\_ACK  
> state and then, after client responds with ACK, into CLOSE state.
> 
> So if the server side is in CLOSE\_WAIT, doesn't that mean that the server  
> received a FIN, but did not send ACK back to client?
> 
> ## Thanks, Otis
> 
> Search Analytics - [http://sematext.com/search-](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> a [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)nalytics/index.html[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> index. [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)html[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> 
> On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:
> 
> > Hello,
> > 
> > I've been fighting with ES that stops working after it gets hit for  
> > several minutes by about 1200 QPS. This is happening with ES 0.19.3 and  
> > 0.19.4 on big boxes (24 cores, 96 GB RAM).
> > 
> > What seems to be happening is that after a while we start seeing more  
> > and more and more CLOSE\_WAIT connections between search clients (~500  
> > frontend PHP apps) and ES, like this:
> > 
> > $ netstat -T | head  
> > Active Internet connections (w/o servers)  
> > Proto Recv-Q Send-Q Local Address Foreign Address  
> > State  
> > tcp 325 0 [11.11.11.11-static.reverse.softlayer.com](http://11.11.11.11-static.reverse.softlayer.com):wap-wsp  
> > 184.184. [http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > 184.184-static.[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > reverse. [http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > [softlayer.com:32035](http://softlayer.com:32035)[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > CLOSE\_WAIT  
> > ...  
> > ...
> > 
> > The number of CLOSE\_WAIT connections goes from being 0 for a while to  
> > going up into thousands. And then at some point ES stops responding on  
> > port 9200 (but still responds on port 9300).
> > 
> > The number of these CLOSE\_WAIT connections seems to _roughly_ correspond  
> > to the "current\_open" HTTP metric:
> > 
> > $ curl --silent 11.11.11.11:9200/\_[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > cluster/[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > nodes/stats?network=[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > transport=true&http=true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > thread\_pool=true&indices=[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > false[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > &pretty=true[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)'  
> > | egrep 'address|current|curr|server\_open'  
> > "transport\_address" : "inet[/ 11.11.11.11 :9300]",  
> > "curr\_estab" : 87,  
> > "server\_open" : 36,  
> > "current\_open" : 7,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 22.22.22.22 :9300]",  
> > "curr\_estab" : 1245,  
> > "server\_open" : 612,  
> > "current\_open" : 14,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 33.33.33.33:9300]",  
> > "curr\_estab" : 93,  
> > "server\_open" : 36,  
> > "current\_open" : 14,  
> > \<== healthy, just restarted  
> > "transport\_address" : "inet[/ 44.44.44.44 :9300]",  
> > "curr\_estab" : 171,  
> > "server\_open" : 36,  
> > "current\_open" : 15776,  
> > \<== baaad, not restarted
> > 
> > I tried using using threadpool (both fixed and blocking with both abort  
> > and client rejection policies) to try stopping this "current\_open" from  
> > growing, e.g.:
> > 
> > threadpool:  
> > search:  
> > type: fixed  
> > size: 120  
> > queue\_size: 100  
> > reject\_policy: abort
> > 
> > But that didn't help.
> > 
> > I should say that the search apps hitting this ES cluster are not using  
> > persistent/keep-alive connections. And while this is clearly not ideal and  
> > not efficient, I think it still shouldn't cause this "leak" that ends up  
> > accumulating connections in CLOSE\_WAIT state and eventually getting ES to  
> > stop being responsive.
> > 
> > Is there anything one can do on the ES side to more aggressively close  
> > connections?
> > 
> > ## Thanks, Otis
> > 
> > Search Analytics - [http://sematext.com/search-](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > a [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)nalytics/index.html[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)
> 
> > Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> > index. [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)html[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> 
> >

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [June 30, 2012, 4:35pm UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/15 "2012-06-30T16:35:00Z")

</div>

Ok, btw, did you manage to do the tests I asked for before?  
Also, can you check what the raw HTTP request the PHP clients are sending.  
A simple dump of a single request (especially the HTTP version passed and  
the headers) would be enough.

On Fri, Jun 29, 2012 at 8:37 PM, Otis Gospodnetic \<  
[otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:

> Hi,
> 
> If it matters - this PHP application is really running on 500 front-end  
> servers. Each of them is connecting to ES independently. And on top of  
> that the PHP process is not persistent, so a new connection is created  
> (with PHP curl lib) for each query sent to ES. (not our design! :))
> 
> ## Otis
> 
> Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Friday, June 29, 2012 3:26:39 AM UTC-4, Rafał Kuć wrote:
> 
> > Hello!
> > 
> > PHP application is running on different machines. We will gist the jstack  
> > as soon as CLOSE\_WAIT situation will happen again.
> > 
> > \*--  
> > Regards,  
> > Rafał Kuć  
> > Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> > Elasticsearch
> > 
> > - 
> > 
> > Another question, are the PHP clients running on the same nodes as  
> > elasticsearch, or on different nodes (I assume on different nodes, just  
> > want to make sure).
> > 
> > I have a suspicion that maybe netty, the networking library we use, does  
> > not get around to actually close the connections because of the load on the  
> > system. Trying to chase that one down with Trustin. When this happens, can  
> > you gist a thread dump (jstack) just so we have it?
> > 
> > But, as suggested, the best thing to do is to use keep alive. nginx by  
> > the way can be used to abstract that nicely as Karmi found out, see here:  
> > [https://gist.github.com/\*\*0a2b0e0df83813a4045f/](https://gist.github.com/**0a2b0e0df83813a4045f/)\*\*  
> > d237b1f2425353edc3b13c593ee2f9\*\*60dfae0fca[https://gist.github.com/0a2b0e0df83813a4045f/d237b1f2425353edc3b13c593ee2f960dfae0fca](https://gist.github.com/0a2b0e0df83813a4045f/d237b1f2425353edc3b13c593ee2f960dfae0fca)  
> > .
> > 
> > On Thu, Jun 28, 2012 at 9:38 AM, Rafał Kuć [r.kuc@solr.pl](mailto:r.kuc@solr.pl) wrote:  
> > Hello!
> > 
> > Right now, after a few hours the number of CLOSE\_WAIT connections are \>  
> > 20k on all nodes in the cluster, which makes Elasticsearch unresponsive to  
> > the API calls on 9200 port. However, clients using Java API are still  
> > working without a problem.
> > 
> > \*--  
> > Regards,  
> > Rafał Kuć  
> > Sematext ::  
> > _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch - Elasticsearch
> > 
> > - 
> > 
> > Hi Shay,
> > 
> > This is what we are using now:
> > 
> > network:  
> > tcp:  
> > keep\_alive: true # also tried setting to false
> > 
> > # timeout: 5s # this didn't help
> > 
> > threadpool:  
> > search:  
> > type: fixed # also tried blocking  
> > size: 40  
> > queue\_size: 10  
> > reject\_policy: abort # also tried client
> > 
> > And ES sees this:
> > 
> > $ curl --silent '11.11.11.11:9200/_cluster/ **nodes/stats?network=true&**  
> > transport=true&http=true&\*\*thread\_pool=true&indices=\*\*false&pretty=true[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)'  
> > | egrep 'address|current|curr|server_\*\*open'  
> > "transport\_address" : "inet[/11.11.11.11:9300]",  
> > "curr\_estab" : 764,  
> > "server\_open" : 360,  
> > "current\_open" : 1053, \<== this is approximately the  
> > number of CLOSE\_WAITs as shown by netstat -T
> > 
> > And this is where we see the threadpool numbers:
> > 
> > $ curl --silent '11.11.11.11:9200/\_cluster/ **nodes/stats?network=true&**  
> > transport=true&http=true&\*\*thread\_pool=true&indices=\*\*false&pretty=true[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)'  
> > | egrep -C 4 search"
> > 
> > ```
> > "search" : {
> > "threads" : 40,
> > "queue" : 0,
> > "active" : 6
> > 
> > ```
> > 
> > The above is actually from a production environment we are trying to  
> > stabilize.
> > 
> > So because of only 40 threads, we do see rejections:
> > 
> > [2012-06-27 21:23:51,541][WARN][search.action] [search 2]  
> > Failed to send release search context  
> > org.elasticsearch.transport.\*\*RemoteTransportException: [search  
> > 4][inet[/11.11.11.11:9300]][\*\*search/freeContext]  
> > Caused by: org.elasticsearch.common.util. **concurrent.**  
> > EsRejectedExecutionException  
> > at org.elasticsearch.common.util. **concurrent.EsAbortPolicy.**  
> > rejectedExecution(\*\*EsAbortPolicy.java:33)  
> > at java.util.concurrent.**ThreadPoolExecutor.reject(**  
> > ThreadPoolExecutor.java:767)
> > 
> > But CLOSE\_WAITs are still piling up.
> > 
> > The tricky thing is that under a certain load (\< 1200 QPS) CLOSE\_WAITs  
> > are nowhere to be found. The number of ESTABLISHED connections is more or  
> > less constant and the number of CLOSE\_WAITs is 0.  
> > But a few minutes after we increase the load to \> 1200 QPS we start  
> > seeing CLOSE\_WAITs and they just keep increasing.  
> > Actually, now that I look at things, I see that even the number of  
> > ESTABLISHED connections start to grow at some point, too. Not as fast as  
> > CLOSE\_WAITs, but growing.
> > 
> > ## Otis
> > 
> > Search Analytics - [http://sematext.com/search-\*\*analytics/index.html](http://sematext.com/search-**analytics/index.html)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/**index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > 
> > On Wednesday, June 27, 2012 6:37:49 PM UTC-4, Rafał Kuć wrote:  
> > Hello!
> > 
> > We tried different settings on nodes - one of it was reject policy abort  
> > with size 500 and 120, which resulted in CLOSE\_WAIT's under a certain load.  
> > Right now I've configured all the nodes in the cluster to 'reject\_policy:  
> > abort'.
> > 
> > As for the tests, I think it won't be a problem, but first, lets see how  
> > Elasticsearch will behave with the current reject policy.
> > 
> > Ah, just to clarify things, we did try 0.19.3, 0.19.4 and now we are  
> > running 0.19.7 🙂
> > 
> > \*--  
> > Thanks,  
> > Rafał Kuć  
> > Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> > Elasticsearch
> > 
> > - 
> > 
> > - Based on otis first post, he was using reject\_policy of abort, can  
> > you clarify what reject\_policy was used with the test with no keep alive  
> > that resulted in many CLOSE\_WAIT?
> > 
> > - If the CLOSE\_WAIT is still a problem with reject policy of abort, can  
> > you run the test without keep alive and the two options I asked for?
> > 
> > On Thu, Jun 28, 2012 at 12:11 AM, Rafał Kuć [r.kuc@solr.pl](mailto:r.kuc@solr.pl) wrote:  
> > Hello!
> > 
> > We didn't see any CLOSE\_WAIT's while we were doing performance testing  
> > using keep alive. I'll change reject policy to abort and will see how that  
> > goes.
> > 
> > \*--  
> > Regards,  
> > Rafał Kuć  
> > Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> > Elasticsearch
> > 
> > - 
> > 
> > First, can you make sure you ran your test with reject policy of abort  
> > for the thread pool?  
> > Second, Can you try two things:
> > 
> > 1. After you stop the load test, do you still have CLOSE\_WAIT?
> > 2. If you run a single "client" load test, do you see CLOSE\_WAIT?
> > 
> > -shay.banon
> > 
> > On Wed, Jun 27, 2012 at 11:12 PM, Otis Gospodnetic \<  
> > [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:  
> > Hi Paul,
> > 
> > On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:
> > 
> > Hi, Otis --
> > 
> > The wikipedia article on TCP has a state chart that may be helpful:
> > 
> > [http://en.wikipedia.org/](http://en.wikipedia.org/)[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> > wiki\*\*/ [http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> > Transmission\_Control\_[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> > Protocol [http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)\*\*
> > 
> > CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I  
> > assume) hasn't done a full job of closing the connection, e.g., closing the  
> > TCP connection but not the underlying socket, so that's where I'd look.  
> > For example, the option CURLOPT\_FORBID\_REUSE might be useful.
> > 
> > Is that really so?  
> > I looked at this diagram: [http://en.wikipedia](http://en.wikipedia).[http://en.wikipedia.org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> > o\*\*rg/wiki/File:TCP\_CLOSE.svg[http://en.wikipedia.org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> > If I read that correctly, it looks like CLOSE\_WAIT happens when the  
> > client (left side) which initiated the connection issues a FIN, which I  
> > understand as the client saying "I want to close this connection". After  
> > that FIN is received by the server/receiver, that server/receiver goes into  
> > the CLOSE\_WAIT state and at that time it is supposed to answer by sending  
> > the ACK and then (after some time?) by sending FIN going going into  
> > LAST\_ACK state and then, after client responds with ACK, into CLOSE state.
> > 
> > So if the server side is in CLOSE\_WAIT, doesn't that mean that the server  
> > received a FIN, but did not send ACK back to client?
> > 
> > ## Thanks, Otis
> > 
> > Search Analytics - [Sematext | IT System Monitoring Tools for DevOps](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > a [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)**nalytics/index.html[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> > inde**x. [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)html[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > 
> > On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:
> > 
> > > Hello,
> > > 
> > > I've been fighting with ES that stops working after it gets hit for  
> > > several minutes by about 1200 QPS. This is happening with ES 0.19.3 and  
> > > 0.19.4 on big boxes (24 cores, 96 GB RAM).
> > > 
> > > What seems to be happening is that after a while we start seeing more  
> > > and more and more CLOSE\_WAIT connections between search clients (~500  
> > > frontend PHP apps) and ES, like this:
> > > 
> > > $ netstat -T | head  
> > > Active Internet connections (w/o servers)  
> > > Proto Recv-Q Send-Q Local Address Foreign Address  
> > > State  
> > > tcp 325 0 11.11.11.11-static.reverse.**[softlayer.com](http://softlayer.com):wap-wsp  
> > > 184.184. [http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)**  
> > > 184.184-static.[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > > reverse. [http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > > softlay\*\*[er.com:32035](http://er.com:32035)[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > > CLOSE\_WAIT  
> > > ...  
> > > ...
> > > 
> > > The number of CLOSE\_WAIT connections goes from being 0 for a while to  
> > > going up into thousands. And then at some point ES stops responding on  
> > > port 9200 (but still responds on port 9300).
> > > 
> > > The number of these CLOSE\_WAIT connections seems to _roughly_  
> > > correspond to the "current\_open" HTTP metric:
> > > 
> > > $ curl --silent 11.11.11.11:9200/_[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > clu\*\*ster/[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > nodes/stats?network=[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > **transport=true&http=true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > threa**d\_pool=true&indices=[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > false[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > &\*\*pretty=true[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)'  
> > > | egrep 'address|current|curr|server_\*\*open'  
> > > "transport\_address" : "inet[/ 11.11.11.11 :9300]",  
> > > "curr\_estab" : 87,  
> > > "server\_open" : 36,  
> > > "current\_open" : 7,  
> > > \<== healthy, just restarted  
> > > "transport\_address" : "inet[/ 22.22.22.22 :9300]",  
> > > "curr\_estab" : 1245,  
> > > "server\_open" : 612,  
> > > "current\_open" : 14,  
> > > \<== healthy, just restarted  
> > > "transport\_address" : "inet[/ 33.33.33.33:9300]",  
> > > "curr\_estab" : 93,  
> > > "server\_open" : 36,  
> > > "current\_open" : 14,  
> > > \<== healthy, just restarted  
> > > "transport\_address" : "inet[/ 44.44.44.44 :9300]",  
> > > "curr\_estab" : 171,  
> > > "server\_open" : 36,  
> > > "current\_open" : 15776,  
> > > \<== baaad, not restarted
> > > 
> > > I tried using using threadpool (both fixed and blocking with both abort  
> > > and client rejection policies) to try stopping this "current\_open" from  
> > > growing, e.g.:
> > > 
> > > threadpool:  
> > > search:  
> > > type: fixed  
> > > size: 120  
> > > queue\_size: 100  
> > > reject\_policy: abort
> > > 
> > > But that didn't help.
> > > 
> > > I should say that the search apps hitting this ES cluster are not using  
> > > persistent/keep-alive connections. And while this is clearly not ideal and  
> > > not efficient, I think it still shouldn't cause this "leak" that ends up  
> > > accumulating connections in CLOSE\_WAIT state and eventually getting ES to  
> > > stop being responsive.
> > > 
> > > Is there anything one can do on the ES side to more aggressively close  
> > > connections?
> > > 
> > > ## Thanks, Otis
> > > 
> > > Search Analytics - [Sematext | IT System Monitoring Tools for DevOps](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > > a [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)\*\*nalytics/index.html[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)
> > 
> > > Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> > > inde\*\*x. [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)html[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > 
> > >

---

<div class="post-metadata">

**Author:** ![Rafal\_Kuc\_3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rafal_kuc_3/32/799_2.png) [@Rafal\_Kuc\_3](https://discuss.elastic.co/u/Rafal_Kuc_3)\
**Post date:** [July 2, 2012, 10:16am UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/16 "2012-07-02T10:16:43Z")

</div>

Hi,

I've managed to do that test today. I've run the following query:

{"size":10,"from":0,"sort":["\_score"],"fields":["\_source"],"query":{"query\_string":{"query":"high","default\_field":"tags"}},"facets":{"tags":{"terms":{"field":"tags","size":40}}}}

in a loop to one of the nodes using a simple simple bash loop. After  
running that query the number of CLOSE\_WAIT's increased from 0 about 150 -  
200, but slowly started to drop as the time was passing. Basically I  
couldn't force a single machine running queries to cause CLOSE\_WAIT's to be  
killing the node.

One more thing we didn't say before - we have Thrift plugin installed and  
working.

Thanks  
Rafał

One thing I noticed, is when the load get bigger the

W dniu sobota, 30 czerwca 2012 18:35:00 UTC+2 użytkownik kimchy napisał:

> Ok, btw, did you manage to do the tests I asked for before?  
> Also, can you check what the raw HTTP request the PHP clients are sending.  
> A simple dump of a single request (especially the HTTP version passed and  
> the headers) would be enough.
> 
> On Fri, Jun 29, 2012 at 8:37 PM, Otis Gospodnetic \<  
> [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:
> 
> > Hi,
> > 
> > If it matters - this PHP application is really running on 500 front-end  
> > servers. Each of them is connecting to ES independently. And on top of  
> > that the PHP process is not persistent, so a new connection is created  
> > (with PHP curl lib) for each query sent to ES. (not our design! :))
> > 
> > ## Otis
> > 
> > Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> > Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > 
> > On Friday, June 29, 2012 3:26:39 AM UTC-4, Rafał Kuć wrote:
> > 
> > > Hello!
> > > 
> > > PHP application is running on different machines. We will gist the  
> > > jstack as soon as CLOSE\_WAIT situation will happen again.
> > > 
> > > \*--  
> > > Regards,  
> > > Rafał Kuć  
> > > Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> > > Elasticsearch
> > > 
> > > - 
> > > 
> > > Another question, are the PHP clients running on the same nodes as  
> > > elasticsearch, or on different nodes (I assume on different nodes, just  
> > > want to make sure).
> > > 
> > > I have a suspicion that maybe netty, the networking library we use, does  
> > > not get around to actually close the connections because of the load on the  
> > > system. Trying to chase that one down with Trustin. When this happens, can  
> > > you gist a thread dump (jstack) just so we have it?
> > > 
> > > But, as suggested, the best thing to do is to use keep alive. nginx by  
> > > the way can be used to abstract that nicely as Karmi found out, see here:  
> > > [https://gist.github.com/\*\*0a2b0e0df83813a4045f/](https://gist.github.com/**0a2b0e0df83813a4045f/)\*\*  
> > > d237b1f2425353edc3b13c593ee2f9\*\*60dfae0fca[https://gist.github.com/0a2b0e0df83813a4045f/d237b1f2425353edc3b13c593ee2f960dfae0fca](https://gist.github.com/0a2b0e0df83813a4045f/d237b1f2425353edc3b13c593ee2f960dfae0fca)  
> > > .
> > > 
> > > On Thu, Jun 28, 2012 at 9:38 AM, Rafał Kuć [r.kuc@solr.pl](mailto:r.kuc@solr.pl) wrote:  
> > > Hello!
> > > 
> > > Right now, after a few hours the number of CLOSE\_WAIT connections are \>  
> > > 20k on all nodes in the cluster, which makes Elasticsearch unresponsive to  
> > > the API calls on 9200 port. However, clients using Java API are still  
> > > working without a problem.
> > > 
> > > \*--  
> > > Regards,  
> > > Rafał Kuć  
> > > Sematext ::  
> > > _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch - Elasticsearch
> > > 
> > > - 
> > > 
> > > Hi Shay,
> > > 
> > > This is what we are using now:
> > > 
> > > network:  
> > > tcp:  
> > > keep\_alive: true # also tried setting to false
> > > 
> > > # timeout: 5s # this didn't help
> > > 
> > > threadpool:  
> > > search:  
> > > type: fixed # also tried blocking  
> > > size: 40  
> > > queue\_size: 10  
> > > reject\_policy: abort # also tried client
> > > 
> > > And ES sees this:
> > > 
> > > $ curl --silent '11.11.11.11:9200/_cluster/ **nodes/stats?network=true&**  
> > > transport=true&http=true&\*\*thread\_pool=true&indices=\*\*false&pretty=true[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)'  
> > > | egrep 'address|current|curr|server_\*\*open'  
> > > "transport\_address" : "inet[/11.11.11.11:9300]",  
> > > "curr\_estab" : 764,  
> > > "server\_open" : 360,  
> > > "current\_open" : 1053, \<== this is approximately the  
> > > number of CLOSE\_WAITs as shown by netstat -T
> > > 
> > > And this is where we see the threadpool numbers:
> > > 
> > > $ curl --silent '11.11.11.11:9200/\_cluster/ **nodes/stats?network=true&**  
> > > transport=true&http=true&\*\*thread\_pool=true&indices=\*\*false&pretty=true[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)'  
> > > | egrep -C 4 search"
> > > 
> > > ```
> > > "search" : {
> > > "threads" : 40,
> > > "queue" : 0,
> > > "active" : 6
> > > 
> > > ```
> > > 
> > > The above is actually from a production environment we are trying to  
> > > stabilize.
> > > 
> > > So because of only 40 threads, we do see rejections:
> > > 
> > > [2012-06-27 21:23:51,541][WARN][search.action] [search 2]  
> > > Failed to send release search context  
> > > org.elasticsearch.transport.\*\*RemoteTransportException: [search  
> > > 4][inet[/11.11.11.11:9300]][\*\*search/freeContext]  
> > > Caused by: org.elasticsearch.common.util. **concurrent.**  
> > > EsRejectedExecutionException  
> > > at org.elasticsearch.common.util. **concurrent.EsAbortPolicy.**  
> > > rejectedExecution(\*\*EsAbortPolicy.java:33)  
> > > at java.util.concurrent.**ThreadPoolExecutor.reject(**  
> > > ThreadPoolExecutor.java:767)
> > > 
> > > But CLOSE\_WAITs are still piling up.
> > > 
> > > The tricky thing is that under a certain load (\< 1200 QPS) CLOSE\_WAITs  
> > > are nowhere to be found. The number of ESTABLISHED connections is more or  
> > > less constant and the number of CLOSE\_WAITs is 0.  
> > > But a few minutes after we increase the load to \> 1200 QPS we start  
> > > seeing CLOSE\_WAITs and they just keep increasing.  
> > > Actually, now that I look at things, I see that even the number of  
> > > ESTABLISHED connections start to grow at some point, too. Not as fast as  
> > > CLOSE\_WAITs, but growing.
> > > 
> > > ## Otis
> > > 
> > > Search Analytics - [http://sematext.com/search-\*\*analytics/index.html](http://sematext.com/search-**analytics/index.html)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > > Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/**index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > 
> > > On Wednesday, June 27, 2012 6:37:49 PM UTC-4, Rafał Kuć wrote:  
> > > Hello!
> > > 
> > > We tried different settings on nodes - one of it was reject policy abort  
> > > with size 500 and 120, which resulted in CLOSE\_WAIT's under a certain load.  
> > > Right now I've configured all the nodes in the cluster to 'reject\_policy:  
> > > abort'.
> > > 
> > > As for the tests, I think it won't be a problem, but first, lets see how  
> > > Elasticsearch will behave with the current reject policy.
> > > 
> > > Ah, just to clarify things, we did try 0.19.3, 0.19.4 and now we are  
> > > running 0.19.7 🙂
> > > 
> > > \*--  
> > > Thanks,  
> > > Rafał Kuć  
> > > Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> > > Elasticsearch
> > > 
> > > - 
> > > 
> > > - Based on otis first post, he was using reject\_policy of abort, can  
> > > you clarify what reject\_policy was used with the test with no keep alive  
> > > that resulted in many CLOSE\_WAIT?
> > > 
> > > - If the CLOSE\_WAIT is still a problem with reject policy of abort, can  
> > > you run the test without keep alive and the two options I asked for?
> > > 
> > > On Thu, Jun 28, 2012 at 12:11 AM, Rafał Kuć [r.kuc@solr.pl](mailto:r.kuc@solr.pl) wrote:  
> > > Hello!
> > > 
> > > We didn't see any CLOSE\_WAIT's while we were doing performance testing  
> > > using keep alive. I'll change reject policy to abort and will see how that  
> > > goes.
> > > 
> > > \*--  
> > > Regards,  
> > > Rafał Kuć  
> > > Sematext :: _[http://sematext.com/](http://sematext.com/)_ :: Solr - Lucene - Nutch -  
> > > Elasticsearch
> > > 
> > > - 
> > > 
> > > First, can you make sure you ran your test with reject policy of abort  
> > > for the thread pool?  
> > > Second, Can you try two things:
> > > 
> > > 1. After you stop the load test, do you still have CLOSE\_WAIT?
> > > 2. If you run a single "client" load test, do you see CLOSE\_WAIT?
> > > 
> > > -shay.banon
> > > 
> > > On Wed, Jun 27, 2012 at 11:12 PM, Otis Gospodnetic \<  
> > > [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:  
> > > Hi Paul,
> > > 
> > > On Tuesday, June 26, 2012 2:50:19 AM UTC-4, Paul Brown wrote:
> > > 
> > > Hi, Otis --
> > > 
> > > The wikipedia article on TCP has a state chart that may be helpful:
> > > 
> > > [http://en.wikipedia.org/](http://en.wikipedia.org/)[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> > > wiki\*\*/ [http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> > > Transmission\_Control\_[http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)  
> > > Protocol [http://en.wikipedia.org/wiki/Transmission\_Control\_Protocol](http://en.wikipedia.org/wiki/Transmission_Control_Protocol)\*\*
> > > 
> > > CLOSE\_WAIT essentially means that the PHP app (libcurl of some sort, I  
> > > assume) hasn't done a full job of closing the connection, e.g., closing the  
> > > TCP connection but not the underlying socket, so that's where I'd look.  
> > > For example, the option CURLOPT\_FORBID\_REUSE might be useful.
> > > 
> > > Is that really so?  
> > > I looked at this diagram: [http://en.wikipedia](http://en.wikipedia).[http://en.wikipedia.org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> > > o\*\*rg/wiki/File:TCP\_CLOSE.svg[http://en.wikipedia.org/wiki/File:TCP\_CLOSE.svg](http://en.wikipedia.org/wiki/File:TCP_CLOSE.svg)  
> > > If I read that correctly, it looks like CLOSE\_WAIT happens when the  
> > > client (left side) which initiated the connection issues a FIN, which I  
> > > understand as the client saying "I want to close this connection". After  
> > > that FIN is received by the server/receiver, that server/receiver goes into  
> > > the CLOSE\_WAIT state and at that time it is supposed to answer by sending  
> > > the ACK and then (after some time?) by sending FIN going going into  
> > > LAST\_ACK state and then, after client responds with ACK, into CLOSE state.
> > > 
> > > So if the server side is in CLOSE\_WAIT, doesn't that mean that the  
> > > server received a FIN, but did not send ACK back to client?
> > > 
> > > ## Thanks, Otis
> > > 
> > > Search Analytics - [Sematext | IT System Monitoring Tools for DevOps](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > > a [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)**nalytics/index.html[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > > Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> > > inde**x. [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)html[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > 
> > > On Jun 25, 2012, at 8:26 PM, Otis Gospodnetic wrote:
> > > 
> > > > Hello,
> > > > 
> > > > I've been fighting with ES that stops working after it gets hit for  
> > > > several minutes by about 1200 QPS. This is happening with ES 0.19.3 and  
> > > > 0.19.4 on big boxes (24 cores, 96 GB RAM).
> > > > 
> > > > What seems to be happening is that after a while we start seeing more  
> > > > and more and more CLOSE\_WAIT connections between search clients (~500  
> > > > frontend PHP apps) and ES, like this:
> > > > 
> > > > $ netstat -T | head  
> > > > Active Internet connections (w/o servers)  
> > > > Proto Recv-Q Send-Q Local Address Foreign Address  
> > > > State  
> > > > tcp 325 0 11.11.11.11-static.reverse.\*\*  
> > > > [softlayer.com](http://softlayer.com):wap-wsp 184.184.[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > > > **184.184-static.[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > > > reverse. [http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > > > softlay**[er.com:32035](http://er.com:32035)[http://184.184.184.184-static.reverse.softlayer.com:32035](http://184.184.184.184-static.reverse.softlayer.com:32035)  
> > > > CLOSE\_WAIT  
> > > > ...  
> > > > ...
> > > > 
> > > > The number of CLOSE\_WAIT connections goes from being 0 for a while to  
> > > > going up into thousands. And then at some point ES stops responding on  
> > > > port 9200 (but still responds on port 9300).
> > > > 
> > > > The number of these CLOSE\_WAIT connections seems to _roughly_  
> > > > correspond to the "current\_open" HTTP metric:
> > > > 
> > > > $ curl --silent 11.11.11.11:9200/_[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > > clu\*\*ster/[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > > nodes/stats?network=[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > > true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > > **transport=true&http=true&[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > > threa**d\_pool=true&indices=[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > > false[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)  
> > > > &\*\*pretty=true[http://11.11.11.11:9200/\_cluster/nodes/stats?network=true&transport=true&http=true&thread\_pool=true&indices=false&pretty=true](http://11.11.11.11:9200/_cluster/nodes/stats?network=true&transport=true&http=true&thread_pool=true&indices=false&pretty=true)'  
> > > > | egrep 'address|current|curr|server_\*\*open'  
> > > > "transport\_address" : "inet[/ 11.11.11.11 :9300]",  
> > > > "curr\_estab" : 87,  
> > > > "server\_open" : 36,  
> > > > "current\_open" : 7,  
> > > > \<== healthy, just restarted  
> > > > "transport\_address" : "inet[/ 22.22.22.22 :9300]",  
> > > > "curr\_estab" : 1245,  
> > > > "server\_open" : 612,  
> > > > "current\_open" : 14,  
> > > > \<== healthy, just restarted  
> > > > "transport\_address" : "inet[/ 33.33.33.33:9300]",  
> > > > "curr\_estab" : 93,  
> > > > "server\_open" : 36,  
> > > > "current\_open" : 14,  
> > > > \<== healthy, just restarted  
> > > > "transport\_address" : "inet[/ 44.44.44.44 :9300]",  
> > > > "curr\_estab" : 171,  
> > > > "server\_open" : 36,  
> > > > "current\_open" : 15776,  
> > > > \<== baaad, not restarted
> > > > 
> > > > I tried using using threadpool (both fixed and blocking with both  
> > > > abort and client rejection policies) to try stopping this "current\_open"  
> > > > from growing, e.g.:
> > > > 
> > > > threadpool:  
> > > > search:  
> > > > type: fixed  
> > > > size: 120  
> > > > queue\_size: 100  
> > > > reject\_policy: abort
> > > > 
> > > > But that didn't help.
> > > > 
> > > > I should say that the search apps hitting this ES cluster are not  
> > > > using persistent/keep-alive connections. And while this is clearly not  
> > > > ideal and not efficient, I think it still shouldn't cause this "leak" that  
> > > > ends up accumulating connections in CLOSE\_WAIT state and eventually getting  
> > > > ES to stop being responsive.
> > > > 
> > > > Is there anything one can do on the ES side to more aggressively close  
> > > > connections?
> > > > 
> > > > ## Thanks, Otis
> > > > 
> > > > Search Analytics - [Sematext | IT System Monitoring Tools for DevOps](http://sematext.com/search-)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > > > a [http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)\*\*nalytics/index.html[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)
> > > 
> > > > Scalable Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)  
> > > > inde\*\*x. [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)html[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > 
> > > >

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:21am UTC](https://discuss.elastic.co/t/increasing-close-wait-connections-and-http-current-open-metric/8223/17 "2017-07-06T03:21:50Z")

</div>


