# Disconnected transport client on ES 0.19.0

**URL:** <https://discuss.elastic.co/t/disconnected-transport-client-on-es-0-19-0/7167>\
**Category:** Elasticsearch\
**Created:** [March 28, 2012, 4:53pm UTC](https://discuss.elastic.co/t/disconnected-transport-client-on-es-0-19-0/7167 "2012-03-28T16:53:34Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![davie](https://avatars.discourse-cdn.com/v4/letter/d/b77776/32.png) [@davie](https://discuss.elastic.co/u/davie)\
**Post date:** [March 28, 2012, 4:53pm UTC](https://discuss.elastic.co/t/disconnected-transport-client-on-es-0-19-0/7167/1 "2012-03-28T16:53:34Z")

</div>

Hi,  
We recently upgraded to elasticsearch 0.19.0 and have been seeing the  
following errors occassionally:

JsQS0K1nRNCb8GNgA][inet[/192.168.1.100:9300]], disconnecting...  
org.elasticsearch.transport.ReceiveTimeoutTransportException:  
[prod-node-prod][inet[/192.168.1.100:9300]][cluster/nodes/info] request\_id  
[143146] timed out after [5002ms]  
at  
org.elasticsearch.transport.TransportService$TimeoutHandler.run(TransportService.java:347)

```
    at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown

```

Source)  
at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown  
Source)

We have 2 nodes in the cluster, and are connecting using the transport  
client.  
We can reproduce this error by running a query that takes tens of seconds  
to complete. Then if we increase the value of  
elasticsearch.client.transport.ping\_timeout on the client to longer than  
the query takes to return we don't get the error.

It looks like this may be related to the node disconnect on timeout  
introduced in  
[https://github.com/elasticsearch/elasticsearch/commit/eb4f6709d97287b0c9de6af9bf5f4a42fcf98991but](https://github.com/elasticsearch/elasticsearch/commit/eb4f6709d97287b0c9de6af9bf5f4a42fcf98991but)  
I don't really understand what that timeout should do - could you  
explain?  
We can't reproduce the issue using curl, the long running queries return  
without any problem.

We realise that long running queries should probably be using scan, so we  
will switch to use that, but it would be nice to understand the timeout  
behaviour too.

Thanks,  
Davie

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 29, 2012, 12:17pm UTC](https://discuss.elastic.co/t/disconnected-transport-client-on-es-0-19-0/7167/2 "2012-03-29T12:17:19Z")

</div>

When you say a query that takes long time to return, do you mean a query  
that returns a very large resultset, or simply a query that takes long to  
execute but returns a "not so large" result set? We can improve things by  
using a different channel to do the ping, I will do it, but still, would be  
interesting to investigate this further.

On Wed, Mar 28, 2012 at 6:53 PM, Davie Moston [daviemoston@gmail.com](mailto:daviemoston@gmail.com) wrote:

> Hi,  
> We recently upgraded to elasticsearch 0.19.0 and have been seeing the  
> following errors occassionally:
> 
> JsQS0K1nRNCb8GNgA][inet[/192.168.1.100:9300]], disconnecting...  
> org.elasticsearch.transport.ReceiveTimeoutTransportException:  
> [prod-node-prod][inet[/192.168.1.100:9300]][cluster/nodes/info]  
> request\_id [143146] timed out after [5002ms]  
> at  
> org.elasticsearch.transport.TransportService$TimeoutHandler.run(TransportService.java:347)
> 
> ```
> at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown
> 
> ```
> 
> Source)  
> at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown  
> Source)
> 
> We have 2 nodes in the cluster, and are connecting using the transport  
> client.  
> We can reproduce this error by running a query that takes tens of seconds  
> to complete. Then if we increase the value of  
> elasticsearch.client.transport.ping\_timeout on the client to longer than  
> the query takes to return we don't get the error.
> 
> It looks like this may be related to the node disconnect on timeout  
> introduced in  
> [https://github.com/elasticsearch/elasticsearch/commit/eb4f6709d97287b0c9de6af9bf5f4a42fcf98991but](https://github.com/elasticsearch/elasticsearch/commit/eb4f6709d97287b0c9de6af9bf5f4a42fcf98991but) I don't really understand what that timeout should do - could you  
> explain?  
> We can't reproduce the issue using curl, the long running queries return  
> without any problem.
> 
> We realise that long running queries should probably be using scan, so we  
> will switch to use that, but it would be nice to understand the timeout  
> behaviour too.
> 
> Thanks,  
> Davie

---

<div class="post-metadata">

**Author:** ![davie](https://avatars.discourse-cdn.com/v4/letter/d/b77776/32.png) [@davie](https://discuss.elastic.co/u/davie)\
**Post date:** [March 29, 2012, 1:12pm UTC](https://discuss.elastic.co/t/disconnected-transport-client-on-es-0-19-0/7167/3 "2012-03-29T13:12:21Z")

</div>

Thanks for the reply.  
The result set itself is fairly large.  
Issuing the same query with curl returns around 145 meg of json, which  
is roughly 15 thousand documents. This takes around a minute over  
http, and data starts being returned almost immediately.  
We also see intermittent failures for index requests using the same  
client which we suspect are related as they happen at the same time.  
In this case we see a nodedisconnected exception.  
For this query we should probably switch to using scan/scroll, but I  
had expected it to use the timeout passed in with the query rather  
than causing a ping timeout.  
This failure is new in 0.19.0, we were previously running 0.18.7.0 and  
it worked fine.  
Let me know if you need any more info.  
Thanks,  
Davie

On Mar 29, 1:17 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> When you say a query that takes long time to return, do you mean a query  
> that returns a very large resultset, or simply a query that takes long to  
> execute but returns a "not so large" result set? We can improve things by  
> using a different channel to do the ping, I will do it, but still, would be  
> interesting to investigate this further.
> 
> On Wed, Mar 28, 2012 at 6:53 PM, Davie Moston [daviemos...@gmail.com](mailto:daviemos...@gmail.com) wrote:
> 
> > Hi,  
> > We recently upgraded to elasticsearch 0.19.0 and have been seeing the  
> > following errors occassionally:
> 
> > JsQS0K1nRNCb8GNgA][inet[/192.168.1.100:9300]], disconnecting...  
> > org.elasticsearch.transport.ReceiveTimeoutTransportException:  
> > [prod-node-prod][inet[/192.168.1.100:9300]][cluster/nodes/info]  
> > request\_id [143146] timed out after [5002ms]  
> > at  
> > org.elasticsearch.transport.TransportService$TimeoutHandler.run(TransportSe rvice.java:347)
> 
> > ```
> > at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown
> > 
> > ```
> > 
> > Source)  
> > at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown  
> > Source)
> 
> > We have 2 nodes in the cluster, and are connecting using the transport  
> > client.  
> > We can reproduce this error by running a query that takes tens of seconds  
> > to complete. Then if we increase the value of  
> > elasticsearch.client.transport.ping\_timeout on the client to longer than  
> > the query takes to return we don't get the error.
> 
> > It looks like this may be related to the node disconnect on timeout  
> > introduced in  
> > [https://github.com/elasticsearch/elasticsearch/commit/eb4f6709d97287b...I](https://github.com/elasticsearch/elasticsearch/commit/eb4f6709d97287b...I) don't really understand what that timeout should do - could you  
> > explain?  
> > We can't reproduce the issue using curl, the long running queries return  
> > without any problem.
> 
> > We realise that long running queries should probably be using scan, so we  
> > will switch to use that, but it would be nice to understand the timeout  
> > behaviour too.
> 
> > Thanks,  
> > Davie

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 29, 2012, 2:21pm UTC](https://discuss.elastic.co/t/disconnected-transport-client-on-es-0-19-0/7167/4 "2012-03-29T14:21:01Z")

</div>

It might take time to move 145mb on the channel, so possibly the ping  
request from the transport client gets delayed because of it (wiht the  
default 5 seconds ping timeout). I pushed a change to use the "high"  
priority transport channel for ping requests to both master and upcoming  
0.19.2.

On Thu, Mar 29, 2012 at 3:12 PM, davie [daviemoston@gmail.com](mailto:daviemoston@gmail.com) wrote:

> Thanks for the reply.  
> The result set itself is fairly large.  
> Issuing the same query with curl returns around 145 meg of json, which  
> is roughly 15 thousand documents. This takes around a minute over  
> http, and data starts being returned almost immediately.  
> We also see intermittent failures for index requests using the same  
> client which we suspect are related as they happen at the same time.  
> In this case we see a nodedisconnected exception.  
> For this query we should probably switch to using scan/scroll, but I  
> had expected it to use the timeout passed in with the query rather  
> than causing a ping timeout.  
> This failure is new in 0.19.0, we were previously running 0.18.7.0 and  
> it worked fine.  
> Let me know if you need any more info.  
> Thanks,  
> Davie
> 
> On Mar 29, 1:17 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > When you say a query that takes long time to return, do you mean a query  
> > that returns a very large resultset, or simply a query that takes long to  
> > execute but returns a "not so large" result set? We can improve things by  
> > using a different channel to do the ping, I will do it, but still, would  
> > be  
> > interesting to investigate this further.
> > 
> > On Wed, Mar 28, 2012 at 6:53 PM, Davie Moston [daviemos...@gmail.com](mailto:daviemos...@gmail.com)  
> > wrote:
> > 
> > > Hi,  
> > > We recently upgraded to elasticsearch 0.19.0 and have been seeing the  
> > > following errors occassionally:
> > 
> > > JsQS0K1nRNCb8GNgA][inet[/192.168.1.100:9300]], disconnecting...  
> > > org.elasticsearch.transport.ReceiveTimeoutTransportException:  
> > > [prod-node-prod][inet[/192.168.1.100:9300]][cluster/nodes/info]  
> > > request\_id [143146] timed out after [5002ms]  
> > > at
> 
> org.elasticsearch.transport.TransportService$TimeoutHandler.run(TransportSe  
> rvice.java:347)
> 
> > > ```
> > > at
> > > 
> > > ```
> 
> java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown
> 
> > > Source)  
> > > at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown  
> > > Source)
> > 
> > > We have 2 nodes in the cluster, and are connecting using the transport  
> > > client.  
> > > We can reproduce this error by running a query that takes tens of  
> > > seconds  
> > > to complete. Then if we increase the value of  
> > > elasticsearch.client.transport.ping\_timeout on the client to longer  
> > > than  
> > > the query takes to return we don't get the error.
> > 
> > > It looks like this may be related to the node disconnect on timeout  
> > > introduced in
> 
> [https://github.com/elasticsearch/elasticsearch/commit/eb4f6709d97287b...Idon't](https://github.com/elasticsearch/elasticsearch/commit/eb4f6709d97287b...Idon't) really understand what that timeout should do - could you
> 
> > > explain?  
> > > We can't reproduce the issue using curl, the long running queries  
> > > return  
> > > without any problem.
> > 
> > > We realise that long running queries should probably be using scan, so  
> > > we  
> > > will switch to use that, but it would be nice to understand the timeout  
> > > behaviour too.
> > 
> > > Thanks,  
> > > Davie

---

<div class="post-metadata">

**Author:** ![davie](https://avatars.discourse-cdn.com/v4/letter/d/b77776/32.png) [@davie](https://discuss.elastic.co/u/davie)\
**Post date:** [March 30, 2012, 12:55pm UTC](https://discuss.elastic.co/t/disconnected-transport-client-on-es-0-19-0/7167/5 "2012-03-30T12:55:45Z")

</div>

Thanks for the quick turnaround.  
I just tried upgrading the elastic search jar on the client side to the latest 0.19 branch in git, but the problem still seems to be there.  
Looking at the change it seems to only affect the client side, but maybe I'm missing something - I'll try upgrading the nodes too and see if that helps.  
Thanks,  
Davie

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 31, 2012, 8:49pm UTC](https://discuss.elastic.co/t/disconnected-transport-client-on-es-0-19-0/7167/6 "2012-03-31T20:49:03Z")

</div>

Yea, it might be a bit trickier to solve... (though we have a step in the  
right direction). Open an issue for it? Btw, can you create a standlone  
testcase for this? It would help speed things around...

On Fri, Mar 30, 2012 at 3:55 PM, davie [daviemoston@gmail.com](mailto:daviemoston@gmail.com) wrote:

> Thanks for the quick turnaround.  
> I just tried upgrading the Elasticsearch jar on the client side to the  
> latest 0.19 branch in git, but the problem still seems to be there.  
> Looking at the change it seems to only affect the client side, but maybe  
> I'm missing something - I'll try upgrading the nodes too and see if that  
> helps.  
> Thanks,  
> Davie

---

<div class="post-metadata">

**Author:** ![davie](https://avatars.discourse-cdn.com/v4/letter/d/b77776/32.png) [@davie](https://discuss.elastic.co/u/davie)\
**Post date:** [April 26, 2012, 6:44pm UTC](https://discuss.elastic.co/t/disconnected-transport-client-on-es-0-19-0/7167/7 "2012-04-26T18:44:31Z")

</div>

I've created an issue for this here  
[transport client disconnected after ping timeout · Issue #1886 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/1886) and a testcase  
here [tescase for https://github.com/elasticsearch/elasticsearch/issues/1886 · GitHub](https://gist.github.com/2501687).  
We've been running with client.transport.ping\_timeout=180s as a workaround  
and haven't seen the problem since.

Thanks,  
Davie

On Saturday, 31 March 2012 21:49:03 UTC+1, kimchy wrote:

> Yea, it might be a bit trickier to solve... (though we have a step in the  
> right direction). Open an issue for it? Btw, can you create a standlone  
> testcase for this? It would help speed things around...
> 
> On Fri, Mar 30, 2012 at 3:55 PM, davie [daviemoston@gmail.com](mailto:daviemoston@gmail.com) wrote:
> 
> > Thanks for the quick turnaround.  
> > I just tried upgrading the Elasticsearch jar on the client side to the  
> > latest 0.19 branch in git, but the problem still seems to be there.  
> > Looking at the change it seems to only affect the client side, but maybe  
> > I'm missing something - I'll try upgrading the nodes too and see if that  
> > helps.  
> > Thanks,  
> > Davie

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [April 27, 2012, 9:53pm UTC](https://discuss.elastic.co/t/disconnected-transport-client-on-es-0-19-0/7167/8 "2012-04-27T21:53:46Z")

</div>

Commented on the issue, lets move the discussion there then (just so we  
have a single place).

On Thu, Apr 26, 2012 at 9:44 PM, davie [daviemoston@gmail.com](mailto:daviemoston@gmail.com) wrote:

> I've created an issue for this here  
> [transport client disconnected after ping timeout · Issue #1886 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/1886) and a testcase  
> here [tescase for https://github.com/elasticsearch/elasticsearch/issues/1886 · GitHub](https://gist.github.com/2501687).  
> We've been running with client.transport.ping\_timeout=180s as a workaround  
> and haven't seen the problem since.
> 
> Thanks,  
> Davie
> 
> On Saturday, 31 March 2012 21:49:03 UTC+1, kimchy wrote:
> 
> > Yea, it might be a bit trickier to solve... (though we have a step in the  
> > right direction). Open an issue for it? Btw, can you create a standlone  
> > testcase for this? It would help speed things around...
> > 
> > On Fri, Mar 30, 2012 at 3:55 PM, davie [daviemoston@gmail.com](mailto:daviemoston@gmail.com) wrote:
> > 
> > > Thanks for the quick turnaround.  
> > > I just tried upgrading the Elasticsearch jar on the client side to the  
> > > latest 0.19 branch in git, but the problem still seems to be there.  
> > > Looking at the change it seems to only affect the client side, but maybe  
> > > I'm missing something - I'll try upgrading the nodes too and see if that  
> > > helps.  
> > > Thanks,  
> > > Davie

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:30am UTC](https://discuss.elastic.co/t/disconnected-transport-client-on-es-0-19-0/7167/9 "2017-07-06T03:30:48Z")

</div>


