# TCP Transport versus Transport Client settings

**URL:** <https://discuss.elastic.co/t/tcp-transport-versus-transport-client-settings/8709>\
**Category:** Elasticsearch\
**Created:** [August 10, 2012, 6:38pm UTC](https://discuss.elastic.co/t/tcp-transport-versus-transport-client-settings/8709 "2012-08-10T18:38:46Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 10, 2012, 6:38pm UTC](https://discuss.elastic.co/t/tcp-transport-versus-transport-client-settings/8709/1 "2012-08-10T18:38:46Z")

</div>

Trying to solve intermittent "NoNodeAvailableException: No node  
available" errors that occur while searching. Cluster consists of 4  
nodes running 0.19.2 using multicast. Client is a singleton  
TransportClient configured with all the nodes defined in the settings.  
client.transport.sniff is set to true. All other settings for either  
client or server are the default. Queries have a timeout of 500ms.

Increasing the timeout limits to hopefully eliminate the problem.  
First question is would it be possible to determine which node was  
trying to be accessed when NoNodeAvailableException was returned?  
Perhaps only one node has issue. For the transport client,  
client.transport.ping\_timeout would be the setting to change, but the  
current default of 5 seconds seems already high. Is the transport  
client communication solely to blame or could inter-node (TCP  
Transport) communication be to blame as well? Should those settings be  
modified as well?

Cheers,

Ivan

--

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 13, 2012, 10:18am UTC](https://discuss.elastic.co/t/tcp-transport-versus-transport-client-settings/8709/2 "2012-08-13T10:18:56Z")

</div>

If you set to debug the client.transport (or org.elasticsearch.client.transport if embedded) do you see disconnections? Can you try and use a newer 0.19 version, the logic of the transport client has been improved in later versions.

On Aug 10, 2012, at 8:38 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> Trying to solve intermittent "NoNodeAvailableException: No node  
> available" errors that occur while searching. Cluster consists of 4  
> nodes running 0.19.2 using multicast. Client is a singleton  
> TransportClient configured with all the nodes defined in the settings.  
> client.transport.sniff is set to true. All other settings for either  
> client or server are the default. Queries have a timeout of 500ms.
> 
> Increasing the timeout limits to hopefully eliminate the problem.  
> First question is would it be possible to determine which node was  
> trying to be accessed when NoNodeAvailableException was returned?  
> Perhaps only one node has issue. For the transport client,  
> client.transport.ping\_timeout would be the setting to change, but the  
> current default of 5 seconds seems already high. Is the transport  
> client communication solely to blame or could inter-node (TCP  
> Transport) communication be to blame as well? Should those settings be  
> modified as well?
> 
> Cheers,
> 
> Ivan
> 
> --

--

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 13, 2012, 5:01pm UTC](https://discuss.elastic.co/t/tcp-transport-versus-transport-client-settings/8709/3 "2012-08-13T17:01:58Z")

</div>

Thanks Shay.

We are using Lucene 3.5 for other parts of the code on the client  
side, so not quite ready to move to Lucene 3.6 (needs testing). The  
issue is not consistent, so it has been difficult to reproduce  
faithfully the problem. Will stress test with debug on. Should both  
the client and server have debug enable. I am assuming the disconnect  
is on the client side.

Ivan

On Mon, Aug 13, 2012 at 3:18 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> If you set to debug the client.transport (or org.elasticsearch.client.transport if embedded) do you see disconnections? Can you try and use a newer 0.19 version, the logic of the transport client has been improved in later versions.
> 
> On Aug 10, 2012, at 8:38 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:
> 
> > Trying to solve intermittent "NoNodeAvailableException: No node  
> > available" errors that occur while searching. Cluster consists of 4  
> > nodes running 0.19.2 using multicast. Client is a singleton  
> > TransportClient configured with all the nodes defined in the settings.  
> > client.transport.sniff is set to true. All other settings for either  
> > client or server are the default. Queries have a timeout of 500ms.
> > 
> > Increasing the timeout limits to hopefully eliminate the problem.  
> > First question is would it be possible to determine which node was  
> > trying to be accessed when NoNodeAvailableException was returned?  
> > Perhaps only one node has issue. For the transport client,  
> > client.transport.ping\_timeout would be the setting to change, but the  
> > current default of 5 seconds seems already high. Is the transport  
> > client communication solely to blame or could inter-node (TCP  
> > Transport) communication be to blame as well? Should those settings be  
> > modified as well?
> > 
> > Cheers,
> > 
> > Ivan
> > 
> > --
> 
> --

--

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 13, 2012, 7:04pm UTC](https://discuss.elastic.co/t/tcp-transport-versus-transport-client-settings/8709/4 "2012-08-13T19:04:29Z")

</div>

Just the client side needs logging. Also, you can safely run the transport client with Lucene 3.5 and not use Lucene 3.6.

On Aug 13, 2012, at 7:01 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> Thanks Shay.
> 
> We are using Lucene 3.5 for other parts of the code on the client  
> side, so not quite ready to move to Lucene 3.6 (needs testing). The  
> issue is not consistent, so it has been difficult to reproduce  
> faithfully the problem. Will stress test with debug on. Should both  
> the client and server have debug enable. I am assuming the disconnect  
> is on the client side.
> 
> Ivan
> 
> On Mon, Aug 13, 2012 at 3:18 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > If you set to debug the client.transport (or org.elasticsearch.client.transport if embedded) do you see disconnections? Can you try and use a newer 0.19 version, the logic of the transport client has been improved in later versions.
> > 
> > On Aug 10, 2012, at 8:38 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:
> > 
> > > Trying to solve intermittent "NoNodeAvailableException: No node  
> > > available" errors that occur while searching. Cluster consists of 4  
> > > nodes running 0.19.2 using multicast. Client is a singleton  
> > > TransportClient configured with all the nodes defined in the settings.  
> > > client.transport.sniff is set to true. All other settings for either  
> > > client or server are the default. Queries have a timeout of 500ms.
> > > 
> > > Increasing the timeout limits to hopefully eliminate the problem.  
> > > First question is would it be possible to determine which node was  
> > > trying to be accessed when NoNodeAvailableException was returned?  
> > > Perhaps only one node has issue. For the transport client,  
> > > client.transport.ping\_timeout would be the setting to change, but the  
> > > current default of 5 seconds seems already high. Is the transport  
> > > client communication solely to blame or could inter-node (TCP  
> > > Transport) communication be to blame as well? Should those settings be  
> > > modified as well?
> > > 
> > > Cheers,
> > > 
> > > Ivan
> > > 
> > > --
> > 
> > --
> 
> --

--

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [August 13, 2012, 11:20pm UTC](https://discuss.elastic.co/t/tcp-transport-versus-transport-client-settings/8709/5 "2012-08-13T23:20:56Z")

</div>

Hi Ivan,

my recommendation is also upgrading from 0.19.2 to a newer version because  
there were issues with TransportClient sniffing. For an example, see #1819.

Best regards,

Jörg

On Friday, August 10, 2012 8:38:46 PM UTC+2, Ivan Brusic wrote:

> Trying to solve intermittent "NoNodeAvailableException: No node  
> available" errors that occur while searching. Cluster consists of 4  
> nodes running 0.19.2 using multicast. Client is a singleton  
> TransportClient configured with all the nodes defined in the settings.  
> client.transport.sniff is set to true. All other settings for either  
> client or server are the default. Queries have a timeout of 500ms.
> 
> Increasing the timeout limits to hopefully eliminate the problem.  
> First question is would it be possible to determine which node was  
> trying to be accessed when NoNodeAvailableException was returned?  
> Perhaps only one node has issue. For the transport client,  
> client.transport.ping\_timeout would be the setting to change, but the  
> current default of 5 seconds seems already high. Is the transport  
> client communication solely to blame or could inter-node (TCP  
> Transport) communication be to blame as well? Should those settings be  
> modified as well?
> 
> Cheers,
> 
> Ivan

--

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 14, 2012, 12:42am UTC](https://discuss.elastic.co/t/tcp-transport-versus-transport-client-settings/8709/6 "2012-08-14T00:42:29Z")

</div>

Hopefully I would never encounter #1819 since running with no nodes is  
not where I want to be!

Ran some stress tests for a few hours without causing an issue. Then  
when I was running some other queries, I managed to get into the  
erroneous state:

> <https://gist.github.com/brusic/88222848a398f813fdb0>

The client timed out although it was executing queries at the time.  
Converted the code to Lucene 3.6 (only one API change) and will test  
later with 0.19.8.

Cheers,

Ivan

On Mon, Aug 13, 2012 at 4:20 PM, Jörg Prante [joergprante@gmail.com](mailto:joergprante@gmail.com) wrote:

> Hi Ivan,
> 
> my recommendation is also upgrading from 0.19.2 to a newer version because  
> there were issues with TransportClient sniffing. For an example, see #1819.
> 
> Best regards,
> 
> Jörg
> 
> On Friday, August 10, 2012 8:38:46 PM UTC+2, Ivan Brusic wrote:
> 
> > Trying to solve intermittent "NoNodeAvailableException: No node  
> > available" errors that occur while searching. Cluster consists of 4  
> > nodes running 0.19.2 using multicast. Client is a singleton  
> > TransportClient configured with all the nodes defined in the settings.  
> > client.transport.sniff is set to true. All other settings for either  
> > client or server are the default. Queries have a timeout of 500ms.
> > 
> > Increasing the timeout limits to hopefully eliminate the problem.  
> > First question is would it be possible to determine which node was  
> > trying to be accessed when NoNodeAvailableException was returned?  
> > Perhaps only one node has issue. For the transport client,  
> > client.transport.ping\_timeout would be the setting to change, but the  
> > current default of 5 seconds seems already high. Is the transport  
> > client communication solely to blame or could inter-node (TCP  
> > Transport) communication be to blame as well? Should those settings be  
> > modified as well?
> > 
> > Cheers,
> > 
> > Ivan
> 
> --

--

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 14, 2012, 7:17pm UTC](https://discuss.elastic.co/t/tcp-transport-versus-transport-client-settings/8709/7 "2012-08-14T19:17:36Z")

</div>

Not only have I been able to replicate the issue, but the problem is  
now consistent. Now we have been getting "RemoteTransportException ...  
OutOfMemoryError" errors as well.

Here are the recent errors: [More "no node available" errors · GitHub](https://gist.github.com/c6728e50b40a34a9c42a)

The only relevant commit that I see to the Transport Client is

> <https://github.com/elastic/elasticsearch/commit/f01acb20e157f9e85859567ba4a84cec3048e5ca>

Will upgrade to 0.19.8 today. Just noticed another commit  
(bdea0e2eddb4373b850e00d8e363c5240d78d180) that I hope gets released  
soon as well (I wrote identical code, but prefer to a standard class  
whenever possible).

Cheers,

Ivan

On Mon, Aug 13, 2012 at 5:42 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> Hopefully I would never encounter #1819 since running with no nodes is  
> not where I want to be!
> 
> Ran some stress tests for a few hours without causing an issue. Then  
> when I was running some other queries, I managed to get into the  
> erroneous state:
> 
> [NoNodeAvailableException: No node available · GitHub](https://gist.github.com/88222848a398f813fdb0)
> 
> The client timed out although it was executing queries at the time.  
> Converted the code to Lucene 3.6 (only one API change) and will test  
> later with 0.19.8.
> 
> Cheers,
> 
> Ivan
> 
> On Mon, Aug 13, 2012 at 4:20 PM, Jörg Prante [joergprante@gmail.com](mailto:joergprante@gmail.com) wrote:
> 
> > Hi Ivan,
> > 
> > my recommendation is also upgrading from 0.19.2 to a newer version because  
> > there were issues with TransportClient sniffing. For an example, see #1819.
> > 
> > Best regards,
> > 
> > Jörg
> > 
> > On Friday, August 10, 2012 8:38:46 PM UTC+2, Ivan Brusic wrote:
> > 
> > > Trying to solve intermittent "NoNodeAvailableException: No node  
> > > available" errors that occur while searching. Cluster consists of 4  
> > > nodes running 0.19.2 using multicast. Client is a singleton  
> > > TransportClient configured with all the nodes defined in the settings.  
> > > client.transport.sniff is set to true. All other settings for either  
> > > client or server are the default. Queries have a timeout of 500ms.
> > > 
> > > Increasing the timeout limits to hopefully eliminate the problem.  
> > > First question is would it be possible to determine which node was  
> > > trying to be accessed when NoNodeAvailableException was returned?  
> > > Perhaps only one node has issue. For the transport client,  
> > > client.transport.ping\_timeout would be the setting to change, but the  
> > > current default of 5 seconds seems already high. Is the transport  
> > > client communication solely to blame or could inter-node (TCP  
> > > Transport) communication be to blame as well? Should those settings be  
> > > modified as well?
> > > 
> > > Cheers,
> > > 
> > > Ivan
> > 
> > --

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:16am UTC](https://discuss.elastic.co/t/tcp-transport-versus-transport-client-settings/8709/8 "2017-07-06T03:16:21Z")

</div>


