# Transportclient retry logic/resiliency

**URL:** <https://discuss.elastic.co/t/transportclient-retry-logic-resiliency/21398>\
**Category:** Elasticsearch\
**Created:** [December 26, 2014, 2:41pm UTC](https://discuss.elastic.co/t/transportclient-retry-logic-resiliency/21398 "2014-12-26T14:41:49Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![abose78](https://avatars.discourse-cdn.com/v4/letter/a/9de0a6/32.png) [@abose78](https://discuss.elastic.co/u/abose78)\
**Post date:** [December 26, 2014, 2:41pm UTC](https://discuss.elastic.co/t/transportclient-retry-logic-resiliency/21398/1 "2014-12-26T14:41:49Z")

</div>

I have added 3 trasportclient nodes while creating a client.

Settings settings = ImmutableSettings.settingsBuilder()  
.put("cluster.name", clusterName)  
.put("client.transport.sniff", true)  
.build();  
TransportClient client = new TransportClient(settings);  
client.addTransportAddresses(new InetSocketTransportAddress(esHos1,  
esPort1));  
client.addTransportAddresses(new InetSocketTransportAddress(esHost2,  
esPort2));  
client.addTransportAddresses(new InetSocketTransportAddress(esHost3,  
esPort3));

esHost1 and esHost2 are down. But esHost3 is running. However, when I try  
to connect, its giving NoNodeAvailableException. What I was expecting as  
below as per the Round Robin logic for each actions:

1. try to connect to esHost1
2. NoNodeAvailableException after ping.timeout
3. try to connect to esHost2
4. NoNodeAvailableException after ping.timeout
5. try to connect to esHost3 - and successfully being able to connect.

So now I am beginning to think that the Round Robin is actually for the  
actions but not in case if there is a NoNodeAvailableException. Is that  
correct?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 26, 2014, 8:30pm UTC](https://discuss.elastic.co/t/transportclient-retry-logic-resiliency/21398/2 "2014-12-26T20:30:35Z")

</div>

If you get a NoNodeAvailableException, none of the hosts are available.

If you have "sniff" on, TransportClient tries to connect to all discovered  
nodes, this may take up to 10-15 seconds.

Steps 1-5 are performed in parallel, not sequential, i.e. each added  
transport address is instantly connected (or not).

The correct method is to add the known host addresses with  
addTransportAddresses() and afterwards check the connectedNodes() method.  
If it returns empty list, no nodes could be found.

Jörg

On Fri, Dec 26, 2014 at 3:41 PM, Arindam Bose [abose78@gmail.com](mailto:abose78@gmail.com) wrote:

> I have added 3 trasportclient nodes while creating a client.
> 
> Settings settings = ImmutableSettings.settingsBuilder()  
> .put("cluster.name", clusterName)  
> .put("client.transport.sniff", true)  
> .build();  
> TransportClient client = new TransportClient(settings);  
> client.addTransportAddresses(new InetSocketTransportAddress(esHos1,  
> esPort1));  
> client.addTransportAddresses(new InetSocketTransportAddress(esHost2,  
> esPort2));  
> client.addTransportAddresses(new InetSocketTransportAddress(esHost3,  
> esPort3));
> 
> esHost1 and esHost2 are down. But esHost3 is running. However, when I try  
> to connect, its giving NoNodeAvailableException. What I was expecting as  
> below as per the Round Robin logic for each actions:
> 
> 1. try to connect to esHost1
> 2. NoNodeAvailableException after ping.timeout
> 3. try to connect to esHost2
> 4. NoNodeAvailableException after ping.timeout
> 5. try to connect to esHost3 - and successfully being able to connect.
> 
> So now I am beginning to think that the Round Robin is actually for the  
> actions but not in case if there is a NoNodeAvailableException. Is that  
> correct?
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGP9%2BvDXaf05bN7nvQE1fCWMybwSF%3DjqBgaw7yOOajFFA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGP9%2BvDXaf05bN7nvQE1fCWMybwSF%3DjqBgaw7yOOajFFA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![abose78](https://avatars.discourse-cdn.com/v4/letter/a/9de0a6/32.png) [@abose78](https://discuss.elastic.co/u/abose78)\
**Post date:** [December 31, 2014, 2:22pm UTC](https://discuss.elastic.co/t/transportclient-retry-logic-resiliency/21398/3 "2014-12-31T14:22:57Z")

</div>

I understand the sniffing and steps 1-5 being parallel.

Now, I am still trying to infer why in my use case, I am getting the  
NoNodeAvailableException! I know esHos1 and esHost2 are down. Only esHost3  
is up and running. Considering what you said (steps 1-5 being parallel), so  
in my case, even before transportclient can sniff around and check that the esHos1  
and esHost2 are down, I am firing the persisting executions/actions. So  
transportclient did not get the time to remove those nodes from the list as  
was created by the initial 'addTransportAddresses'. So as a result as per  
the round robin logic the first node ehos1 was chosen to carry out the  
requested executions. As this host was down so I got the  
NoNodeAvailableException. Is that correct?

If that is true, then isnt there any retry logic in the transportclient to  
say, if the execution has failed in 1 node, to propagate the same execution  
to anyother node?

On Friday, December 26, 2014 2:30:41 PM UTC-6, Jörg Prante wrote:

> If you get a NoNodeAvailableException, none of the hosts are available.
> 
> If you have "sniff" on, TransportClient tries to connect to all discovered  
> nodes, this may take up to 10-15 seconds.
> 
> Steps 1-5 are performed in parallel, not sequential, i.e. each added  
> transport address is instantly connected (or not).
> 
> The correct method is to add the known host addresses with  
> addTransportAddresses() and afterwards check the connectedNodes() method.  
> If it returns empty list, no nodes could be found.
> 
> Jörg
> 
> On Fri, Dec 26, 2014 at 3:41 PM, Arindam Bose \<[abo...@gmail.com](mailto:abo...@gmail.com)  
> \<javascript:\>\> wrote:
> 
> > I have added 3 trasportclient nodes while creating a client.
> > 
> > Settings settings = ImmutableSettings.settingsBuilder()  
> > .put("cluster.name", clusterName)  
> > .put("client.transport.sniff", true)  
> > .build();  
> > TransportClient client = new TransportClient(settings);  
> > client.addTransportAddresses(new InetSocketTransportAddress(esHos1,  
> > esPort1));  
> > client.addTransportAddresses(new InetSocketTransportAddress(esHost2,  
> > esPort2));  
> > client.addTransportAddresses(new InetSocketTransportAddress(esHost3,  
> > esPort3));
> > 
> > esHost1 and esHost2 are down. But esHost3 is running. However, when I try  
> > to connect, its giving NoNodeAvailableException. What I was expecting as  
> > below as per the Round Robin logic for each actions:
> > 
> > 1. try to connect to esHost1
> > 2. NoNodeAvailableException after ping.timeout
> > 3. try to connect to esHost2
> > 4. NoNodeAvailableException after ping.timeout
> > 5. try to connect to esHost3 - and successfully being able to connect.
> > 
> > So now I am beginning to think that the Round Robin is actually for the  
> > actions but not in case if there is a NoNodeAvailableException. Is that  
> > correct?
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/b8560629-00d0-40f3-80e9-4367b6025c2c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/b8560629-00d0-40f3-80e9-4367b6025c2c%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [January 2, 2015, 5:25pm UTC](https://discuss.elastic.co/t/transportclient-retry-logic-resiliency/21398/4 "2015-01-02T17:25:38Z")

</div>

The TransportClient does not perform a retry/error logic on connected nodes.

When an action is executed, the TransportClient picks a connection from the  
pool and executes the action once on this connection. When there is no  
connection, the NoNodeAvailableException is thrown. When the action fails,  
it is reported straight to the user. There would be not much sense in  
retrying actions silently without user interaction, for example if an index  
operation fails, the user should be in charge of deciding what to do next.

Jörg

On Wed, Dec 31, 2014 at 3:22 PM, Arindam Bose [abose78@gmail.com](mailto:abose78@gmail.com) wrote:

> I understand the sniffing and steps 1-5 being parallel.
> 
> Now, I am still trying to infer why in my use case, I am getting the  
> NoNodeAvailableException! I know esHos1 and esHost2 are down. Only esHost3  
> is up and running. Considering what you said (steps 1-5 being parallel),  
> so in my case, even before transportclient can sniff around and check that  
> the esHos1 and esHost2 are down, I am firing the persisting  
> executions/actions. So transportclient did not get the time to remove those  
> nodes from the list as was created by the initial 'addTransportAddresses'.  
> So as a result as per the round robin logic the first node ehos1 was chosen  
> to carry out the requested executions. As this host was down so I got the  
> NoNodeAvailableException. Is that correct?
> 
> If that is true, then isnt there any retry logic in the transportclient to  
> say, if the execution has failed in 1 node, to propagate the same execution  
> to anyother node?
> 
> On Friday, December 26, 2014 2:30:41 PM UTC-6, Jörg Prante wrote:
> 
> > If you get a NoNodeAvailableException, none of the hosts are available.
> > 
> > If you have "sniff" on, TransportClient tries to connect to all  
> > discovered nodes, this may take up to 10-15 seconds.
> > 
> > Steps 1-5 are performed in parallel, not sequential, i.e. each added  
> > transport address is instantly connected (or not).
> > 
> > The correct method is to add the known host addresses with  
> > addTransportAddresses() and afterwards check the connectedNodes() method.  
> > If it returns empty list, no nodes could be found.
> > 
> > Jörg
> > 
> > On Fri, Dec 26, 2014 at 3:41 PM, Arindam Bose [abo...@gmail.com](mailto:abo...@gmail.com) wrote:
> > 
> > > I have added 3 trasportclient nodes while creating a client.
> > > 
> > > Settings settings = ImmutableSettings.settingsBuilder()  
> > > .put("cluster.name", clusterName)  
> > > .put("client.transport.sniff", true)  
> > > .build();  
> > > TransportClient client = new TransportClient(settings);  
> > > client.addTransportAddresses(new InetSocketTransportAddress(esHos1,  
> > > esPort1));  
> > > client.addTransportAddresses(new InetSocketTransportAddress(esHost2,  
> > > esPort2));  
> > > client.addTransportAddresses(new InetSocketTransportAddress(esHost3,  
> > > esPort3));
> > > 
> > > esHost1 and esHost2 are down. But esHost3 is running. However, when I  
> > > try to connect, its giving NoNodeAvailableException. What I was  
> > > expecting as below as per the Round Robin logic for each actions:
> > > 
> > > 1. try to connect to esHost1
> > > 2. NoNodeAvailableException after ping.timeout
> > > 3. try to connect to esHost2
> > > 4. NoNodeAvailableException after ping.timeout
> > > 5. try to connect to esHost3 - and successfully being able to connect.
> > > 
> > > So now I am beginning to think that the Round Robin is actually for the  
> > > actions but not in case if there is a NoNodeAvailableException. Is that  
> > > correct?
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > > To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> > > msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%  
> > > [40googlegroups.com](http://40googlegroups.com)  
> > > [https://groups.google.com/d/msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/1a5ef4f3-1046-4db6-9803-a308f51cf79b%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > .  
> > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/b8560629-00d0-40f3-80e9-4367b6025c2c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/b8560629-00d0-40f3-80e9-4367b6025c2c%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/b8560629-00d0-40f3-80e9-4367b6025c2c%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/b8560629-00d0-40f3-80e9-4367b6025c2c%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHfROh0%2BxbBDZ1VLentFJr3wNyEW%2B86ZaW5\_xHrNe0inA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHfROh0%2BxbBDZ1VLentFJr3wNyEW%2B86ZaW5_xHrNe0inA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:41am UTC](https://discuss.elastic.co/t/transportclient-retry-logic-resiliency/21398/5 "2017-07-06T00:41:08Z")

</div>


