# Random search errors from one node of three on a previously working cluster

**URL:** <https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788>\
**Category:** Elasticsearch\
**Created:** [July 6, 2011, 8:46am UTC](https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788 "2011-07-06T08:46:21Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![ian\_clark](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ian_clark/32/3129_2.png) [@ian\_clark](https://discuss.elastic.co/u/ian_clark)\
**Post date:** [July 6, 2011, 8:46am UTC](https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788/1 "2011-07-06T08:46:21Z")

</div>

Hi,

I'm getting some random errors that started happening today. This is  
one error:-

{"error":"ReduceSearchPhaseException[Failed to execute phase [fetch],  
[reduce] ]; nested: ArrayIndexOutOfBoundsException[1]; ","status":500}

and this is another:-

{"error":"ReduceSearchPhaseException[Failed to execute phase [fetch],  
[reduce] ; shardFailures {SearchContextMissingException[No search  
context found for id [36813]]}{RemoteTransportException[[Aardwolf]  
[inet[/10.206.38.97:9300]][search/phase/fetch/id]]; nested:  
SearchContextMissingException[No search context found for id  
[55197]]; }]; nested: IndexOutOfBoundsException[index (0) must be less  
than size (0)]; ","status":500}

I am load testing my cluster by sending pseudo-random generated search  
requests. I can take the search requests that fail, and replay them  
manually and they work. Strange, as this test was working for the past  
few days.... I have found exceptions in my logs on only one node, I  
have copied them here:-

> <https://gist.github.com/ianAndrewClark/1066838>

Any ideas what could be going wrong? Shall I raise a bug?

cheers

Ian

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [July 6, 2011, 9:04am UTC](https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788/2 "2011-07-06T09:04:59Z")

</div>

Hi Ian

> I am load testing my cluster by sending pseudo-random generated search  
> requests. I can take the search requests that fail, and replay them  
> manually and they work. Strange, as this test was working for the past  
> few days.... I have found exceptions in my logs on only one node, I  
> have copied them here:-
> 
> [search response errors · GitHub](https://gist.github.com/1066838)

I see that the string of errors start with this one:

Caught exception while handling client http traffic, closing connection  
java.lang.IllegalStateException: cannot send more responses than  
requests

I see this error a few times a day, and it starts a run of bad  
responses, eg I search index A and the results contain docs from index A  
and index B, or the wrong types are returned etc

It's like the responses from each node are being mixed up.

Do you have any way of replicating this? even infrequently?

It would help greatly in the debugging process

thanks

Clint

---

<div class="post-metadata">

**Author:** ![ian\_clark](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ian_clark/32/3129_2.png) [@ian\_clark](https://discuss.elastic.co/u/ian_clark)\
**Post date:** [July 6, 2011, 9:39am UTC](https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788/3 "2011-07-06T09:39:04Z")

</div>

After a period of getting those responses, I restarted the test to  
capture responses to file, and damn it has not happened since. ☹

This was the first test run after I restarted the nodes one by one  
without issuing any shutdowns, maybe that could cause that same state?  
I will try that again.

Ian

On Jul 6, 10:04 am, Clinton Gormley [clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) wrote:

> Hi Ian
> 
> > I am load testing my cluster by sending pseudo-random generated search  
> > requests. I can take the search requests that fail, and replay them  
> > manually and they work. Strange, as this test was working for the past  
> > few days.... I have found exceptions in my logs on only one node, I  
> > have copied them here:-
> 
> > [search response errors · GitHub](https://gist.github.com/1066838)
> 
> I see that the string of errors start with this one:
> 
> Caught exception while handling client http traffic, closing connection  
> java.lang.IllegalStateException: cannot send more responses than  
> requests
> 
> I see this error a few times a day, and it starts a run of bad  
> responses, eg I search index A and the results contain docs from index A  
> and index B, or the wrong types are returned etc
> 
> It's like the responses from each node are being mixed up.
> 
> Do you have any way of replicating this? even infrequently?
> 
> It would help greatly in the debugging process
> 
> thanks
> 
> Clint

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 6, 2011, 4:08pm UTC](https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788/4 "2011-07-06T16:08:46Z")

</div>

If there is a way to recreate it, I would love to try and nail it...

On Wednesday, July 6, 2011 at 12:39 PM, Ian wrote:

> After a period of getting those responses, I restarted the test to  
> capture responses to file, and damn it has not happened since. ☹
> 
> This was the first test run after I restarted the nodes one by one  
> without issuing any shutdowns, maybe that could cause that same state?  
> I will try that again.
> 
> Ian
> 
> On Jul 6, 10:04 am, Clinton Gormley \<[clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) ([http://iannounce.co.uk](http://iannounce.co.uk))\> wrote:
> 
> > Hi Ian
> > 
> > > I am load testing my cluster by sending pseudo-random generated search  
> > > requests. I can take the search requests that fail, and replay them  
> > > manually and they work. Strange, as this test was working for the past  
> > > few days.... I have found exceptions in my logs on only one node, I  
> > > have copied them here:-
> > 
> > > [search response errors · GitHub](https://gist.github.com/1066838)
> > 
> > I see that the string of errors start with this one:
> > 
> > Caught exception while handling client http traffic, closing connection  
> > java.lang.IllegalStateException: cannot send more responses than  
> > requests
> > 
> > I see this error a few times a day, and it starts a run of bad  
> > responses, eg I search index A and the results contain docs from index A  
> > and index B, or the wrong types are returned etc
> > 
> > It's like the responses from each node are being mixed up.
> > 
> > Do you have any way of replicating this? even infrequently?
> > 
> > It would help greatly in the debugging process
> > 
> > thanks
> > 
> > Clint

---

<div class="post-metadata">

**Author:** ![ian\_clark](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ian_clark/32/3129_2.png) [@ian\_clark](https://discuss.elastic.co/u/ian_clark)\
**Post date:** [July 7, 2011, 10:59am UTC](https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788/5 "2011-07-07T10:59:58Z")

</div>

Hi Shay,

I'm still trying to reproduce this reliably, I repeated the same  
steps, process kill the nodes and bring them back, and I haven't been  
able to reliably reproduce it. I will keep trying as I need to get to  
the bottom of this.

It has happened once today however, is there any useful log settings I  
could add? (I'll try setting them all to debug)

Ian

On Jul 6, 5:08 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> If there is a way to recreate it, I would love to try and nail it...
> 
> On Wednesday, July 6, 2011 at 12:39 PM, Ian wrote:
> 
> > After a period of getting those responses, I restarted the test to  
> > capture responses to file, and damn it has not happened since. ☹
> 
> > This was the first test run after I restarted the nodes one by one  
> > without issuing any shutdowns, maybe that could cause that same state?  
> > I will try that again.
> 
> > Ian
> 
> > On Jul 6, 10:04 am, Clinton Gormley \<[clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) ([http://iannounce.co.uk](http://iannounce.co.uk))\> wrote:
> > 
> > > Hi Ian
> 
> > > > I am load testing my cluster by sending pseudo-random generated search  
> > > > requests. I can take the search requests that fail, and replay them  
> > > > manually and they work. Strange, as this test was working for the past  
> > > > few days.... I have found exceptions in my logs on only one node, I  
> > > > have copied them here:-
> 
> > > > [search response errors · GitHub](https://gist.github.com/1066838)
> 
> > > I see that the string of errors start with this one:
> 
> > > Caught exception while handling client http traffic, closing connection  
> > > java.lang.IllegalStateException: cannot send more responses than  
> > > requests
> 
> > > I see this error a few times a day, and it starts a run of bad  
> > > responses, eg I search index A and the results contain docs from index A  
> > > and index B, or the wrong types are returned etc
> 
> > > It's like the responses from each node are being mixed up.
> 
> > > Do you have any way of replicating this? even infrequently?
> 
> > > It would help greatly in the debugging process
> 
> > > thanks
> 
> > > Clint

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 7, 2011, 2:42pm UTC](https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788/6 "2011-07-07T14:42:35Z")

</div>

Ian, nothing specific... . Which client lib to talk to ES are you using?

On Thursday, July 7, 2011 at 1:59 PM, Ian wrote:

> Hi Shay,
> 
> I'm still trying to reproduce this reliably, I repeated the same  
> steps, process kill the nodes and bring them back, and I haven't been  
> able to reliably reproduce it. I will keep trying as I need to get to  
> the bottom of this.
> 
> It has happened once today however, is there any useful log settings I  
> could add? (I'll try setting them all to debug)
> 
> Ian
> 
> On Jul 6, 5:08 pm, Shay Banon \<[shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) ([http://elasticsearch.com](http://elasticsearch.com))\> wrote:
> 
> > If there is a way to recreate it, I would love to try and nail it...
> > 
> > On Wednesday, July 6, 2011 at 12:39 PM, Ian wrote:
> > 
> > > After a period of getting those responses, I restarted the test to  
> > > capture responses to file, and damn it has not happened since. ☹
> > 
> > > This was the first test run after I restarted the nodes one by one  
> > > without issuing any shutdowns, maybe that could cause that same state?  
> > > I will try that again.
> > 
> > > Ian
> > 
> > > On Jul 6, 10:04 am, Clinton Gormley \<[clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) ([http://iannounce.co.uk](http://iannounce.co.uk))\> wrote:
> > > 
> > > > Hi Ian
> > 
> > > > > I am load testing my cluster by sending pseudo-random generated search  
> > > > > requests. I can take the search requests that fail, and replay them  
> > > > > manually and they work. Strange, as this test was working for the past  
> > > > > few days.... I have found exceptions in my logs on only one node, I  
> > > > > have copied them here:-
> > 
> > > > > [search response errors · GitHub](https://gist.github.com/1066838)
> > 
> > > > I see that the string of errors start with this one:
> > 
> > > > Caught exception while handling client http traffic, closing connection  
> > > > java.lang.IllegalStateException: cannot send more responses than  
> > > > requests
> > 
> > > > I see this error a few times a day, and it starts a run of bad  
> > > > responses, eg I search index A and the results contain docs from index A  
> > > > and index B, or the wrong types are returned etc
> > 
> > > > It's like the responses from each node are being mixed up.
> > 
> > > > Do you have any way of replicating this? even infrequently?
> > 
> > > > It would help greatly in the debugging process
> > 
> > > > thanks
> > 
> > > > Clint

---

<div class="post-metadata">

**Author:** ![ian\_clark](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ian_clark/32/3129_2.png) [@ian\_clark](https://discuss.elastic.co/u/ian_clark)\
**Post date:** [July 7, 2011, 4:45pm UTC](https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788/7 "2011-07-07T16:45:34Z")

</div>

I'm just load testing it at the moment, using jmeter. I can seem to  
reproduce it with a trivial 3 node cluster, and if I forcibly kill a  
node whilst throwing lot's of searches at the cluster. That wasn't my  
original scenario, then the cluster was healthy and I was resuming  
tests only, but I get a similar stack trace:-

Caused by: java.lang.IndexOutOfBoundsException: index (0) must be less  
than size (0)  
at  
org.elasticsearch.common.base.Preconditions.checkElementIndex(Preconditions.java:  
301)  
at  
org.elasticsearch.common.base.Preconditions.checkElementIndex(Preconditions.java:  
280)  
at org.elasticsearch.common.collect.Iterables.get(Iterables.java:649)  
at  
org.elasticsearch.search.controller.SearchPhaseController.merge(SearchPhaseController.java:  
259)  
at  
org.elasticsearch.action.search.type.TransportSearchQueryThenFetchAction  
$AsyncAction.innerFinishHim(TransportSearchQueryThenFetchAction.java:  
179)  
at  
org.elasticsearch.action.search.type.TransportSearchQueryThenFetchAction  
$AsyncAction.finishHim(TransportSearchQueryThenFetchAction.java:164)

Could this exception be causing valid search contexts to be released?

[https://github.com/elasticsearch/elasticsearch/blob/master/modules/elasticsearch/src/main/java/org/elasticsearch/action/search/type/TransportSearchQueryThenFetchAction.java#L172](https://github.com/elasticsearch/elasticsearch/blob/master/modules/elasticsearch/src/main/java/org/elasticsearch/action/search/type/TransportSearchQueryThenFetchAction.java#L172)

(A wild guess, I'm not even sure what the search context is!!)

IC

On Jul 7, 3:42 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Ian, nothing specific... . Which client lib to talk to ES are you using?
> 
> On Thursday, July 7, 2011 at 1:59 PM, Ian wrote:
> 
> > Hi Shay,
> 
> > I'm still trying to reproduce this reliably, I repeated the same  
> > steps, process kill the nodes and bring them back, and I haven't been  
> > able to reliably reproduce it. I will keep trying as I need to get to  
> > the bottom of this.
> 
> > It has happened once today however, is there any useful log settings I  
> > could add? (I'll try setting them all to debug)
> 
> > Ian
> 
> > On Jul 6, 5:08 pm, Shay Banon \<[shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) ([http://elasticsearch.com](http://elasticsearch.com))\> wrote:
> > 
> > > If there is a way to recreate it, I would love to try and nail it...
> 
> > > On Wednesday, July 6, 2011 at 12:39 PM, Ian wrote:
> > > 
> > > > After a period of getting those responses, I restarted the test to  
> > > > capture responses to file, and damn it has not happened since. ☹
> 
> > > > This was the first test run after I restarted the nodes one by one  
> > > > without issuing any shutdowns, maybe that could cause that same state?  
> > > > I will try that again.
> 
> > > > Ian
> 
> > > > On Jul 6, 10:04 am, Clinton Gormley \<[clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) ([http://iannounce.co.uk](http://iannounce.co.uk))\> wrote:
> > > > 
> > > > > Hi Ian
> 
> > > > > > I am load testing my cluster by sending pseudo-random generated search  
> > > > > > requests. I can take the search requests that fail, and replay them  
> > > > > > manually and they work. Strange, as this test was working for the past  
> > > > > > few days.... I have found exceptions in my logs on only one node, I  
> > > > > > have copied them here:-
> 
> > > > > > [search response errors · GitHub](https://gist.github.com/1066838)
> 
> > > > > I see that the string of errors start with this one:
> 
> > > > > Caught exception while handling client http traffic, closing connection  
> > > > > java.lang.IllegalStateException: cannot send more responses than  
> > > > > requests
> 
> > > > > I see this error a few times a day, and it starts a run of bad  
> > > > > responses, eg I search index A and the results contain docs from index A  
> > > > > and index B, or the wrong types are returned etc
> 
> > > > > It's like the responses from each node are being mixed up.
> 
> > > > > Do you have any way of replicating this? even infrequently?
> 
> > > > > It would help greatly in the debugging process
> 
> > > > > thanks
> 
> > > > > Clint

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 7, 2011, 5:24pm UTC](https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788/8 "2011-07-07T17:24:51Z")

</div>

This failure has been fixed in master, but it is "guarded" in terms of things being properly released because of that.

On Thursday, July 7, 2011 at 7:45 PM, Ian wrote:

> I'm just load testing it at the moment, using jmeter. I can seem to  
> reproduce it with a trivial 3 node cluster, and if I forcibly kill a  
> node whilst throwing lot's of searches at the cluster. That wasn't my  
> original scenario, then the cluster was healthy and I was resuming  
> tests only, but I get a similar stack trace:-
> 
> Caused by: java.lang.IndexOutOfBoundsException: index (0) must be less  
> than size (0)  
> at  
> org.elasticsearch.common.base.Preconditions.checkElementIndex(Preconditions.java:  
> 301)  
> at  
> org.elasticsearch.common.base.Preconditions.checkElementIndex(Preconditions.java:  
> 280)  
> at org.elasticsearch.common.collect.Iterables.get(Iterables.java:649)  
> at  
> org.elasticsearch.search.controller.SearchPhaseController.merge(SearchPhaseController.java:  
> 259)  
> at  
> org.elasticsearch.action.search.type.TransportSearchQueryThenFetchAction  
> $AsyncAction.innerFinishHim(TransportSearchQueryThenFetchAction.java:  
> 179)  
> at  
> org.elasticsearch.action.search.type.TransportSearchQueryThenFetchAction  
> $AsyncAction.finishHim(TransportSearchQueryThenFetchAction.java:164)
> 
> Could this exception be causing valid search contexts to be released?
> 
> [https://github.com/elasticsearch/elasticsearch/blob/master/modules/elasticsearch/src/main/java/org/elasticsearch/action/search/type/TransportSearchQueryThenFetchAction.java#L172](https://github.com/elasticsearch/elasticsearch/blob/master/modules/elasticsearch/src/main/java/org/elasticsearch/action/search/type/TransportSearchQueryThenFetchAction.java#L172)
> 
> (A wild guess, I'm not even sure what the search context is!!)
> 
> IC
> 
> On Jul 7, 3:42 pm, Shay Banon \<[shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) ([http://elasticsearch.com](http://elasticsearch.com))\> wrote:
> 
> > Ian, nothing specific... . Which client lib to talk to ES are you using?
> > 
> > On Thursday, July 7, 2011 at 1:59 PM, Ian wrote:
> > 
> > > Hi Shay,
> > 
> > > I'm still trying to reproduce this reliably, I repeated the same  
> > > steps, process kill the nodes and bring them back, and I haven't been  
> > > able to reliably reproduce it. I will keep trying as I need to get to  
> > > the bottom of this.
> > 
> > > It has happened once today however, is there any useful log settings I  
> > > could add? (I'll try setting them all to debug)
> > 
> > > Ian
> > 
> > > On Jul 6, 5:08 pm, Shay Banon \<[shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) ([http://elasticsearch.com](http://elasticsearch.com)) ([http://elasticsearch.com](http://elasticsearch.com))\> wrote:
> > > 
> > > > If there is a way to recreate it, I would love to try and nail it...
> > 
> > > > On Wednesday, July 6, 2011 at 12:39 PM, Ian wrote:
> > > > 
> > > > > After a period of getting those responses, I restarted the test to  
> > > > > capture responses to file, and damn it has not happened since. ☹
> > 
> > > > > This was the first test run after I restarted the nodes one by one  
> > > > > without issuing any shutdowns, maybe that could cause that same state?  
> > > > > I will try that again.
> > 
> > > > > Ian
> > 
> > > > > On Jul 6, 10:04 am, Clinton Gormley \<[clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) ([http://iannounce.co.uk](http://iannounce.co.uk)) ([http://iannounce.co.uk](http://iannounce.co.uk))\> wrote:
> > > > > 
> > > > > > Hi Ian
> > 
> > > > > > > I am load testing my cluster by sending pseudo-random generated search  
> > > > > > > requests. I can take the search requests that fail, and replay them  
> > > > > > > manually and they work. Strange, as this test was working for the past  
> > > > > > > few days.... I have found exceptions in my logs on only one node, I  
> > > > > > > have copied them here:-
> > 
> > > > > > > [search response errors · GitHub](https://gist.github.com/1066838)
> > 
> > > > > > I see that the string of errors start with this one:
> > 
> > > > > > Caught exception while handling client http traffic, closing connection  
> > > > > > java.lang.IllegalStateException: cannot send more responses than  
> > > > > > requests
> > 
> > > > > > I see this error a few times a day, and it starts a run of bad  
> > > > > > responses, eg I search index A and the results contain docs from index A  
> > > > > > and index B, or the wrong types are returned etc
> > 
> > > > > > It's like the responses from each node are being mixed up.
> > 
> > > > > > Do you have any way of replicating this? even infrequently?
> > 
> > > > > > It would help greatly in the debugging process
> > 
> > > > > > thanks
> > 
> > > > > > Clint

---

<div class="post-metadata">

**Author:** ![ian\_clark](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ian_clark/32/3129_2.png) [@ian\_clark](https://discuss.elastic.co/u/ian_clark)\
**Post date:** [July 8, 2011, 10:54am UTC](https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788/9 "2011-07-08T10:54:05Z")

</div>

What does "guarded" mean? Shall I test with master or keep trying to  
reliably reproduce using 0.16.2?

IC

On Jul 7, 6:24 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> This failure has been fixed in master, but it is "guarded" in terms of things being properly released because of that.
> 
> On Thursday, July 7, 2011 at 7:45 PM, Ian wrote:
> 
> > I'm just load testing it at the moment, using jmeter. I can seem to  
> > reproduce it with a trivial 3 node cluster, and if I forcibly kill a  
> > node whilst throwing lot's of searches at the cluster. That wasn't my  
> > original scenario, then the cluster was healthy and I was resuming  
> > tests only, but I get a similar stack trace:-
> 
> > Caused by: java.lang.IndexOutOfBoundsException: index (0) must be less  
> > than size (0)  
> > at  
> > org.elasticsearch.common.base.Preconditions.checkElementIndex(Preconditions .java:  
> > 301)  
> > at  
> > org.elasticsearch.common.base.Preconditions.checkElementIndex(Preconditions .java:  
> > 280)  
> > at org.elasticsearch.common.collect.Iterables.get(Iterables.java:649)  
> > at  
> > org.elasticsearch.search.controller.SearchPhaseController.merge(SearchPhase Controller.java:  
> > 259)  
> > at  
> > org.elasticsearch.action.search.type.TransportSearchQueryThenFetchAction  
> > $AsyncAction.innerFinishHim(TransportSearchQueryThenFetchAction.java:  
> > 179)  
> > at  
> > org.elasticsearch.action.search.type.TransportSearchQueryThenFetchAction  
> > $AsyncAction.finishHim(TransportSearchQueryThenFetchAction.java:164)
> 
> > Could this exception be causing valid search contexts to be released?
> 
> > [https://github.com/elasticsearch/elasticsearch/blob/master/modules/el](https://github.com/elasticsearch/elasticsearch/blob/master/modules/el)...
> 
> > (A wild guess, I'm not even sure what the search context is!!)
> 
> > IC
> 
> > On Jul 7, 3:42 pm, Shay Banon \<[shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) ([http://elasticsearch.com](http://elasticsearch.com))\> wrote:
> > 
> > > Ian, nothing specific... . Which client lib to talk to ES are you using?
> 
> > > On Thursday, July 7, 2011 at 1:59 PM, Ian wrote:
> > > 
> > > > Hi Shay,
> 
> > > > I'm still trying to reproduce this reliably, I repeated the same  
> > > > steps, process kill the nodes and bring them back, and I haven't been  
> > > > able to reliably reproduce it. I will keep trying as I need to get to  
> > > > the bottom of this.
> 
> > > > It has happened once today however, is there any useful log settings I  
> > > > could add? (I'll try setting them all to debug)
> 
> > > > Ian
> 
> > > > On Jul 6, 5:08 pm, Shay Banon \<[shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) ([http://elasticsearch.com](http://elasticsearch.com)) ([http://elasticsearch.com](http://elasticsearch.com))\> wrote:
> > > > 
> > > > > If there is a way to recreate it, I would love to try and nail it...
> 
> > > > > On Wednesday, July 6, 2011 at 12:39 PM, Ian wrote:
> > > > > 
> > > > > > After a period of getting those responses, I restarted the test to  
> > > > > > capture responses to file, and damn it has not happened since. ☹
> 
> > > > > > This was the first test run after I restarted the nodes one by one  
> > > > > > without issuing any shutdowns, maybe that could cause that same state?  
> > > > > > I will try that again.
> 
> > > > > > Ian
> 
> > > > > > On Jul 6, 10:04 am, Clinton Gormley \<[clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) ([http://iannounce.co.uk](http://iannounce.co.uk)) ([http://iannounce.co.uk](http://iannounce.co.uk))\> wrote:
> > > > > > 
> > > > > > > Hi Ian
> 
> > > > > > > > I am load testing my cluster by sending pseudo-random generated search  
> > > > > > > > requests. I can take the search requests that fail, and replay them  
> > > > > > > > manually and they work. Strange, as this test was working for the past  
> > > > > > > > few days.... I have found exceptions in my logs on only one node, I  
> > > > > > > > have copied them here:-
> 
> > > > > > > > [search response errors · GitHub](https://gist.github.com/1066838)
> 
> > > > > > > I see that the string of errors start with this one:
> 
> > > > > > > Caught exception while handling client http traffic, closing connection  
> > > > > > > java.lang.IllegalStateException: cannot send more responses than  
> > > > > > > requests
> 
> > > > > > > I see this error a few times a day, and it starts a run of bad  
> > > > > > > responses, eg I search index A and the results contain docs from index A  
> > > > > > > and index B, or the wrong types are returned etc
> 
> > > > > > > It's like the responses from each node are being mixed up.
> 
> > > > > > > Do you have any way of replicating this? even infrequently?
> 
> > > > > > > It would help greatly in the debugging process
> 
> > > > > > > thanks
> 
> > > > > > > Clint

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 8, 2011, 8:46pm UTC](https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788/10 "2011-07-08T20:46:59Z")

</div>

Sorry, by guarded I mean that its properly handled. Either master or 0.16.2 are fine, if you manage to recreate it. I can run the relevant version once you do reproduce it. Thanks!

On Friday, July 8, 2011 at 1:54 PM, Ian wrote:

> What does "guarded" mean? Shall I test with master or keep trying to  
> reliably reproduce using 0.16.2?
> 
> IC
> 
> On Jul 7, 6:24 pm, Shay Banon \<[shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) ([http://elasticsearch.com](http://elasticsearch.com))\> wrote:
> 
> > This failure has been fixed in master, but it is "guarded" in terms of things being properly released because of that.
> > 
> > On Thursday, July 7, 2011 at 7:45 PM, Ian wrote:
> > 
> > > I'm just load testing it at the moment, using jmeter. I can seem to  
> > > reproduce it with a trivial 3 node cluster, and if I forcibly kill a  
> > > node whilst throwing lot's of searches at the cluster. That wasn't my  
> > > original scenario, then the cluster was healthy and I was resuming  
> > > tests only, but I get a similar stack trace:-
> > 
> > > Caused by: java.lang.IndexOutOfBoundsException: index (0) must be less  
> > > than size (0)  
> > > at  
> > > org.elasticsearch.common.base.Preconditions.checkElementIndex(Preconditions .java:  
> > > 301)  
> > > at  
> > > org.elasticsearch.common.base.Preconditions.checkElementIndex(Preconditions .java:  
> > > 280)  
> > > at org.elasticsearch.common.collect.Iterables.get(Iterables.java:649)  
> > > at  
> > > org.elasticsearch.search.controller.SearchPhaseController.merge(SearchPhase Controller.java:  
> > > 259)  
> > > at  
> > > org.elasticsearch.action.search.type.TransportSearchQueryThenFetchAction  
> > > $AsyncAction.innerFinishHim(TransportSearchQueryThenFetchAction.java:  
> > > 179)  
> > > at  
> > > org.elasticsearch.action.search.type.TransportSearchQueryThenFetchAction  
> > > $AsyncAction.finishHim(TransportSearchQueryThenFetchAction.java:164)
> > 
> > > Could this exception be causing valid search contexts to be released?
> > 
> > > [https://github.com/elasticsearch/elasticsearch/blob/master/modules/el](https://github.com/elasticsearch/elasticsearch/blob/master/modules/el)...
> > 
> > > (A wild guess, I'm not even sure what the search context is!!)
> > 
> > > IC
> > 
> > > On Jul 7, 3:42 pm, Shay Banon \<[shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) ([http://elasticsearch.com](http://elasticsearch.com))\> wrote:
> > > 
> > > > Ian, nothing specific... . Which client lib to talk to ES are you using?
> > 
> > > > On Thursday, July 7, 2011 at 1:59 PM, Ian wrote:
> > > > 
> > > > > Hi Shay,
> > 
> > > > > I'm still trying to reproduce this reliably, I repeated the same  
> > > > > steps, process kill the nodes and bring them back, and I haven't been  
> > > > > able to reliably reproduce it. I will keep trying as I need to get to  
> > > > > the bottom of this.
> > 
> > > > > It has happened once today however, is there any useful log settings I  
> > > > > could add? (I'll try setting them all to debug)
> > 
> > > > > Ian
> > 
> > > > > On Jul 6, 5:08 pm, Shay Banon \<[shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) ([http://elasticsearch.com](http://elasticsearch.com)) ([http://elasticsearch.com](http://elasticsearch.com))\> wrote:
> > > > > 
> > > > > > If there is a way to recreate it, I would love to try and nail it...
> > 
> > > > > > On Wednesday, July 6, 2011 at 12:39 PM, Ian wrote:
> > > > > > 
> > > > > > > After a period of getting those responses, I restarted the test to  
> > > > > > > capture responses to file, and damn it has not happened since. ☹
> > 
> > > > > > > This was the first test run after I restarted the nodes one by one  
> > > > > > > without issuing any shutdowns, maybe that could cause that same state?  
> > > > > > > I will try that again.
> > 
> > > > > > > Ian
> > 
> > > > > > > On Jul 6, 10:04 am, Clinton Gormley \<[clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) ([http://iannounce.co.uk](http://iannounce.co.uk)) ([http://iannounce.co.uk](http://iannounce.co.uk))\> wrote:
> > > > > > > 
> > > > > > > > Hi Ian
> > 
> > > > > > > > > I am load testing my cluster by sending pseudo-random generated search  
> > > > > > > > > requests. I can take the search requests that fail, and replay them  
> > > > > > > > > manually and they work. Strange, as this test was working for the past  
> > > > > > > > > few days.... I have found exceptions in my logs on only one node, I  
> > > > > > > > > have copied them here:-
> > 
> > > > > > > > > [search response errors · GitHub](https://gist.github.com/1066838)
> > 
> > > > > > > > I see that the string of errors start with this one:
> > 
> > > > > > > > Caught exception while handling client http traffic, closing connection  
> > > > > > > > java.lang.IllegalStateException: cannot send more responses than  
> > > > > > > > requests
> > 
> > > > > > > > I see this error a few times a day, and it starts a run of bad  
> > > > > > > > responses, eg I search index A and the results contain docs from index A  
> > > > > > > > and index B, or the wrong types are returned etc
> > 
> > > > > > > > It's like the responses from each node are being mixed up.
> > 
> > > > > > > > Do you have any way of replicating this? even infrequently?
> > 
> > > > > > > > It would help greatly in the debugging process
> > 
> > > > > > > > thanks
> > 
> > > > > > > > Clint

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:01am UTC](https://discuss.elastic.co/t/random-search-errors-from-one-node-of-three-on-a-previously-working-cluster/4788/11 "2017-07-06T04:01:15Z")

</div>


