# Courier Fetch: X of Y shards failed

**URL:** <https://discuss.elastic.co/t/courier-fetch-x-of-y-shards-failed/36655>\
**Category:** Elasticsearch\
**Created:** [December 8, 2015, 4:41pm UTC](https://discuss.elastic.co/t/courier-fetch-x-of-y-shards-failed/36655 "2015-12-08T16:41:49Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![cburkins](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cburkins/32/4703_2.png) [@cburkins](https://discuss.elastic.co/u/cburkins)\
**Post date:** [December 8, 2015, 4:41pm UTC](https://discuss.elastic.co/t/courier-fetch-x-of-y-shards-failed/36655/1 "2015-12-08T16:41:49Z")

</div>

Back when I was using elasticsearch 1.5, I experienced the following error in Kibana 4.2:

Courier Fetch: X of Y shards failed

To compensate for this, I inserted the following line into /etc/elasticsearch/elasticsearch.yml

```
# Allows for unbounded queue for reads
threadpool.search.type: cached

```

I don't thoroughly understand the setting, but my impression is that it forces the elasticsearch query (from Kibana) to wait for a response, rather than timing out. Performance isn't critical to me, so that's just fine.

I recently upgraded to elasticsearch 2.1. I received the same "Courier Fetch: X of Y shards failed", so I tried to use the same setting. Unfortunately, I got an error this time. Has this setting been removed from elasticsearch ?

-Chad

---

<div class="post-metadata">

**Author:** ![cburkins](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cburkins/32/4703_2.png) [@cburkins](https://discuss.elastic.co/u/cburkins)\
**Post date:** [December 8, 2015, 4:58pm UTC](https://discuss.elastic.co/t/courier-fetch-x-of-y-shards-failed/36655/2 "2015-12-08T16:58:11Z")

</div>

The error that I receive is:

> setting threadpool.search.type to cached is not permitted; must be fixed

---

<div class="post-metadata">

**Author:** ![cburkins](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cburkins/32/4703_2.png) [@cburkins](https://discuss.elastic.co/u/cburkins)\
**Post date:** [December 8, 2015, 7:49pm UTC](https://discuss.elastic.co/t/courier-fetch-x-of-y-shards-failed/36655/3 "2015-12-08T19:49:34Z")

</div>

I looked in /var/log/elasticsearch/.log

> Caused by: EsRejectedExecutionException[rejected execution of org.elasticsearch.transport.TransportService$4@100f9843 on EsThreadPoolExecutor[search, queue capacity = 1000, org.elasticsearch.common.util.concurrent.EsThreadPoolExecutor@3a3dd1dd[Running, pool size = 4, active threads = 4, queued tasks = 1000, completed tasks = 9343]]]  
> RemoteTransportException[[Robert Bruce Banner][127.0.0.1:9300][indices:data/read/search[phase/query]]]; nested: EsRejectedExecutionException[rejected execution of org.elasticsearch.transport.TransportService$4@4ce31bbe on EsThreadPoolExecutor[search, queue capacity = 1000, org.elasticsearch.common.util.concurrent.EsThreadPoolExecutor@3a3dd1dd[Running, pool size = 4, active threads = 4, queued tasks = 1000, completed tasks = 9343]]];

The key part seems to be ?

> search, queue capacity = 1000

So, took a guess, and this _seems_ to solve my problem

I added the following to /etc/elasticsearch/elasticsearch.yml:

> threadpool.search.queue\_size: 2000

I'm not entirely sure what this does.....''

---

<div class="post-metadata">

**Author:** ![jasontedor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jasontedor/32/66992_2.png) [@jasontedor](https://discuss.elastic.co/u/jasontedor)\
**Post date:** [December 9, 2015, 1:51am UTC](https://discuss.elastic.co/t/courier-fetch-x-of-y-shards-failed/36655/4 "2015-12-09T01:51:44Z")

</div>

> [@cburkins](#):
>
> To compensate for this, I inserted the following line into /etc/elasticsearch/elasticsearch.yml
> 
> `# Allows for unbounded queue for reads threadpool.search.type: cached`

First, this is addressing the symptom, not the underlying problem. It's important to note here that your cluster was under duress and reporting `EsRejectedExecutionException`s; buried in those exception messages were messages that Elasticsearch's search queue was stuffed. What you did is change the thread pool type for the search thread pool from `fixed` (a fixed number of workers, bounded work queue) to `cached` (an unbounded number of works, unbounded work queue). The thread pool types are covered in the [documentation](https://www.elastic.co/guide/en/elasticsearch/reference/current/modules-threadpool.html).

This change is _incredibly_ dangerous. If your node is having trouble keeping up with work, allowing more work in is not the solution. The `EsRejectedExecutionException` is a backpressure mechanism desperately trying to signal to the clients: stop, I can not keep up! The fervent hope would be the clients would apply some kind of backoff-retry mechanism until the work queue drains. Clients could either exponentially backoff-retry, or introspect Elasticsearch and check the number of tasks in the work queue before sending the rejected request again.

In fact, the `cached` thread pool type is so dangerous that in Elasticsearch 2.1.0, Elasticsearch stopped allowing thread pools to be set to type `cached`. That is why you saw this message:

> [@cburkins](#):
>
> `setting threadpool.search.type to cached is not permitted; must be fixed`

An unbounded thread pool is reserved for extremely special circumstances, namely requests that absolutely must be served immediately lest Elasticsearch blocks.

Well, that's not quite truthful. Elasticsearch as of 2.1.0 actually completely forbids changing thread pool types at all. The reasoning for this is because changing the thread pool type is very risky with extremely little real-world benefit; it's was deemed not worth the cost to users and the complexity to Elasticsearch to enable changing thread pool types.

> [@cburkins](#):
>
> the key part seems to be ?
> 
> `search, queue capacity = 1000`
> 
> So, took a guess, and this seems to solve my problem
> 
> I added the following to /etc/elasticsearch/elasticsearch.yml:
> 
> `threadpool.search.queue_size: 2000`

Again, you have addressed the symptom and not the problem. It is possible that your workload constantly hovers around 2000 search tasks. But maybe something else is going on? Maybe your cluster is undersized? Maybe your nodes are undersized? Maybe something is wrong with your clients and they sending too many requests?

> [@cburkins](#):
>
> I'm not entirely sure what this does.....'

Increasing the queue size without discerning the cause of the stuffed queue, and without considering the effects of increasing the thread pool could just be postponing the day of reckoning. You most definitely do not want to keep increasing the queue size without finding out why your queues are so constantly stuffed, and whether or not your cluster can handle the workload that it is currently under.

I hope that helps. Happy to help more if the need arises. 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:32pm UTC](https://discuss.elastic.co/t/courier-fetch-x-of-y-shards-failed/36655/6 "2017-07-05T22:32:14Z")

</div>


