# Elasticsearch cluster in Yellow state and 1 Unassigned Shard

**URL:** https://discuss.elastic.co/t/elasticsearch-cluster-in-yellow-state-and-1-unassigned-shard/247290
**Category:** Elasticsearch
**Created:** [September 2, 2020, 9:26pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-in-yellow-state-and-1-unassigned-shard/247290 "2020-09-02T21:26:19Z")
**Posts on this page:** 8
**Page:** 2

<div class="post-metadata">

### Author: ![zaeemmasood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zaeemmasood/32/102383_2.png) [@zaeemmasood](https://discuss.elastic.co/u/zaeemmasood)
#### Post date: [September 3, 2020, 2:13pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-in-yellow-state-and-1-unassigned-shard/247290/21 "2020-09-03T14:13:21Z")

</div>

Executed reroute on today's index which was showing Yellow:

`POST _cluster/reroute?retry_failed=true`

On running GET (command below) on recovery process, I could see recovery taking place in terms of percentage.

`GET _cat/recovery/prod_test_access-2020.09.03?v`

After it completed 100%, health still shows as Yellow with one unassigned shard for the same (today's) index.

Shall I drop the replica for today's index and readd?

---

<div class="post-metadata">

### Author: ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)
#### Post date: [September 3, 2020, 9:18pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-in-yellow-state-and-1-unassigned-shard/247290/22 "2020-09-03T21:18:58Z")

</div>

Why is it showing as not allocated?

---

<div class="post-metadata">

### Author: ![zaeemmasood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zaeemmasood/32/102383_2.png) [@zaeemmasood](https://discuss.elastic.co/u/zaeemmasood)
#### Post date: [September 3, 2020, 11:30pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-in-yellow-state-and-1-unassigned-shard/247290/23 "2020-09-03T23:30:42Z")

</div>

Sorry. Where do you see it unallocated?

---

<div class="post-metadata">

### Author: ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)
#### Post date: [September 3, 2020, 11:33pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-in-yellow-state-and-1-unassigned-shard/247290/24 "2020-09-03T23:33:09Z")

</div>

What does a `_cluster/allocation/explain` say on why it's unallocated?

---

<div class="post-metadata">

### Author: ![zaeemmasood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zaeemmasood/32/102383_2.png) [@zaeemmasood](https://discuss.elastic.co/u/zaeemmasood)
#### Post date: [September 4, 2020, 12:04am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-in-yellow-state-and-1-unassigned-shard/247290/25 "2020-09-04T00:04:13Z")

</div>

Here is the output of allocation/explain:

```auto
{
  "index" : "prod_test_access-2020.09.03",
  "shard" : 0,
  "primary" : false,
  "current_state" : "unassigned",
  "unassigned_info" : {
    "reason" : "ALLOCATION_FAILED",
    "at" : "2020-09-03T14:03:38.063Z",
    "failed_allocation_attempts" : 5,
    "details" : "failed shard on node [2ULg0RFsSZiaUYm4SUpDOQ]: failed to perform indices:data/write/bulk[s] on replica [prod_test_access-2020.09.03][0], node[2ULg0RFsSZiaUYm4SUpDOQ], [R], recovery_source[peer recovery], s[INITIALIZING], a[id=LT4uKYoPSBCYiUIm4xCibA], unassigned_info[[reason=ALLOCATION_FAILED], at[2020-09-03T14:03:06.722Z], failed_attempts[4], failed_nodes[[2ULg0RFsSZiaUYm4SUpDOQ, y1s252NOR5amHRRjWpu5jA]], delayed=false, details[failed shard on node [y1s252NOR5amHRRjWpu5jA]: failed to perform indices:data/write/bulk[s] on replica [prod_test_access-2020.09.03][0], node[y1s252NOR5amHRRjWpu5jA], [R], recovery_source[peer recovery], s[INITIALIZING], a[id=6HZ6qmCTRsecCYZudceSDg], unassigned_info[[reason=ALLOCATION_FAILED], at[2020-09-03T14:02:21.461Z], failed_attempts[3], failed_nodes[[2ULg0RFsSZiaUYm4SUpDOQ, y1s252NOR5amHRRjWpu5jA]], delayed=false, details[failed shard on node [2ULg0RFsSZiaUYm4SUpDOQ]: failed to perform indices:data/write/bulk[s] on replica [prod_test_access-2020.09.03][0], node[2ULg0RFsSZiaUYm4SUpDOQ], [R], recovery_source[peer recovery], s[INITIALIZING], a[id=MNGK3GTQR2CWU2XtTm5NnQ], unassigned_info[[reason=ALLOCATION_FAILED], at[2020-09-03T14:01:01.986Z], failed_attempts[2], failed_nodes[[2ULg0RFsSZiaUYm4SUpDOQ, y1s252NOR5amHRRjWpu5jA]], delayed=false, details[failed shard on node [y1s252NOR5amHRRjWpu5jA]: failed recovery, failure RecoveryFailedException[[prod_test_access-2020.09.03][0]: Recovery failed from {xxx-xxx-xxx}{acF0z4ZvQNy0lE5fykyf_Q}{0Jl-faJ_R26v8dlzzJ1OaQ}{xxx-xxx-xxx}{123-345-567:9300}{dil}{ml.machine_memory=67378692096, ml.max_open_jobs=20, xpack.installed=true} into {xxx-xxx-xxx}{y1s252NOR5amHRRjWpu5jA}{Sw0AaQSwQR6qJEVc5mHaqA}{xxx-xxx-xxx}{123-345-567:9300}{dil}{ml.machine_memory=67378692096, xpack.installed=true, ml.max_open_jobs=20}]; nested: RemoteTransportException[[xxx-xxx-xxx][123-345-567:9300][internal:index/shard/recovery/start_recovery]]; nested: RemoteTransportException[[xxx-xxx-xxx][123-345-567:9300][internal:index/shard/recovery/translog_ops]]; nested: CircuitBreakingException[[parent] Data too large, data for [<transport_request>] would be [993359896/947.3mb], which is larger than the limit of [986061209/940.3mb], real usage: [992302552/946.3mb], new bytes reserved: [1057344/1mb], usages [request=0/0b, fielddata=18581/18.1kb, in_flight_requests=1057344/1mb, accounting=210380623/200.6mb]]; ], allocation_status[no_attempt]], expected_shard_size[40551605435], failure RemoteTransportException[[xxx-xxx-xxx][123-345-567:9300][indices:data/write/bulk[s][r]]]; nested: CircuitBreakingException[[parent] Data too large, data for [<transport_request>] would be [991720182/945.7mb], which is larger than the limit of [986061209/940.3mb], real usage: [991595760/945.6mb], new bytes reserved: [124422/121.5kb], usages [request=0/0b, fielddata=16859/16.4kb, in_flight_requests=1495582/1.4mb, accounting=225438133/214.9mb]]; ], allocation_status[no_attempt]], expected_shard_size[34150060094], failure RemoteTransportException[[xxx-xxx-xxx][123-345-567:9300][indices:data/write/bulk[s][r]]]; nested: CircuitBreakingException[[parent] Data too large, data for [<transport_request>] would be [989038996/943.2mb], which is larger than the limit of [986061209/940.3mb], real usage: [988874824/943mb], new bytes reserved: [164172/160.3kb], usages [request=0/0b, fielddata=18581/18.1kb, in_flight_requests=1219938/1.1mb, accounting=205318354/195.8mb]]; ], allocation_status[no_attempt]], expected_shard_size[33902587303], failure RemoteTransportException[[xxx-xxx-xxx][123-345-567:9300][indices:data/write/bulk[s][r]]]; nested: CircuitBreakingException[[parent] Data too large, data for [<transport_request>] would be [986870246/941.1mb], which is larger than the limit of [986061209/940.3mb], real usage: [986791752/941mb], new bytes reserved: [78494/76.6kb], usages [request=0/0b, fielddata=16859/16.4kb, in_flight_requests=1137206/1mb, accounting=225693585/215.2mb]]; ",
    "last_allocation_status" : "no_attempt"
  },
  "can_allocate" : "no",
  "allocate_explanation" : "cannot allocate because allocation is not permitted to any of the nodes",
  "node_allocation_decisions" : [
    {
      "node_id" : "2ULg0RFsSZiaUYm4SUpDOQ",
      "node_name" : "xxx-xxx-xxx",
      "transport_address" : "123-345-567:9300",
      "node_attributes" : {
        "ml.machine_memory" : "67378692096",
        "ml.max_open_jobs" : "20",
        "xpack.installed" : "true"
      },
      "node_decision" : "no",
      "deciders" : [
        {
          "decider" : "max_retry",
          "decision" : "NO",
          "explanation" : "shard has exceeded the maximum number of retries [5] on failed allocation attempts - manually call [/_cluster/reroute?retry_failed=true] to retry, [unassigned_info[[reason=ALLOCATION_FAILED], at[2020-09-03T14:03:38.063Z], failed_attempts[5], failed_nodes[[2ULg0RFsSZiaUYm4SUpDOQ, y1s252NOR5amHRRjWpu5jA]], delayed=false, details[failed shard on node [2ULg0RFsSZiaUYm4SUpDOQ]: failed to perform indices:data/write/bulk[s] on replica [prod_test_access-2020.09.03][0], node[2ULg0RFsSZiaUYm4SUpDOQ], [R], recovery_source[peer recovery], s[INITIALIZING], a[id=LT4uKYoPSBCYiUIm4xCibA], unassigned_info[[reason=ALLOCATION_FAILED], at[2020-09-03T14:03:06.722Z], failed_attempts[4], failed_nodes[[2ULg0RFsSZiaUYm4SUpDOQ, y1s252NOR5amHRRjWpu5jA]], delayed=false, details[failed shard on node [y1s252NOR5amHRRjWpu5jA]: failed to perform indices:data/write/bulk[s] on replica [prod_test_access-2020.09.03][0], node[y1s252NOR5amHRRjWpu5jA], [R], recovery_source[peer recovery], s[INITIALIZING], a[id=6HZ6qmCTRsecCYZudceSDg], unassigned_info[[reason=ALLOCATION_FAILED], at[2020-09-03T14:02:21.461Z], failed_attempts[3], failed_nodes[[2ULg0RFsSZiaUYm4SUpDOQ, y1s252NOR5amHRRjWpu5jA]], delayed=false, details[failed shard on node [2ULg0RFsSZiaUYm4SUpDOQ]: failed to perform indices:data/write/bulk[s] on replica [prod_test_access-2020.09.03][0], node[2ULg0RFsSZiaUYm4SUpDOQ], [R], recovery_source[peer recovery], s[INITIALIZING], a[id=MNGK3GTQR2CWU2XtTm5NnQ], unassigned_info[[reason=ALLOCATION_FAILED], at[2020-09-03T14:01:01.986Z], failed_attempts[2], failed_nodes[[2ULg0RFsSZiaUYm4SUpDOQ, y1s252NOR5amHRRjWpu5jA]], delayed=false, details[failed shard on node [y1s252NOR5amHRRjWpu5jA]: failed recovery, failure RecoveryFailedException[[prod_test_access-2020.09.03][0]: Recovery failed from {xxx-xxx-xxx}{acF0z4ZvQNy0lE5fykyf_Q}{0Jl-faJ_R26v8dlzzJ1OaQ}{xxx-xxx-xxx}{123-345-567:9300}{dil}{ml.machine_memory=67378692096, ml.max_open_jobs=20, xpack.installed=true} into {xxx-xxx-xxx}{y1s252NOR5amHRRjWpu5jA}{Sw0AaQSwQR6qJEVc5mHaqA}{xxx-xxx-xxx}{123-345-567:9300}{dil}{ml.machine_memory=67378692096, xpack.installed=true, ml.max_open_jobs=20}]; nested: RemoteTransportException[[xxx-xxx-xxx][123-345-567:9300][internal:index/shard/recovery/start_recovery]]; nested: RemoteTransportException[[xxx-xxx-xxx][123-345-567:9300][internal:index/shard/recovery/translog_ops]]; nested: CircuitBreakingException[[parent] Data too large, data for [<transport_request>] would be [993359896/947.3mb], which is larger than the limit of [986061209/940.3mb], real usage: [992302552/946.3mb], new bytes reserved: [1057344/1mb], usages [request=0/0b, fielddata=18581/18.1kb, in_flight_requests=1057344/1mb, accounting=210380623/200.6mb]]; ], allocation_status[no_attempt]], expected_shard_size[40551605435], failure RemoteTransportException[[xxx-xxx-xxx][123-345-567:9300][indices:data/write/bulk[s][r]]]; nested: CircuitBreakingException[[parent] Data too large, data for [<transport_request>] would be [991720182/945.7mb], which is larger than the limit of [986061209/940.3mb], real usage: [991595760/945.6mb], new bytes reserved: [124422/121.5kb], usages [request=0/0b, fielddata=16859/16.4kb, in_flight_requests=1495582/1.4mb, accounting=225438133/214.9mb]]; ], allocation_status[no_attempt]], expected_shard_size[34150060094], failure RemoteTransportException[[xxx-xxx-xxx][123-345-567:9300][indices:data/write/bulk[s][r]]]; nested: CircuitBreakingException[[parent] Data too large, data for [<transport_request>] would be [989038996/943.2mb], which is larger than the limit of [986061209/940.3mb], real usage: [988874824/943mb], new bytes reserved: [164172/160.3kb], usages [request=0/0b, fielddata=18581/18.1kb, in_flight_requests=1219938/1.1mb, accounting=205318354/195.8mb]]; ], allocation_status[no_attempt]], expected_shard_size[33902587303], failure RemoteTransportException[[xxx-xxx-xxx][123-345-567:9300][indices:data/write/bulk[s][r]]]; nested: CircuitBreakingException[[parent] Data too large, data for [<transport_request>] would be [986870246/941.1mb], which is larger than the limit of [986061209/940.3mb], real usage: [986791752/941mb], new bytes reserved: [78494/76.6kb], usages [request=0/0b, fielddata=16859/16.4kb, in_flight_requests=1137206/1mb, accounting=225693585/215.2mb]]; ], allocation_status[no_attempt]]]"
        }
      ]
    }

```

There is more but it is same as above.

Primarily `CircuitBreakingException[[parent] Data too large`

---

<div class="post-metadata">

### Author: ![zaeemmasood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zaeemmasood/32/102383_2.png) [@zaeemmasood](https://discuss.elastic.co/u/zaeemmasood)
#### Post date: [September 4, 2020, 1:13pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-in-yellow-state-and-1-unassigned-shard/247290/26 "2020-09-04T13:13:20Z")

</div>

Today the unassigned number has risen to 4. See below

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/4/0/4030aa42b1843394d3c1fb73a80042b7518d1ca8.png)

```auto
{
  "cluster_name" : "elkcluster-prod",
  "status" : "yellow",
  "timed_out" : false,
  "number_of_nodes" : 8,
  "number_of_data_nodes" : 5,
  "active_primary_shards" : 195,
  "active_shards" : 386,
  "relocating_shards" : 0,
  "initializing_shards" : 0,
  "unassigned_shards" : 4,
  "delayed_unassigned_shards" : 0,
  "number_of_pending_tasks" : 0,
  "number_of_in_flight_fetch" : 0,
  "task_max_waiting_in_queue_millis" : 0,
  "active_shards_percent_as_number" : 98.97435897435898
}

```

---

<div class="post-metadata">

### Author: ![zaeemmasood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zaeemmasood/32/102383_2.png) [@zaeemmasood](https://discuss.elastic.co/u/zaeemmasood)
#### Post date: [September 4, 2020, 7:00pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-in-yellow-state-and-1-unassigned-shard/247290/27 "2020-09-04T19:00:13Z")

</div>

Hi All,

I think I know where the issue is. "CircuitBreakingException[[parent] Data too large" error lead me to issues related to memory. I increased the heap size to just less than 50% of the physical RAM as advised in [https://www.elastic.co/guide/en/elasticsearch/reference/current/heap-size.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/heap-size.html)

Did `_cluster/reroute?retry_failed=true` one more time and the health is now back to Green with no unassigned shard.

Thanks @Warkolm for the useful tips as I got to learn new things in the process!

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [October 2, 2020, 7:00pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-in-yellow-state-and-1-unassigned-shard/247290/28 "2020-10-02T19:00:15Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.

[Previous page](https://discuss.elastic.co/t/elasticsearch-cluster-in-yellow-state-and-1-unassigned-shard/247290.md?page=1)
