# Large amounts of shards failing

**URL:** <https://discuss.elastic.co/t/large-amounts-of-shards-failing/27216>\
**Category:** Elasticsearch\
**Created:** [August 11, 2015, 7:12pm UTC](https://discuss.elastic.co/t/large-amounts-of-shards-failing/27216 "2015-08-11T19:12:41Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![DC64](https://avatars.discourse-cdn.com/v4/letter/d/919ad9/32.png) [@DC64](https://discuss.elastic.co/u/DC64)\
**Post date:** [August 11, 2015, 7:12pm UTC](https://discuss.elastic.co/t/large-amounts-of-shards-failing/27216/1 "2015-08-11T19:12:41Z")

</div>

Earlier today, for the first time, I have been receiving errors in Kibana saying that x/185 shards have failed. Depending on how large my date filter is (last 7 days or last hour) the number and total shards will change.

I am running ES 1.4 and Kibana 4.1

I can run `_cat/recovery/` and get something like this: `observables-2015.07.16 3 26 gateway done CIF-2 CPU2 n/a n/a 73 100.0% 54557248 100.0%`

I can still visualize my data in my dashboard, but the errors are concerning. What should I do next?

Update: Some visualizations have less data showing up than before.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 11, 2015, 11:37pm UTC](https://discuss.elastic.co/t/large-amounts-of-shards-failing/27216/2 "2015-08-11T23:37:06Z")

</div>

Something is up in ES.

Have a look at some of the other cat endpoints, like allocation, to see what is happening.

---

<div class="post-metadata">

**Author:** ![DC64](https://avatars.discourse-cdn.com/v4/letter/d/919ad9/32.png) [@DC64](https://discuss.elastic.co/u/DC64)\
**Post date:** [August 12, 2015, 3:50pm UTC](https://discuss.elastic.co/t/large-amounts-of-shards-failing/27216/3 "2015-08-12T15:50:09Z")

</div>

(Maybe this belongs in Elasticsearch section, sorry...)

Running `curl -XGET http://localhost:9200/_cluster/health?pretty` I get:  
`{ "cluster_name" : "elasticsearch", "status" : "yellow", "timed_out" : false, "number_of_nodes" : 1, "number_of_data_nodes" : 1, "active_primary_shards" : 201, "active_shards" : 201, "relocating_shards" : 0, "initializing_shards" : 0, "unassigned_shards" : 201 }`  
When I run `curl -XGET http://localhost:9200/_cat/shards` I get  
`observables-2015.08.01 4 p STARTED 55421 49.4mb 127.0.1.1 Onyxx observables-2015.08.01 4 r UNASSIGNED` and this list repeats until all shards are read thrugh. It looks like they are being assigned incorrectly, or not at all. The "4" changes from 0, 1,2, 3, 4, in any order, not sure why this is.

Running `curl -XGET 'http://localhost:9200/_cat/allocation'` I get:  
`201 16.3gb 22.2gb 38.6gb 42 CPU2 127.0.1.1 Onyxx 201 UNASSIGNED`  
Running `_shards` with `| grep UNASSIGNED` I see nearly all of my shards as unassigned, as expected.

I have hourly and daily pieces of data that are read, so my logs and shards are in big numbers.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 12, 2015, 9:31pm UTC](https://discuss.elastic.co/t/large-amounts-of-shards-failing/27216/4 "2015-08-12T21:31:40Z")

</div>

Don't worry about the yellow, those shards won't be assigned as you only have one node. You can get rid of them with `curl -XPUT localhost:9200/*/_settings -d '{ "index" : { "number_of_replicas" : 0 } }'` to get your cluster to green.

What's happening in your ES logs?

(I'll also move this to the ES area)

---

<div class="post-metadata">

**Author:** ![DC64](https://avatars.discourse-cdn.com/v4/letter/d/919ad9/32.png) [@DC64](https://discuss.elastic.co/u/DC64)\
**Post date:** [August 13, 2015, 3:01pm UTC](https://discuss.elastic.co/t/large-amounts-of-shards-failing/27216/5 "2015-08-13T15:01:32Z")

</div>

Looking through my logs, elasticsearch.log.2015-08-10 looks fine with 6 lines of logs, but in elasticsearch.log.2015-08-11, I find in a few hundred lines of errors that `java.lang.OutOfMemoryError: Java heap space`, looks like a memory issue with java. This is running in a VM server with 16 gigs of memory, and 11 gigs of shards, though the hourly downloads of data could be making java upset.

I also get a common error: `[FIELDDATA] Data too large, data for [timezone] would be larger than limit of [633785548/604.4mb]`

What I don't understand is that my shards (over 200) from a month ago are also listed as unassigned. Ideally I should be able to look at my data from a week ago just fine, but that's not the case.

---

<div class="post-metadata">

**Author:** ![DC64](https://avatars.discourse-cdn.com/v4/letter/d/919ad9/32.png) [@DC64](https://discuss.elastic.co/u/DC64)\
**Post date:** [August 17, 2015, 2:38pm UTC](https://discuss.elastic.co/t/large-amounts-of-shards-failing/27216/6 "2015-08-17T14:38:55Z")

</div>

I found out the reality of how Elasticsearch _really_ loves RAM. I had to move some shards into another folder so that some memory is free'd up. Until I get more than one server, moving shards around manually is the current option.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:55pm UTC](https://discuss.elastic.co/t/large-amounts-of-shards-failing/27216/7 "2017-07-05T23:55:30Z")

</div>


