# Shard failures

**URL:** <https://discuss.elastic.co/t/shard-failures/17184>\
**Category:** Elasticsearch\
**Created:** [April 24, 2014, 7:35pm UTC](https://discuss.elastic.co/t/shard-failures/17184 "2014-04-24T19:35:17Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Drew\_Blessing](https://avatars.discourse-cdn.com/v4/letter/d/838e76/32.png) [@Drew\_Blessing](https://discuss.elastic.co/u/Drew_Blessing)\
**Post date:** [April 24, 2014, 7:35pm UTC](https://discuss.elastic.co/t/shard-failures/17184/1 "2014-04-24T19:35:17Z")

</div>

We have been struggling with this issue for a few months. We've experienced  
it in versions 0.90.6 - 0.90.13 and now in 1.1, too.

A shard (sometimes 2) will fail within a single index. We get this error  
during or after our data loader indexes data. Sometimes it takes a day or  
two to occur but most recently it's been immediately on/after first index.  
The shard that fails is always in the same index. This is a 2-node cluster  
running on CentOS 6.5 with Oracle Java 1.7.0u51. In addition, there is 1  
non-data node for the process that handles indexing data and 2 non-data  
nodes serving the front-end. All non-data nodes are java clients using  
spring-data-elasticsearch library. All nodes are on 1.1 now. We understand  
there's a probability that our loader application is causing this but we  
can't see how or where. Also, it seems like a bug if a client can cause  
shards to fail on the server. We are grasping at straws now and appreciate  
any ideas on what could be causing this.

In this gist, the first log message is from node 1 and happens at the same  
time that the shard failure occurs on node 2. See node 2 for the stack(s)  
that occur when the shard fails. It's interesting that node 1 says it's  
closing the connection because of it. Someone on Twitter noted that these  
are WARN level messages and don't signify a failure. However, it is causing  
queries against this index to totally fail, so there's definitely more than  
a WARN scenario going on here. Any thoughts?

> <https://gist.github.com/dblessing/11266650>

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/371aa3d3-1b02-4fb5-bad7-b6217e09fb6a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/371aa3d3-1b02-4fb5-bad7-b6217e09fb6a%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [April 26, 2014, 12:32am UTC](https://discuss.elastic.co/t/shard-failures/17184/2 "2014-04-26T00:32:53Z")

</div>

Hey,

do you run any additional plugins or is this stock elasticsearch? Can you  
tell what happens on your cluster? Do you have long running  
queries/operations? Can you tell more, what/how this loader executes?

Minor operational hint: Upgrade your JVM version to the latest one or \_25,  
the one your are using could lead to data corruption with lucene.

--Alex

On Thu, Apr 24, 2014 at 3:35 PM, Drew Blessing [blessing.drew@gmail.com](mailto:blessing.drew@gmail.com)wrote:

> We have been struggling with this issue for a few months. We've  
> experienced it in versions 0.90.6 - 0.90.13 and now in 1.1, too.
> 
> A shard (sometimes 2) will fail within a single index. We get this error  
> during or after our data loader indexes data. Sometimes it takes a day or  
> two to occur but most recently it's been immediately on/after first index.  
> The shard that fails is always in the same index. This is a 2-node cluster  
> running on CentOS 6.5 with Oracle Java 1.7.0u51. In addition, there is 1  
> non-data node for the process that handles indexing data and 2 non-data  
> nodes serving the front-end. All non-data nodes are java clients using  
> spring-data-elasticsearch library. All nodes are on 1.1 now. We understand  
> there's a probability that our loader application is causing this but we  
> can't see how or where. Also, it seems like a bug if a client can cause  
> shards to fail on the server. We are grasping at straws now and appreciate  
> any ideas on what could be causing this.
> 
> In this gist, the first log message is from node 1 and happens at the same  
> time that the shard failure occurs on node 2. See node 2 for the stack(s)  
> that occur when the shard fails. It's interesting that node 1 says it's  
> closing the connection because of it. Someone on Twitter noted that these  
> are WARN level messages and don't signify a failure. However, it is causing  
> queries against this index to totally fail, so there's definitely more than  
> a WARN scenario going on here. Any thoughts?
> 
> [Shard failure · GitHub](https://gist.github.com/dblessing/11266650)
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/371aa3d3-1b02-4fb5-bad7-b6217e09fb6a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/371aa3d3-1b02-4fb5-bad7-b6217e09fb6a%40googlegroups.com)[https://groups.google.com/d/msgid/elasticsearch/371aa3d3-1b02-4fb5-bad7-b6217e09fb6a%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/371aa3d3-1b02-4fb5-bad7-b6217e09fb6a%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAGCwEM8E7spY%2BOU9PHfMBuLujqCZeCLXPmM4fdDasMvGyXnAgg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAGCwEM8E7spY%2BOU9PHfMBuLujqCZeCLXPmM4fdDasMvGyXnAgg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:33am UTC](https://discuss.elastic.co/t/shard-failures/17184/3 "2017-07-06T01:33:21Z")

</div>


