# Getting occasional unassigned replica shards

**URL:** <https://discuss.elastic.co/t/getting-occasional-unassigned-replica-shards/14178>\
**Category:** Elasticsearch\
**Created:** [October 30, 2013, 10:14pm UTC](https://discuss.elastic.co/t/getting-occasional-unassigned-replica-shards/14178 "2013-10-30T22:14:07Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mike\_Turner](https://avatars.discourse-cdn.com/v4/letter/m/ba8739/32.png) [@Mike\_Turner](https://discuss.elastic.co/u/Mike_Turner)\
**Post date:** [October 30, 2013, 10:14pm UTC](https://discuss.elastic.co/t/getting-occasional-unassigned-replica-shards/14178/1 "2013-10-30T22:14:07Z")

</div>

Hi there,

I'm running a cluster that's fairly busy, and after doing a full cluster  
restart and a bit of tuning to the configs, I'm getting an occasional  
unassigned replica shard that never seems to get assigned. We generate a  
new index daily, and last night I saw for the first time a new index get  
created with a missing replica shard and show up as unassigned. I've since  
deleted that index and let it recreate from inbound data, and it created a  
second time the same way, with one replica unassigned.

So really, two issues here:

1. On a cluster restart, four indices came up with an unassigned replica  
shard, and the cluster did not self-correct that condition.
2. On a new index being generated, one replica shard is missing,  
repeatable 2x, and the cluster did not self-correct that condition.

I found a solution with the help of a user in the #elasticsearch channel  
that showed me how to reindex the older indices to get the missing replica  
shard back (using the Tire library's index.reindex method in combination  
with an alias) and I am in the process of reindexing the four older indices  
that had missing replicas. That process will take more than a week (~2.5  
days per index) with the size of things and level of IO activity we  
presently have.

I really want to get to the bottom of why this happened, and even more  
importantly, why a new index would get created without all of the required  
shards.

We haven't seen this happen before, so I suspect that it's a product of the  
tuning that I recently did. Here's what changed:

1. Doubled the size of the cluster from 3 nodes to 6.
2. Increased the shard and replica count from 3 shards to 6, and from 1  
replica to 2 at the same time.
3. "index.routing.allocation.total\_shards\_per\_node" : 3
4. discovery.zen.minimum\_master\_nodes: 4
5. gateway.recover\_after\_nodes: 4
6. gateway.recover\_after\_time: 10m
7. gateway.expected\_data\_nodes: 2
8. gateway.expected\_master\_nodes: 6

Is there anything in those settings that stand out as a misconfiguration or  
a potential culprit for the behavior I'm seeing? I haven't seen anything  
in logging so far to indicate an issue. Are there other data points that  
would be useful in troubleshooting this? I don't know how to reproduce it  
so I'm skipping over creating the gist that the website requests for the  
moment until I get a bit of feedback on what's actually useful.

Thanks in advance for your help with this.

Michael Turner

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Mike\_Turner](https://avatars.discourse-cdn.com/v4/letter/m/ba8739/32.png) [@Mike\_Turner](https://discuss.elastic.co/u/Mike_Turner)\
**Post date:** [October 30, 2013, 10:15pm UTC](https://discuss.elastic.co/t/getting-occasional-unassigned-replica-shards/14178/2 "2013-10-30T22:15:05Z")

</div>

I should also note -- 0.90.3 running on OEL6U2.

On Wednesday, October 30, 2013 3:14:07 PM UTC-7, Mike Turner wrote:

> Hi there,
> 
> I'm running a cluster that's fairly busy, and after doing a full cluster  
> restart and a bit of tuning to the configs, I'm getting an occasional  
> unassigned replica shard that never seems to get assigned. We generate a  
> new index daily, and last night I saw for the first time a new index get  
> created with a missing replica shard and show up as unassigned. I've since  
> deleted that index and let it recreate from inbound data, and it created a  
> second time the same way, with one replica unassigned.
> 
> So really, two issues here:
> 
> 1. On a cluster restart, four indices came up with an unassigned replica  
> shard, and the cluster did not self-correct that condition.
> 2. On a new index being generated, one replica shard is missing,  
> repeatable 2x, and the cluster did not self-correct that condition.
> 
> I found a solution with the help of a user in the #elasticsearch channel  
> that showed me how to reindex the older indices to get the missing replica  
> shard back (using the Tire library's index.reindex method in combination  
> with an alias) and I am in the process of reindexing the four older indices  
> that had missing replicas. That process will take more than a week (~2.5  
> days per index) with the size of things and level of IO activity we  
> presently have.
> 
> I really want to get to the bottom of why this happened, and even more  
> importantly, why a new index would get created without all of the required  
> shards.
> 
> We haven't seen this happen before, so I suspect that it's a product of  
> the tuning that I recently did. Here's what changed:
> 
> 1. Doubled the size of the cluster from 3 nodes to 6.
> 2. Increased the shard and replica count from 3 shards to 6, and from 1  
> replica to 2 at the same time.
> 3. "index.routing.allocation.total\_shards\_per\_node" : 3
> 4. discovery.zen.minimum\_master\_nodes: 4
> 5. gateway.recover\_after\_nodes: 4
> 6. gateway.recover\_after\_time: 10m
> 7. gateway.expected\_data\_nodes: 2
> 8. gateway.expected\_master\_nodes: 6
> 
> Is there anything in those settings that stand out as a misconfiguration  
> or a potential culprit for the behavior I'm seeing? I haven't seen  
> anything in logging so far to indicate an issue. Are there other data  
> points that would be useful in troubleshooting this? I don't know how to  
> reproduce it so I'm skipping over creating the gist that the website  
> requests for the moment until I get a bit of feedback on what's actually  
> useful.
> 
> Thanks in advance for your help with this.
> 
> Michael Turner

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:09am UTC](https://discuss.elastic.co/t/getting-occasional-unassigned-replica-shards/14178/3 "2017-07-06T02:09:42Z")

</div>


