# Elasticsearch replica shard distribution

**URL:** <https://discuss.elastic.co/t/elasticsearch-replica-shard-distribution/95371>\
**Category:** Elasticsearch\
**Created:** [August 1, 2017, 3:12pm UTC](https://discuss.elastic.co/t/elasticsearch-replica-shard-distribution/95371 "2017-08-01T15:12:04Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![monkegoist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/monkegoist/32/20684_2.png) [@monkegoist](https://discuss.elastic.co/u/monkegoist)\
**Post date:** [August 1, 2017, 3:12pm UTC](https://discuss.elastic.co/t/elasticsearch-replica-shard-distribution/95371/1 "2017-08-01T15:12:05Z")

</div>

Hello,

We recently came across an issue with users reporting their queries being slow. There are 5 nodes in a cluster, index has 10 shards, replication factor is 1. Upon investigation, I found out that each of the nodes has 4 shards allocated to it, but for one node all of the shards are replicas while others contain only 1 or 2 replica shards.

As far as I understand, search queries are executed against replica shards which would explain why this node suffers the most (judging by slow queries log and kopf cluster health info).

Is there a way to adjust cluster settings in such a way that both primary and replica shards are allocated evenly? Or is there another approach that we could use in our case?

Thanks in advance for your help!

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [August 1, 2017, 3:40pm UTC](https://discuss.elastic.co/t/elasticsearch-replica-shard-distribution/95371/2 "2017-08-01T15:40:00Z")

</div>

> [@monkegoist](#):
>
> As far as I understand, search queries are executed against replica shards which would explain why this node suffers the most (judging by slow queries log and kopf cluster health info).

I think that is not true, according to the documentation, [replicas and primaries do the same amount of work](https://www.elastic.co/guide/en/elasticsearch/guide/master/replica-shards.html#img-three-nodes).

The cause of the slow queries must be something else. It is only one node that shows performance issue?

---

<div class="post-metadata">

**Author:** ![monkegoist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/monkegoist/32/20684_2.png) [@monkegoist](https://discuss.elastic.co/u/monkegoist)\
**Post date:** [August 3, 2017, 2:33pm UTC](https://discuss.elastic.co/t/elasticsearch-replica-shard-distribution/95371/3 "2017-08-03T14:33:35Z")

</div>

Hi Leandro,

You're right, my statement about query execution against replica shards seems to be false. I checked Elastic source code and found out that if preference isn't specified, random order seems to be used (`IndexShardRoutingTable#activeInitializingShardsRandomIt()`).

Here's what I have in slow logs for the query in question:

node 1:

```
[2017-08-01 14:01:06,286][WARN][index.search.slowlog.query] [Snowfall] [index][0] took[5.2s], took_millis[5242] - primary
[2017-08-01 14:01:07,351][WARN][index.search.slowlog.query] [Snowfall] [index][2] took[6.2s], took_millis[6293] - primary
[2017-08-01 14:01:07,374][WARN][index.search.slowlog.query] [Snowfall] [index][1] took[6s], took_millis[6085] - primary

```

node 2:

```
[2017-08-01 14:02:23,111][WARN][index.search.slowlog.query] [Luchino Nefaria] [index][7] took[43.2s], took_millis[43201] - replica
[2017-08-01 14:02:49,445][WARN][index.search.slowlog.query] [Luchino Nefaria] [index][9] took[1.7m], took_millis[106090] - replica
[2017-08-01 14:02:54,152][WARN][index.search.slowlog.query] [Luchino Nefaria] [index][3] took[1.8m], took_millis[113128] - replica
[2017-08-01 14:03:27,645][WARN][index.search.slowlog.query] [Luchino Nefaria] [index][5] took[2.4m], took_millis[144610] - replica

```

node 3: nothing

node 4:

```
[2017-08-01 14:01:17,866][WARN][index.search.slowlog.query] [Xemnu the Titan] [index][4] took[16.7s], took_millis[16747] - primary
[2017-08-01 14:01:21,942][WARN][index.search.slowlog.query] [Xemnu the Titan] [index][6] took[20.8s], took_millis[20837] - primary

```

node 5:

`[2017-08-01 14:01:05,856][WARN][index.search.slowlog.query] [Gideon] [index][8] took[4.8s], took_millis[4815] - replica`

As you can see, node 2 had the most work to do. Although node 1 had a lot of work too but managed to complete it much faster.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 31, 2017, 2:34pm UTC](https://discuss.elastic.co/t/elasticsearch-replica-shard-distribution/95371/4 "2017-08-31T14:34:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
