# Skewed primary shards distribution leads to performance issues

**URL:** <https://discuss.elastic.co/t/skewed-primary-shards-distribution-leads-to-performance-issues/89108>\
**Category:** Elasticsearch\
**Created:** [June 12, 2017, 8:37pm UTC](https://discuss.elastic.co/t/skewed-primary-shards-distribution-leads-to-performance-issues/89108 "2017-06-12T20:37:30Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dan\_Markhasin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_markhasin/32/14187_2.png) [@Dan\_Markhasin](https://discuss.elastic.co/u/Dan_Markhasin)\
**Post date:** [June 12, 2017, 8:37pm UTC](https://discuss.elastic.co/t/skewed-primary-shards-distribution-leads-to-performance-issues/89108/1 "2017-06-12T20:37:31Z")

</div>

We have a relatively large monthly index (reaches approx. 1TB by the end of the month, so about 30GB are added daily) that is doing hundreds of updates per second - it has 10 shards with 1 replica, and is spread out evenly across 10 physical machines - 2 shards on each server.

I've noticed however that the load distribution is far from balanced - some machines have 2 primary shards for the index and they seem to be doing most of the update work with a full GC cycle every 7 minutes or so.  
Machines with 1 primary shard and 1 replica experience a full GC approx. every 30 minutes and machines that only hold replica shards are the least utilized with approx. 45 minutes between full GC cycles.

Machine with 2 primary shards:

![](https://us1.discourse-cdn.com/elastic/original/3X/7/8/78c85bd8d6b3418b804273c5f5b12ca1aa656a6e.png)

Machine with 1 primary and 1 replica:

![](https://us1.discourse-cdn.com/elastic/original/3X/9/5/95bdcdd92c53a137e08a2f0430bf13ca6f094838.png)

Machine with 2 replica shards:

![](https://us1.discourse-cdn.com/elastic/original/3X/0/4/042aff76539a691766986e26b083275b260fc2f9.png)

Is there any way to rebalance the primary shards so that there is no more than 1 primary shard on each machine?

---

<div class="post-metadata">

**Author:** ![Dan\_Markhasin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_markhasin/32/14187_2.png) [@Dan\_Markhasin](https://discuss.elastic.co/u/Dan_Markhasin)\
**Post date:** [June 14, 2017, 10:23am UTC](https://discuss.elastic.co/t/skewed-primary-shards-distribution-leads-to-performance-issues/89108/2 "2017-06-14T10:23:27Z")

</div>

Any suggestions will be appreciated 😑

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 18, 2017, 11:40am UTC](https://discuss.elastic.co/t/skewed-primary-shards-distribution-leads-to-performance-issues/89108/3 "2017-06-18T11:40:08Z")

</div>

> [@Dan\_Markhasin](#):
>
> Machine with 2 primary shards:

Is this a node or a host?  
Are you using allocation awareness?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 18, 2017, 12:00pm UTC](https://discuss.elastic.co/t/skewed-primary-shards-distribution-leads-to-performance-issues/89108/4 "2017-06-18T12:00:28Z")

</div>

What type of issues is this causing?

---

<div class="post-metadata">

**Author:** ![Dan\_Markhasin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_markhasin/32/14187_2.png) [@Dan\_Markhasin](https://discuss.elastic.co/u/Dan_Markhasin)\
**Post date:** [July 1, 2017, 6:59pm UTC](https://discuss.elastic.co/t/skewed-primary-shards-distribution-leads-to-performance-issues/89108/5 "2017-07-01T18:59:30Z")

</div>

Apologies for the late reply, I was on vacation.

It is not causing any issue at the moment, but it means that we can't really scale out; the more volume of data we push into this index, the more load the machines with the primary shards will need to handle, until at some point they will crash. In this case adding more shards or more machines is not going to help at all, since we can't guarantee even load distribution; even if we double the amount of shards and machines, we may remain in the same situation with some machines handling most of the load while others being mostly idle.

To answer @warkolm's question, we are not using allocation awareness (I'm not sure how it would help), and I'm not sure what you mean by "node or a host"? These are screen captures from Kibana showing the JVM heap utilization of the different data nodes in the cluster.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [July 1, 2017, 9:00pm UTC](https://discuss.elastic.co/t/skewed-primary-shards-distribution-leads-to-performance-issues/89108/6 "2017-07-01T21:00:10Z")

</div>

Looking at the GC graphs they seem fine, you have a very nice sawtooth that has good gaps - ie it's not every minute.

Have a look at [https://www.elastic.co/guide/en/elasticsearch/reference/5.4/allocation-awareness.html](https://www.elastic.co/guide/en/elasticsearch/reference/5.4/allocation-awareness.html), but it won't stop ES from putting multiple primaries on the same node.

---

<div class="post-metadata">

**Author:** ![Dan\_Markhasin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_markhasin/32/14187_2.png) [@Dan\_Markhasin](https://discuss.elastic.co/u/Dan_Markhasin)\
**Post date:** [July 2, 2017, 11:37am UTC](https://discuss.elastic.co/t/skewed-primary-shards-distribution-leads-to-performance-issues/89108/8 "2017-07-02T11:37:26Z")

</div>

Well that's my point really 🙂  
The GC graphs look fine now, but if we increase the load they will start being much more frequent on the nodes with two primary shards, because the load is not properly distributed.

I do wonder if this kind of load distribution is specific to the update use case (where primary shard does more work than replica)?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [July 2, 2017, 9:09pm UTC](https://discuss.elastic.co/t/skewed-primary-shards-distribution-leads-to-performance-issues/89108/9 "2017-07-02T21:09:18Z")

</div>

> [@Dan\_Markhasin](#):
>
> I do wonder if this kind of load distribution is specific to the update use case (where primary shard does more work than replica)?

How does it do more?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 2, 2017, 9:47pm UTC](https://discuss.elastic.co/t/skewed-primary-shards-distribution-leads-to-performance-issues/89108/10 "2017-07-02T21:47:25Z")

</div>

The primary executes indeed the computation of the new version of the document and then send the result to the replica.

In that sense, primary shard does a bit more job than the replica I guess.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 31, 2017, 3:24am UTC](https://discuss.elastic.co/t/skewed-primary-shards-distribution-leads-to-performance-issues/89108/12 "2017-07-31T03:24:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
