# Heavy load on one node (1 index)

**URL:** <https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034>\
**Category:** Elasticsearch\
**Created:** [March 17, 2012, 2:10am UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034 "2012-03-17T02:10:05Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![John\_Cwikla](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@John\_Cwikla](https://discuss.elastic.co/u/John_Cwikla)\
**Post date:** [March 17, 2012, 2:10am UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/1 "2012-03-17T02:10:05Z")

</div>

We have a 4 node cluster with 1 index, 25m docs or so. We have many  
processes reading/writing to our cluster through  
pyes which we've verified randomly chooses an initial node to talk to.

Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is taking  
the brunt -  
40-50 load (16 core).

We can't figure out why - any pointers? Anything I can do to debug what is  
going on?

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 17, 2012, 11:29am UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/2 "2012-03-17T11:29:06Z")

</div>

Can you gist a stack trace on the busy node? You can use jstack to do so:  
[jstack - Stack Trace](http://docs.oracle.com/javase/1.5.0/docs/tooldocs/share/jstack.html).

On Sat, Mar 17, 2012 at 4:10 AM, John Cwikla [cwikla@fwix.com](mailto:cwikla@fwix.com) wrote:

> We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> processes reading/writing to our cluster through  
> pyes which we've verified randomly chooses an initial node to talk to.
> 
> Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is taking  
> the brunt -  
> 40-50 load (16 core).
> 
> We can't figure out why - any pointers? Anything I can do to debug what is  
> going on?

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [March 17, 2012, 12:09pm UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/3 "2012-03-17T12:09:28Z")

</div>

On Fri, 2012-03-16 at 19:10 -0700, John Cwikla wrote:

> We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> processes reading/writing to our cluster through  
> pyes which we've verified randomly chooses an initial node to talk to.
> 
> Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is  
> taking the brunt -  
> 40-50 load (16 core).

Is it continuous? or just for a period? If the latter, it was probably  
doing a merge of big segments. So the load is actually IO, not CPU.

I've been seeing the same thing on one of my indexes. It starts to do a  
single merge of GBs of data, and the machine is too overloaded to  
respond to requests.

> 

You can check if there is a merge happening by looking at the index  
status:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

clint

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [March 19, 2012, 9:35am UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/4 "2012-03-19T09:35:11Z")

</div>

John,

Can you check if your shards are well balanced?  
You may be able to tell what's going on using ES head and SPM for ES:

[http://mobz.github.com/elasticsearch-head/](http://mobz.github.com/elasticsearch-head/)  
[http://apps.sematext.com/](http://apps.sematext.com/)

The former you run on your own servers, the latter is a service. Both are  
free.

## Otis

Hiring Elasticsearch Engineers World-Wide --

> **[Jobs](https://sematext.com/jobs/)**
>
> We’re Hiring We are always looking for smart, passionate, motivated, and independent people regardless of where on the planet they may be. Learn more about the company Agent & Backend Engineer Full Stack Developer Backend Engineer Frontend...

On Saturday, March 17, 2012 10:10:05 AM UTC+8, John Cwikla wrote:

> We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> processes reading/writing to our cluster through  
> pyes which we've verified randomly chooses an initial node to talk to.
> 
> Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is taking  
> the brunt -  
> 40-50 load (16 core).
> 
> We can't figure out why - any pointers? Anything I can do to debug what is  
> going on?

---

<div class="post-metadata">

**Author:** ![John\_Cwikla](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@John\_Cwikla](https://discuss.elastic.co/u/John_Cwikla)\
**Post date:** [March 19, 2012, 5:53pm UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/5 "2012-03-19T17:53:24Z")

</div>

Thanks for all the pointers - I'll start diving into all these and let you  
know what  
I come up with.

On Monday, March 19, 2012 2:35:11 AM UTC-7, Otis Gospodnetic wrote:

> John,
> 
> Can you check if your shards are well balanced?  
> You may be able to tell what's going on using ES head and SPM for ES:
> 
> [http://mobz.github.com/elasticsearch-head/](http://mobz.github.com/elasticsearch-head/)  
> [http://apps.sematext.com/](http://apps.sematext.com/)
> 
> The former you run on your own servers, the latter is a service. Both are  
> free.
> 
> ## Otis
> 
> Hiring Elasticsearch Engineers World-Wide --  
> [Jobs - Sematext](http://sematext.com/about/jobs.html#search)
> 
> On Saturday, March 17, 2012 10:10:05 AM UTC+8, John Cwikla wrote:
> 
> > We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> > processes reading/writing to our cluster through  
> > pyes which we've verified randomly chooses an initial node to talk to.
> > 
> > Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is  
> > taking the brunt -  
> > 40-50 load (16 core).
> > 
> > We can't figure out why - any pointers? Anything I can do to debug what  
> > is going on?

---

<div class="post-metadata">

**Author:** ![John\_Cwikla](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@John\_Cwikla](https://discuss.elastic.co/u/John_Cwikla)\
**Post date:** [March 19, 2012, 8:34pm UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/6 "2012-03-19T20:34:08Z")

</div>

Elasticsearch head gave me some interesting results. So we have 3 indexes  
on our 4 node cluster.

The first index is just a test, so ignore it.

The second is 25m docs (19G), and has 3 shards on 1 node and 1 on each other

The third is 16m docs (30G) and has 5 shards on 1, 4 on another, 2 on  
another, and 1 on the  
last.

The one with 5 shards is the one I was having the problem with. Seems like  
I need to have  
the nodes rebalance somehow to take advantage of all the hardware I have...

At this point our load is average across the machines, so I haven't used  
jstack, but I will  
next time it pops up.

On Friday, March 16, 2012 7:10:05 PM UTC-7, John Cwikla wrote:

> We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> processes reading/writing to our cluster through  
> pyes which we've verified randomly chooses an initial node to talk to.
> 
> Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is taking  
> the brunt -  
> 40-50 load (16 core).
> 
> We can't figure out why - any pointers? Anything I can do to debug what is  
> going on?

---

<div class="post-metadata">

**Author:** ![John\_Cwikla](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@John\_Cwikla](https://discuss.elastic.co/u/John_Cwikla)\
**Post date:** [March 21, 2012, 6:52pm UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/7 "2012-03-21T18:52:26Z")

</div>

Hmm, so I've been looking around and saw that rebalancing replica shards  
isn't something that can be done?  
It is the case the primaries are balanced, but the replicas are overloaded  
on 1 machine in a 4 machine cluster.

Is there something I need to do, or can do to make them rebalance?

On Friday, March 16, 2012 7:10:05 PM UTC-7, John Cwikla wrote:

> We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> processes reading/writing to our cluster through  
> pyes which we've verified randomly chooses an initial node to talk to.
> 
> Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is taking  
> the brunt -  
> 40-50 load (16 core).
> 
> We can't figure out why - any pointers? Anything I can do to debug what is  
> going on?

On Friday, March 16, 2012 7:10:05 PM UTC-7, John Cwikla wrote:

> We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> processes reading/writing to our cluster through  
> pyes which we've verified randomly chooses an initial node to talk to.
> 
> Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is taking  
> the brunt -  
> 40-50 load (16 core).
> 
> We can't figure out why - any pointers? Anything I can do to debug what is  
> going on?

On Friday, March 16, 2012 7:10:05 PM UTC-7, John Cwikla wrote:

> We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> processes reading/writing to our cluster through  
> pyes which we've verified randomly chooses an initial node to talk to.
> 
> Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is taking  
> the brunt -  
> 40-50 load (16 core).
> 
> We can't figure out why - any pointers? Anything I can do to debug what is  
> going on?

---

<div class="post-metadata">

**Author:** ![John\_Cwikla](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@John\_Cwikla](https://discuss.elastic.co/u/John_Cwikla)\
**Post date:** [March 21, 2012, 7:17pm UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/8 "2012-03-21T19:17:56Z")

</div>

And here's my shard list:

{

- index: {
  - primary\_size: 30.5gb
  - primary\_size\_in\_bytes: 32762267538
  - size: 62.5gb
  - size\_in\_bytes: 67185556471  
}

- translog: {
  - operations: 28487  
}

- docs: {
  - num\_docs: 16378524
  - max\_doc: 24589536
  - deleted\_docs: 8211012  
}

- merges: {
  - current: 0
  - current\_docs: 0
  - current\_size: 0b
  - current\_size\_in\_bytes: 0
  - total: 943760
  - total\_time: 1.5d
  - total\_time\_in\_millis: 129794974
  - total\_docs: 2083671491
  - total\_size: 2704.7gb
  - total\_size\_in\_bytes: 2904189840468  
}

- refresh: {
  - total: 8353127
  - total\_time: 1.9d
  - total\_time\_in\_millis: 171464206  
}

- flush: {
  - total: 131156
  - total\_time: 5.5h
  - total\_time\_in\_millis: 19850157  
}

- shards: {
  - 0: [
    - {
      - routing: {
        - state: STARTED
        - primary: false
        - node: nELxlGs6Tzii12JRCNxpGw
        - relocating\_node: null
        - shard: 0
        - index: places  
}

      - state: STARTED
      - index: {
        - size: 4.9gb
        - size\_in\_bytes: 5352243736  
}

      - translog: {
        - id: 1326141078235
        - operations: 3211  
}

      - docs: {
        - num\_docs: 2647990
        - max\_doc: 4003535
        - deleted\_docs: 1355545  
}

      - merges: {
        - current: 0
        - current\_docs: 0
        - current\_size: 0b
        - current\_size\_in\_bytes: 0
        - total: 98373
        - total\_time: 4.6h
        - total\_time\_in\_millis: 16647902
        - total\_docs: 213859847
        - total\_size: 276.4gb
        - total\_size\_in\_bytes: 296867190719  
}

      - refresh: {
        - total: 871137
        - total\_time: 5.4h
        - total\_time\_in\_millis: 19523847  
}

      - flush: {
        - total: 13594
        - total\_time: 49.3m
        - total\_time\_in\_millis: 2963485  
}  
}

    - {
      - routing: {
        - state: STARTED
        - primary: true
        - node: ab4uo\_iyTw2WIT2-eb\_F6Q
        - relocating\_node: null
        - shard: 0
        - index: places  
}

      - state: STARTED
      - index: {
        - size: 4.9gb
        - size\_in\_bytes: 5319904685  
}

      - translog: {
        - id: 1326141078245
        - operations: 2483  
}

      - docs: {
        - num\_docs: 2647990
        - max\_doc: 3981357
        - deleted\_docs: 1333367  
}

      - merges: {
        - current: 0
        - current\_docs: 0
        - current\_size: 0b
        - current\_size\_in\_bytes: 0
        - total: 37155
        - total\_time: 57.1m
        - total\_time\_in\_millis: 3426101
        - total\_docs: 78881622
        - total\_size: 101.8gb
        - total\_size\_in\_bytes: 109329911898  
}

      - refresh: {
        - total: 328715
        - total\_time: 1.5h
        - total\_time\_in\_millis: 5755051  
}

      - flush: {
        - total: 4796
        - total\_time: 5.7m
        - total\_time\_in\_millis: 344400  
}  
}  
]

  - 1: [
    - {
      - routing: {
        - state: STARTED
        - primary: false
        - node: ab4uo\_iyTw2WIT2-eb\_F6Q
        - relocating\_node: null
        - shard: 1
        - index: places  
}

      - state: STARTED
      - index: {
        - size: 5.8gb
        - size\_in\_bytes: 6262820543  
}

      - translog: {
        - id: 1326141078607
        - operations: 56  
}

      - docs: {
        - num\_docs: 2803977
        - max\_doc: 4754923
        - deleted\_docs: 1950946  
}

      - merges: {
        - current: 0
        - current\_docs: 0
        - current\_size: 0b
        - current\_size\_in\_bytes: 0
        - total: 37336
        - total\_time: 57.3m
        - total\_time\_in\_millis: 3443330
        - total\_docs: 79711670
        - total\_size: 103.8gb
        - total\_size\_in\_bytes: 111496656320  
}

      - refresh: {
        - total: 330123
        - total\_time: 1.6h
        - total\_time\_in\_millis: 5859695  
}

      - flush: {
        - total: 4921
        - total\_time: 6m
        - total\_time\_in\_millis: 360131  
}  
}

    - {
      - routing: {
        - state: STARTED
        - primary: true
        - node: nELxlGs6Tzii12JRCNxpGw
        - relocating\_node: null
        - shard: 1
        - index: places  
}

      - state: STARTED
      - index: {
        - size: 5.8gb
        - size\_in\_bytes: 6296790810  
}

      - translog: {
        - id: 1326141078607
        - operations: 51  
}

      - docs: {
        - num\_docs: 2803978
        - max\_doc: 4786774
        - deleted\_docs: 1982796  
}

      - merges: {
        - current: 0
        - current\_docs: 0
        - current\_size: 0b
        - current\_size\_in\_bytes: 0
        - total: 99070
        - total\_time: 4.8h
        - total\_time\_in\_millis: 17343659
        - total\_docs: 218051849
        - total\_size: 283gb
        - total\_size\_in\_bytes: 303888376066  
}

      - refresh: {
        - total: 876293
        - total\_time: 5.8h
        - total\_time\_in\_millis: 20913727  
}

      - flush: {
        - total: 13835
        - total\_time: 53.2m
        - total\_time\_in\_millis: 3192033  
}  
}  
]

  - 2: [
    - {
      - routing: {
        - state: STARTED
        - primary: true
        - node: ab4uo\_iyTw2WIT2-eb\_F6Q
        - relocating\_node: null
        - shard: 2
        - index: places  
}

      - state: STARTED
      - index: {
        - size: 5.2gb
        - size\_in\_bytes: 5589513799  
}

      - translog: {
        - id: 1326141079557
        - operations: 2653  
}

      - docs: {
        - num\_docs: 2737540
        - max\_doc: 4189158
        - deleted\_docs: 1451618  
}

      - merges: {
        - current: 0
        - current\_docs: 0
        - current\_size: 0b
        - current\_size\_in\_bytes: 0
        - total: 37295
        - total\_time: 1h
        - total\_time\_in\_millis: 3602012
        - total\_docs: 86544448
        - total\_size: 111.6gb
        - total\_size\_in\_bytes: 119899765102  
}

      - refresh: {
        - total: 328507
        - total\_time: 1.6h
        - total\_time\_in\_millis: 6103834  
}

      - flush: {
        - total: 5502
        - total\_time: 6.5m
        - total\_time\_in\_millis: 395750  
}  
}

    - {
      - routing: {
        - state: STARTED
        - primary: false
        - node: nELxlGs6Tzii12JRCNxpGw
        - relocating\_node: null
        - shard: 2
        - index: places  
}

      - state: STARTED
      - index: {
        - size: 5.4gb
        - size\_in\_bytes: 5820645020  
}

      - translog: {
        - id: 1326141079550
        - operations: 4779  
}

      - docs: {
        - num\_docs: 2737540
        - max\_doc: 4368769
        - deleted\_docs: 1631229  
}

      - merges: {
        - current: 0
        - current\_docs: 0
        - current\_size: 0b
        - current\_size\_in\_bytes: 0
        - total: 98407
        - total\_time: 4.9h
        - total\_time\_in\_millis: 17647596
        - total\_docs: 222704925
        - total\_size: 289.4gb
        - total\_size\_in\_bytes: 310757299702  
}

      - refresh: {
        - total: 870872
        - total\_time: 5.6h
        - total\_time\_in\_millis: 20188623  
}

      - flush: {
        - total: 14195
        - total\_time: 52.5m
        - total\_time\_in\_millis: 3154127  
}  
}  
]

  - 3: [
    - {
      - routing: {
        - state: STARTED
        - primary: true
        - node: T5\_oKWa3RTO7osracRtPoQ
        - relocating\_node: null
        - shard: 3
        - index: places  
}

      - state: STARTED
      - index: {
        - size: 4.9gb
        - size\_in\_bytes: 5343294260  
}

      - translog: {
        - id: 1326141078245
        - operations: 2124  
}

      - docs: {
        - num\_docs: 2648125
        - max\_doc: 3993220
        - deleted\_docs: 1345095  
}

      - merges: {
        - current: 0
        - current\_docs: 0
        - current\_size: 0b
        - current\_size\_in\_bytes: 0
        - total: 100192
        - total\_time: 2.7h
        - total\_time\_in\_millis: 9740764
        - total\_docs: 216533766
        - total\_size: 280.8gb
        - total\_size\_in\_bytes: 301514882857  
}

      - refresh: {
        - total: 888382
        - total\_time: 4.3h
        - total\_time\_in\_millis: 15669718  
}

      - flush: {
        - total: 13576
        - total\_time: 15.2m
        - total\_time\_in\_millis: 917058  
}  
}

    - {
      - routing: {
        - state: STARTED
        - primary: false
        - node: ab4uo\_iyTw2WIT2-eb\_F6Q
        - relocating\_node: null
        - shard: 3
        - index: places  
}

      - state: STARTED
      - index: {
        - size: 4.9gb
        - size\_in\_bytes: 5349875537  
}

      - translog: {
        - id: 1326141078241
        - operations: 3886  
}

      - docs: {
        - num\_docs: 2648125
        - max\_doc: 3994418
        - deleted\_docs: 1346293  
}

      - merges: {
        - current: 0
        - current\_docs: 0
        - current\_size: 0b
        - current\_size\_in\_bytes: 0
        - total: 37161
        - total\_time: 56.5m
        - total\_time\_in\_millis: 3393844
        - total\_docs: 78031855
        - total\_size: 101.5gb
        - total\_size\_in\_bytes: 108996560347  
}

      - refresh: {
        - total: 328585
        - total\_time: 1.6h
        - total\_time\_in\_millis: 5779601  
}

      - flush: {
        - total: 4780
        - total\_time: 5.8m
        - total\_time\_in\_millis: 351341  
}  
}  
]

  - 4: [
    - {
      - routing: {
        - state: STARTED
        - primary: true
        - node: T5\_oKWa3RTO7osracRtPoQ
        - relocating\_node: null
        - shard: 4
        - index: places  
}

      - state: STARTED
      - index: {
        - size: 5.3gb
        - size\_in\_bytes: 5757997414  
}

      - translog: {
        - id: 1326141078676
        - operations: 2550  
}

      - docs: {
        - num\_docs: 2804287
        - max\_doc: 4323833
        - deleted\_docs: 1519546  
}

      - merges: {
        - current: 0
        - current\_docs: 0
        - current\_size: 0b
        - current\_size\_in\_bytes: 0
        - total: 100919
        - total\_time: 2.7h
        - total\_time\_in\_millis: 9958928
        - total\_docs: 221717044
        - total\_size: 288.5gb
        - total\_size\_in\_bytes: 309792436202  
}

      - refresh: {
        - total: 893623
        - total\_time: 4.3h
        - total\_time\_in\_millis: 15799989  
}

      - flush: {
        - total: 13890
        - total\_time: 15.4m
        - total\_time\_in\_millis: 927816  
}  
}

    - {
      - routing: {
        - state: STARTED
        - primary: false
        - node: nELxlGs6Tzii12JRCNxpGw
        - relocating\_node: null
        - shard: 4
        - index: places  
}

      - state: STARTED
      - index: {
        - size: 5.3gb
        - size\_in\_bytes: 5746398501  
}

      - translog: {
        - id: 1326141078670
        - operations: 852  
}

      - docs: {
        - num\_docs: 2804287
        - max\_doc: 4316365
        - deleted\_docs: 1512078  
}

      - merges: {
        - current: 0
        - current\_docs: 0
        - current\_size: 0b
        - current\_size\_in\_bytes: 0
        - total: 99169
        - total\_time: 4.7h
        - total\_time\_in\_millis: 17225939
        - total\_docs: 222259294
        - total\_size: 288.5gb
        - total\_size\_in\_bytes: 309856763814  
}

      - refresh: {
        - total: 877186
        - total\_time: 5.5h
        - total\_time\_in\_millis: 19947688  
}

      - flush: {
        - total: 13880
        - total\_time: 51.7m
        - total\_time\_in\_millis: 3106099  
}  
}  
]

  - 5: [
    - {
      - routing: {
        - state: STARTED
        - primary: false
        - node: nELxlGs6Tzii12JRCNxpGw
        - relocating\_node: null
        - shard: 5
        - index: places  
}

      - state: STARTED
      - index: {
        - size: 5.4gb
        - size\_in\_bytes: 5891305596  
}

      - translog: {
        - id: 1326141079451
        - operations: 388  
}

      - docs: {
        - num\_docs: 2736604
        - max\_doc: 4468212
        - deleted\_docs: 1731608  
}

      - merges: {
        - current: 0
        - current\_docs: 0
        - current\_size: 0b
        - current\_size\_in\_bytes: 0
        - total: 98430
        - total\_time: 4.8h
        - total\_time\_in\_millis: 17328442
        - total\_docs: 218687019
        - total\_size: 284.5gb
        - total\_size\_in\_bytes: 305584707796  
}

      - refresh: {
        - total: 871118
        - total\_time: 5.5h
        - total\_time\_in\_millis: 20154030  
}

      - flush: {
        - total: 14086
        - total\_time: 54m
        - total\_time\_in\_millis: 3242524  
}  
}

    - {
      - routing: {
        - state: STARTED
        - primary: true
        - node: XGfCby6mT0Ox2zESwvBPsw
        - relocating\_node: null
        - shard: 5
        - index: places  
}

      - state: STARTED
      - index: {
        - size: 4.1gb
        - size\_in\_bytes: 4454766570  
}

      - translog: {
        - id: 1326141079457
        - operations: 5454  
}

      - docs: {
        - num\_docs: 2736604
        - max\_doc: 3315194
        - deleted\_docs: 578590  
}

      - merges: {
        - current: 0
        - current\_docs: 0
        - current\_size: 0b
        - current\_size\_in\_bytes: 0
        - total: 100253
        - total\_time: 2.7h
        - total\_time\_in\_millis: 10036457
        - total\_docs: 226688152
        - total\_size: 294.4gb
        - total\_size\_in\_bytes: 316205289645  
}

      - refresh: {
        - total: 888586
        - total\_time: 4.3h
        - total\_time\_in\_millis: 15768403  
}

      - flush: {
        - total: 14101
        - total\_time: 14.9m
        - total\_time\_in\_millis: 895393  
}  
}  
]  
}

}

On Friday, March 16, 2012 7:10:05 PM UTC-7, John Cwikla wrote:

> We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> processes reading/writing to our cluster through  
> pyes which we've verified randomly chooses an initial node to talk to.
> 
> Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is taking  
> the brunt -  
> 40-50 load (16 core).
> 
> We can't figure out why - any pointers? Anything I can do to debug what is  
> going on?

---

<div class="post-metadata">

**Author:** ![John\_Cwikla](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@John\_Cwikla](https://discuss.elastic.co/u/John_Cwikla)\
**Post date:** [March 22, 2012, 1:52am UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/9 "2012-03-22T01:52:56Z")

</div>

Ha, I figured out what was going on. I was totally ignoring a test index  
that had been balanced onto two of the nodes and  
so the 3 indexes' shards were balanced across ALL the nodes, my assumption  
is that each index would be balanced by shard across  
all the machines not that ALL shards are balanced across the nodes,  
regardless of index.

I'd love that as an option (to balance shards across the nodes by index),  
as it's moving my two indexes to be split pretty much 1 index on 2  
machines, 1 index on the other  
2 machines. Since one of my indexes is a legacy index (we use it about  
1/100th the other one), I'm still not utilizing my machines.

Now I'll just need to figure out how to trick it 🙂

On Friday, March 16, 2012 7:10:05 PM UTC-7, John Cwikla wrote:

> We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> processes reading/writing to our cluster through  
> pyes which we've verified randomly chooses an initial node to talk to.
> 
> Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is taking  
> the brunt -  
> 40-50 load (16 core).
> 
> We can't figure out why - any pointers? Anything I can do to debug what is  
> going on?

---

<div class="post-metadata">

**Author:** ![John\_Cwikla](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@John\_Cwikla](https://discuss.elastic.co/u/John_Cwikla)\
**Post date:** [March 22, 2012, 2:53am UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/10 "2012-03-22T02:53:34Z")

</div>

And just to finish up - I ended up setting my legacy index to have it's  
replica's to node-1 instances -  
it's only read-only, so it's just disk cost - and this is forcing my new  
index to balance across  
my machines evenly like I want.

Sweet!

On Friday, March 16, 2012 7:10:05 PM UTC-7, John Cwikla wrote:

> We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> processes reading/writing to our cluster through  
> pyes which we've verified randomly chooses an initial node to talk to.
> 
> Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is taking  
> the brunt -  
> 40-50 load (16 core).
> 
> We can't figure out why - any pointers? Anything I can do to debug what is  
> going on?

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [March 22, 2012, 11:04pm UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/11 "2012-03-22T23:04:49Z")

</div>

I wonder if you are coming across this issue:  
[http://elasticsearch-users.115913.n3.nabble.com/Shard-Balancing-td3676576.html](http://elasticsearch-users.115913.n3.nabble.com/Shard-Balancing-td3676576.html)

On Wed, Mar 21, 2012 at 7:53 PM, John Cwikla [cwikla@fwix.com](mailto:cwikla@fwix.com) wrote:

> And just to finish up - I ended up setting my legacy index to have it's  
> replica's to node-1 instances -  
> it's only read-only, so it's just disk cost - and this is forcing my new  
> index to balance across  
> my machines evenly like I want.
> 
> Sweet!
> 
> On Friday, March 16, 2012 7:10:05 PM UTC-7, John Cwikla wrote:
> 
> > We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> > processes reading/writing to our cluster through  
> > pyes which we've verified randomly chooses an initial node to talk to.
> > 
> > Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is taking  
> > the brunt -  
> > 40-50 load (16 core).
> > 
> > We can't figure out why - any pointers? Anything I can do to debug what is  
> > going on?

---

<div class="post-metadata">

**Author:** ![John\_Cwikla](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@John\_Cwikla](https://discuss.elastic.co/u/John_Cwikla)\
**Post date:** [March 23, 2012, 1:18am UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/12 "2012-03-23T01:18:02Z")

</div>

Yep, that looks like it. I like the pointer at the ability to create your  
own allocator, definitely will look into that.

On Friday, March 16, 2012 7:10:05 PM UTC-7, John Cwikla wrote:

> We have a 4 node cluster with 1 index, 25m docs or so. We have many  
> processes reading/writing to our cluster through  
> pyes which we've verified randomly chooses an initial node to talk to.
> 
> Strange thing is, 3 of the 4 nodes are bored (0.25 load), 1 node is taking  
> the brunt -  
> 40-50 load (16 core).
> 
> We can't figure out why - any pointers? Anything I can do to debug what is  
> going on?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:35am UTC](https://discuss.elastic.co/t/heavy-load-on-one-node-1-index/7034/13 "2017-07-06T03:35:06Z")

</div>


