# Transfering indices marked cold to the cold nodes cause GC overhead and other issues

**URL:** <https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529>\
**Category:** Elasticsearch\
**Created:** [September 6, 2017, 7:48am UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529 "2017-09-06T07:48:28Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![John\_ax](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@John\_ax](https://discuss.elastic.co/u/John_ax)\
**Post date:** [September 6, 2017, 7:48am UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/1 "2017-09-06T07:48:29Z")

</div>

Hi,I am using ES 5.5 with the hot-cold architecture, by marking indices cold, he transfer indices to the cold nodes from hot nodes automatically.

the cold nodes using 30GB jvm memory setting, still having the the following log .

```auto
[2017-09-06T15:23:23,583][INFO][o.e.m.j.JvmGcMonitorService] [_GgVD4J] [gc][old][158124][26565] duration [5.2s], collections [1]/[5.6s], total [5.2s]/[2h], memory [29.8gb]->[28.9gb]/[29.8gb], all_pools {[young] [1.1gb]->[430.9mb]/[1.1gb]}{[survivor] [107.7mb]->[0b]/[149.7mb]}{[old] [28.5gb]->[28.5gb]/[28.5gb]}
[2017-09-06T15:23:36,217][INFO][o.e.m.j.JvmGcMonitorService] [_GgVD4J] [gc][old][158128][26567] duration [5.2s], collections [1]/[5.6s], total [5.2s]/[2h], memory [29.8gb]->[28.9gb]/[29.8gb], all_pools {[young] [1.1gb]->[388.9mb]/[1.1gb]}{[survivor] [115.3mb]->[0b]/[149.7mb]}{[old] [28.5gb]->[28.5gb]/[28.5gb]}
[2017-09-06T15:23:48,736][INFO][o.e.m.j.JvmGcMonitorService] [_GgVD4J] [gc][old][158132][26569] duration [5.1s], collections [1]/[5.6s], total [5.1s]/[2h], memory [29.7gb]->[28.9gb]/[29.8gb], all_pools {[young] [1.1gb]->[394.6mb]/[1.1gb]}{[survivor] [100.3mb]->[0b]/[149.7mb]}{[old] [28.5gb]->[28.5gb]/[28.5gb]}
[2017-09-06T15:23:48,737][WARN][o.e.m.j.JvmGcMonitorService] [_GgVD4J] [gc][158132] overhead, spent [5.1s] collecting in the last [5.6s]
[2017-09-06T15:23:54,639][WARN][o.e.m.j.JvmGcMonitorService] [_GgVD4J] [gc][158134] overhead, spent [4.4s] collecting in the last [4.9s]
[2017-09-06T15:24:01,120][WARN][o.e.m.j.JvmGcMonitorService] [_GgVD4J] [gc][158136] overhead, spent [4.6s] collecting in the last [5.4s]
[2017-09-06T15:24:07,129][WARN][o.e.m.j.JvmGcMonitorService] [_GgVD4J] [gc][158138] overhead, spent [4.2s] collecting in the last [5s]

```

and now the ES crash:

```auto
No Monitoring Data Found
No Monitoring data is available for the selected time period. This could be because no data is being sent to the cluster or data was not received during that time.

Try adjusting the time filter controls to a time range where the Monitoring data is expected.

```

Questions:  
1.using 30GB jvm memory is the right opition?  
2.the recovery speed too fast?

```json
  "transient": {
    "cluster": {
      "routing": {
        "rebalance": {
          "enable": "none"
        },
        "allocation": {
          "cluster_concurrent_rebalance": "24",
          "enable": "all"
        }
      }
    },
    "indices": {
      "recovery": {
        "max_bytes_per_sec": "240mb"
      }
    },

```

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 6, 2017, 8:03am UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/2 "2017-09-06T08:03:07Z")

</div>

How many shards are we looking at?

---

<div class="post-metadata">

**Author:** ![John\_ax](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@John\_ax](https://discuss.elastic.co/u/John_ax)\
**Post date:** [September 6, 2017, 9:08am UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/3 "2017-09-06T09:08:09Z")

</div>

68 indices, 21 shards each indices cause i have 21 hot nodes with SSD, and the store size is 1.5T everyday.(I found one index is huge:1T above.) , I also have 5 cold node with 48T hdd hard disk.

I just restart the cold nodes ,now I got the marvel monitor data back with some unassigned indices, where did the unassigned indices store temporarily? will i loss data cause the restart and having replication num 0?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 6, 2017, 9:14am UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/4 "2017-09-06T09:14:57Z")

</div>

That's quite a lot.

Are you doing the move all at once? Are you shrinking before at all\>?

---

<div class="post-metadata">

**Author:** ![John\_ax](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@John\_ax](https://discuss.elastic.co/u/John_ax)\
**Post date:** [September 6, 2017, 10:34am UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/5 "2017-09-06T10:34:06Z")

</div>

what do you mean shrinking ? I move the indices by the curator , will forcemerge after moving.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 6, 2017, 10:39am UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/6 "2017-09-06T10:39:18Z")

</div>

Is that 68 indices being generated per day, each with 21 primary shards and possibly the same number of replica shards? What is your average and maximum shard size? What is your total retention period?

---

<div class="post-metadata">

**Author:** ![John\_ax](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@John\_ax](https://discuss.elastic.co/u/John_ax)\
**Post date:** [September 7, 2017, 2:52pm UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/7 "2017-09-07T14:52:54Z")

</div>

the indices were generated everyday, but replica shards is 0. as i said before , the store size is 1.5T everyday.(I found one index is huge:1T above.) ,

can you tell me what will cause the long time gc even the jvm memory setting is 30gb ?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 7, 2017, 2:54pm UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/8 "2017-09-07T14:54:50Z")

</div>

It looks like you are suffering from heap pressure on the cold nodes. How many shards do you have per cold node? How much data do you have on each cold node?

It seems like your average shard size is just around 1GB on the hot nodes. This is quite small and will result in more overhead than if you used larger shards. An average shard size around 20GB or 30GB is not uncommon. I would therefore recommend dramatically reducing the number of shards for smaller indices as they probably waste a lot of resources.

---

<div class="post-metadata">

**Author:** ![John\_ax](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@John\_ax](https://discuss.elastic.co/u/John_ax)\
**Post date:** [September 8, 2017, 1:58am UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/9 "2017-09-08T01:58:28Z")

</div>

you mean too large shards moved from hot nodes to cold nodes will cause long time gc ,or jvm oom?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 8, 2017, 5:08am UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/10 "2017-09-08T05:08:16Z")

</div>

No, I suspect that having too many small shards on the cold nodes is what is causing problems.

---

<div class="post-metadata">

**Author:** ![John\_ax](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@John\_ax](https://discuss.elastic.co/u/John_ax)\
**Post date:** [September 8, 2017, 6:00am UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/11 "2017-09-08T06:00:43Z")

</div>

help ,help , I think it's the same root cause ,now the cluster is dead, the master keep telling :

```json
[2017-09-08T13:58:14,426][DEBUG][o.e.a.a.i.m.p.TransportPutMappingAction] [BrSJ2NM] failed to put mappings on indices [[[pdc_20170908/37e0NIffQ6-Y_Cr5Bau5qQ]]], type [log]
org.elasticsearch.cluster.metadata.ProcessClusterEventTimeoutException: failed to process cluster event (put-mapping) within 30s
        at org.elasticsearch.cluster.service.ClusterService$ClusterServiceTaskBatcher.lambda$null$0(ClusterService.java:255) [elasticsearch-5.5.0.jar:5.5.0]
        at org.elasticsearch.cluster.service.ClusterService$ClusterServiceTaskBatcher$$Lambda$2434/1961004673.accept(Unknown Source) [elasticsearch-5.5.0.jar:5.5.0]
        at java.util.ArrayList.forEach(ArrayList.java:1249) [?:1.8.0_40]
        at org.elasticsearch.cluster.service.ClusterService$ClusterServiceTaskBatcher.lambda$onTimeout$1(ClusterService.java:254) [elasticsearch-5.5.0.jar:5.5.0]
        at org.elasticsearch.cluster.service.ClusterService$ClusterServiceTaskBatcher$$Lambda$2433/1790943430.run(Unknown Source) [elasticsearch-5.5.0.jar:5.5.0]
        at org.elasticsearch.common.util.concurrent.ThreadContext$ContextPreservingRunnable.run(ThreadContext.java:569) [elasticsearch-5.5.0.jar:5.5.0]
        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) [?:1.8.0_40]
        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [?:1.8.0_40]
        at java.lang.Thread.run(Thread.java:745) [?:1.8.0_40]

```

how to recovery the cluster ASAP?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 19, 2017, 6:01am UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/12 "2017-09-19T06:01:00Z")

</div>

Having too many indices and shards is such a common problem that I created a [blog post](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster) with some guidance and best practices.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 17, 2017, 6:01am UTC](https://discuss.elastic.co/t/transfering-indices-marked-cold-to-the-cold-nodes-cause-gc-overhead-and-other-issues/99529/13 "2017-10-17T06:01:26Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
