# \[ES v7.5.1\] Speed up snapshot restore from GCS Snapshot Repository

**URL:** <https://discuss.elastic.co/t/es-v7-5-1-speed-up-snapshot-restore-from-gcs-snapshot-repository/278974>\
**Category:** Elasticsearch\
**Tags:** snapshot-and-restore\
**Created:** [July 17, 2021, 1:40pm UTC](https://discuss.elastic.co/t/es-v7-5-1-speed-up-snapshot-restore-from-gcs-snapshot-repository/278974 "2021-07-17T13:40:19Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![amitsh728](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amitsh728/32/91780_2.png) [@amitsh728](https://discuss.elastic.co/u/amitsh728)\
**Post date:** [July 17, 2021, 1:40pm UTC](https://discuss.elastic.co/t/es-v7-5-1-speed-up-snapshot-restore-from-gcs-snapshot-repository/278974/1 "2021-07-17T13:40:19Z")

</div>

Hello,  
We are trying to speed up our snapshot restoration speed in our ES cluster hosted on GCP Compute instances.

**TL;DR** :

- Current Performance: 56 MBps per data node
- Infra Capable of 500 MBps
- We want to improve our restore speed up to the maximum disk throughput (no throttling from infrastructure).

**Infra Details** :

- 3 Master | 10 Data Nodes
- Each node has 8core/16gb config (heap size: 8gb)
- Each Data instance supports max 15k disk IOPS (500 MBps Throughput)

Current Restore Performance: 4.5 Gbps on the whole cluster (56 MBps per node).

We are currently getting 10% of the total disk write throughput. We are looking for options to improve it.

**Note:** Have already tested any infra-related throttling. Using `gsutil -m`, we saw the download speed reach to 450 MBps on one of the data nodes.

**What we have already tried:**

1. Setting `max_restore_bytes_per_sec` to `0` in our gcs snapshot repository.
2. Setting `indices.recovery.max_bytes_per_sec` to `0`.

We are unable to figure out the config that is throttling the network performance.

* * *

**Update:**

We tried to increase the number of data nodes, to check if the throttling is on some data nodes:

Changed Data nodes count from 10 to 20.  
**Result:** Speed still throttled at 4.5 Gbps.

---

<div class="post-metadata">

**Author:** ![Akshay\_KN](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/akshay_kn/32/76976_2.png) [@Akshay\_KN](https://discuss.elastic.co/u/Akshay_KN)\
**Post date:** [July 20, 2021, 10:02am UTC](https://discuss.elastic.co/t/es-v7-5-1-speed-up-snapshot-restore-from-gcs-snapshot-repository/278974/2 "2021-07-20T10:02:10Z")

</div>

Hi Team,

Please note the below config changes that we have already tried but not getting the expected speed.

**At Cluster Level**

```auto
indices.recovery.max_bytes_per_sec: "10gb"
indices.recovery.max_concurrent_file_chunks: 5

```

**At Index Level:**

```auto
"refresh_interval" : "-1",
"merge.scheduler.max_thread_count" : "1"

```

**At Snapshot Repository Level:**

```auto
"max_restore_bytes_per_sec": "10gb",
"max_snapshot_bytes_per_sec": "10gb"

```

---

<div class="post-metadata">

**Author:** ![xeraa](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xeraa/32/48181_2.png) [@xeraa](https://discuss.elastic.co/u/xeraa)\
**Post date:** [July 20, 2021, 11:52am UTC](https://discuss.elastic.co/t/es-v7-5-1-speed-up-snapshot-restore-from-gcs-snapshot-repository/278974/3 "2021-07-20T11:52:13Z")

</div>

If you're not already doing it, I would set the replicas (temporarily) to 0: [Restore a snapshot | Elasticsearch Guide [7.13] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/snapshots-restore-snapshot.html#change-index-settings-during-restore)

Also I'm not sure how to read the "Infra Capable of 500 MBps". Isn't that what you're getting with "Current Restore Performance: 4.5 Gbps"?

---

<div class="post-metadata">

**Author:** ![amitsh728](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amitsh728/32/91780_2.png) [@amitsh728](https://discuss.elastic.co/u/amitsh728)\
**Post date:** [July 20, 2021, 12:32pm UTC](https://discuss.elastic.co/t/es-v7-5-1-speed-up-snapshot-restore-from-gcs-snapshot-repository/278974/4 "2021-07-20T12:32:13Z")

</div>

Hi @xerrad Thanks for the reply...  
For the "replicas (temporarily) to 0": Yes we are setting it (But that's not helping us with the speed).

"Infra Capable of 500 MBps" -\> That is for one data node.

To clarify (Using Gbps for all the stats):

- Per node, the infra is capable to get 5Gbps.
- For 10 nodes, the infra is capable to get 5Gbps x 10.
- What we are getting for 10 nodes: 4.5 Gbps.

> 4.5 Gbps is over 10 nodes. We should be able to get about 50Gbps over 10 nodes (our need is to get at least 25Gbps).

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 20, 2021, 12:49pm UTC](https://discuss.elastic.co/t/es-v7-5-1-speed-up-snapshot-restore-from-gcs-snapshot-repository/278974/5 "2021-07-20T12:49:31Z")

</div>

haw many shards are you recovering per node? Have you tried increasing `cluster.routing.allocation.node_concurrent_recoveries` to increase the level of papallelism as described [here](https://www.elastic.co/guide/en/elasticsearch/reference/7.13/modules-cluster.html#cluster-shard-allocation-settings)?

---

<div class="post-metadata">

**Author:** ![amitsh728](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amitsh728/32/91780_2.png) [@amitsh728](https://discuss.elastic.co/u/amitsh728)\
**Post date:** [July 20, 2021, 1:22pm UTC](https://discuss.elastic.co/t/es-v7-5-1-speed-up-snapshot-restore-from-gcs-snapshot-repository/278974/6 "2021-07-20T13:22:26Z")

</div>

Hi @Christian_Dahlqvist Thanks for replying.

We have total 32 shared in our index, and 10 nodes. So I can see that 3 shards are going on one node (apart from 2 nodes, where there are 4 shared).

We did try to set `cluster.routing.allocation.node_concurrent_recoveries` to 4 on the cluster (Updated elasticsearch.yml file on all nodes + restarted service on all), although that didn't make any difference on the restoration performance.

We also tried the below cluster configs (Just to see if it impacts speed):

```auto
  thread_pool.snapshot.core: 4
  thread_pool.snapshot.max: 8
  indices.recovery.max_concurrent_file_chunks: 5
  thread_pool.write.queue_size: 10000
  thread_pool.get.queue_size: 10000
  transport.connections_per_node.reg: 50
  transport.connections_per_node.bulk: 50
  thread_pool.fetch_shard_started.core: 4
  thread_pool.fetch_shard_store.core: 4
  cluster.routing.allocation.node_concurrent_recoveries: 4
  indices.query.bool.max_clause_count: 2048

```

But we didn't saw any difference.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 20, 2021, 1:44pm UTC](https://discuss.elastic.co/t/es-v7-5-1-speed-up-snapshot-restore-from-gcs-snapshot-repository/278974/7 "2021-07-20T13:44:58Z")

</div>

What is the size of the shards? Does GCS impose any performance limits for this type of data which includes a large number of sometimes small files?

---

<div class="post-metadata">

**Author:** ![amitsh728](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amitsh728/32/91780_2.png) [@amitsh728](https://discuss.elastic.co/u/amitsh728)\
**Post date:** [July 21, 2021, 5:52am UTC](https://discuss.elastic.co/t/es-v7-5-1-speed-up-snapshot-restore-from-gcs-snapshot-repository/278974/8 "2021-07-21T05:52:41Z")

</div>

Hi @Christian_Dahlqvist , each shard is about 100GB.

---

<div class="post-metadata">

**Author:** ![amitsh728](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amitsh728/32/91780_2.png) [@amitsh728](https://discuss.elastic.co/u/amitsh728)\
**Post date:** [July 21, 2021, 5:54am UTC](https://discuss.elastic.co/t/es-v7-5-1-speed-up-snapshot-restore-from-gcs-snapshot-repository/278974/9 "2021-07-21T05:54:44Z")

</div>

Update:  
After trying multiple options, we weren't able to find the root-cause of this.

We decided to use an updated version of ES (version: 7.10.2).

We are now able to reach 33gbps speed with the updated cluster.

Intrestingly, we are using similar configs (similar ansible playbooks) for both the version, but with version 7.10.2, we were able to reach the target speed.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 18, 2021, 5:55am UTC](https://discuss.elastic.co/t/es-v7-5-1-speed-up-snapshot-restore-from-gcs-snapshot-repository/278974/10 "2021-08-18T05:55:10Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
