# S3 snapshotting timed out requiring Elasticsearch process restart - 6.3.0

**URL:** <https://discuss.elastic.co/t/s3-snapshotting-timed-out-requiring-elasticsearch-process-restart-6-3-0/166205>\
**Category:** Elasticsearch\
**Created:** [January 29, 2019, 4:12pm UTC](https://discuss.elastic.co/t/s3-snapshotting-timed-out-requiring-elasticsearch-process-restart-6-3-0/166205 "2019-01-29T16:12:49Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Martin\_Brennan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/martin_brennan/32/40339_2.png) [@Martin\_Brennan](https://discuss.elastic.co/u/Martin_Brennan)\
**Post date:** [January 29, 2019, 4:12pm UTC](https://discuss.elastic.co/t/s3-snapshotting-timed-out-requiring-elasticsearch-process-restart-6-3-0/166205/1 "2019-01-29T16:12:49Z")

</div>

We ran into an issue where we were unable to read or write snapshots to s3. The error is listed below. We mitigated the issue by restarting the Elasticsearch process on the master node.

On an independent cluster we also had a similar timeout problem but this time it was reported by the data nodes and restarting the data nodes fixed the problem on this cluster.

```auto
curl -s localhost:9200/_cat/snapshots/repo_name
{
  "error": {
    "caused_by": {
      "caused_by": {
        "reason": "connect timed out",
        "type": "i_o_exception"
      },
      "reason": "Connect to repo_name.s3.amazonaws.com:443 [repo_name.s3.amazonaws.com/52.217.0.51] failed: connect timed out",
      "type": "i_o_exception"
    },
    "reason": "sdk_client_exception: Unable to execute HTTP request: Connect to intersearch-snapshots-usersearch-v11-0.s3.amazonaws.com:443 [repo_name.s3.amazonaws.com/52.217.0.51] failed: connect timed out",
    "root_cause": [
      {
        "reason": "sdk_client_exception: Unable to execute HTTP request: Connect to repo_name.s3.amazonaws.com:443 [repo_name.s3.amazonaws.com/52.217.0.51] failed: connect timed out",
        "type": "sdk_client_exception"
      }
    ],
    "type": "sdk_client_exception"
  },
  "status": 500
}

```

Restarting the elasticsearch process with `/etc/init.d/elasticsearch restart` fixed the issue.

```auto
curl -s localhost:9200/_cat/snapshots/repo_name
repo_name_2018-12-26t04:00:06z SUCCESS 1545796807 04:00:07 1545796837 04:00:37 30.6s 4 40 0 40

```

The clusters this happened on were running Elasticsearch version 6.3.0

Our hypothesis is that there was a cached DNS record somewhere in the java process and when something on the AWS side changed this cached record was no longer valid.

This issue might be related - [Cannot list snapshots in S3 repository (previously working fine)](https://discuss.elastic.co/t/cannot-list-snapshots-in-s3-repository-previously-working-fine/104249/3)

---

<div class="post-metadata">

**Author:** ![CathalCoffey](https://avatars.discourse-cdn.com/v4/letter/c/3e96dc/32.png) [@CathalCoffey](https://discuss.elastic.co/u/CathalCoffey)\
**Post date:** [February 2, 2019, 7:00pm UTC](https://discuss.elastic.co/t/s3-snapshotting-timed-out-requiring-elasticsearch-process-restart-6-3-0/166205/2 "2019-02-02T19:00:42Z")

</div>

I'm from this same company as Martin\_Brennan.

We just experienced this problem again today on multiple clusters. Our search clusters are quite large (40 data-nodes) and under a lot of ingestion and query load so preforming a rolling restart of all nodes is not something we are keen to do. So far we have just be restarting the nodes that have exhibited this timeout issue during snapshot creation.

For example, today we saw the following error during snapshotting

```auto
{
  "duration_in_millis": 345001,
  "end_time": "2019-02-02T00:05:48.480Z",
  "end_time_in_millis": 1549065948480,
  "failures": [
    {
      "index": "redacted",
      "index_uuid": "redacted",
      "node_id": "lxtigl9JRvm1dLX0RurmUg",
      "reason": "IndexShardSnapshotFailedException[com.amazonaws.SdkClientException: Unable to execute HTTP request: Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: SdkClientException[Unable to execute HTTP request: Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: ConnectTimeoutException[Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: SocketTimeoutException[connect timed out]; ",
      "shard_id": 11,
      "status": "INTERNAL_SERVER_ERROR"
    },
    {
      "index": "redacted",
      "index_uuid": "redacted",
      "node_id": "lxtigl9JRvm1dLX0RurmUg",
      "reason": "IndexShardSnapshotFailedException[com.amazonaws.SdkClientException: Unable to execute HTTP request: Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: SdkClientException[Unable to execute HTTP request: Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: ConnectTimeoutException[Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: SocketTimeoutException[connect timed out]; ",
      "shard_id": 13,
      "status": "INTERNAL_SERVER_ERROR"
    },
    {
      "index": "redacted",
      "index_uuid": "redacted",
      "node_id": "lxtigl9JRvm1dLX0RurmUg",
      "reason": "IndexShardSnapshotFailedException[com.amazonaws.SdkClientException: Unable to execute HTTP request: Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: SdkClientException[Unable to execute HTTP request: Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: ConnectTimeoutException[Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: SocketTimeoutException[connect timed out]; ",
      "shard_id": 11,
      "status": "INTERNAL_SERVER_ERROR"
    },
    {
      "index": "redacted",
      "index_uuid": "redacted",
      "node_id": "lxtigl9JRvm1dLX0RurmUg",
      "reason": "IndexShardSnapshotFailedException[com.amazonaws.SdkClientException: Unable to execute HTTP request: Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: SdkClientException[Unable to execute HTTP request: Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: ConnectTimeoutException[Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: SocketTimeoutException[connect timed out]; ",
      "shard_id": 14,
      "status": "INTERNAL_SERVER_ERROR"
    },
    {
      "index": "redacted",
      "index_uuid": "redacted",
      "node_id": "lxtigl9JRvm1dLX0RurmUg",
      "reason": "IndexShardSnapshotFailedException[com.amazonaws.SdkClientException: Unable to execute HTTP request: Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: SdkClientException[Unable to execute HTTP request: Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: ConnectTimeoutException[Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: SocketTimeoutException[connect timed out]; ",
      "shard_id": 12,
      "status": "INTERNAL_SERVER_ERROR"
    },
    {
      "index": "redacted",
      "index_uuid": "redacted",
      "node_id": "lxtigl9JRvm1dLX0RurmUg",
      "reason": "IndexShardSnapshotFailedException[com.amazonaws.SdkClientException: Unable to execute HTTP request: Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: SdkClientException[Unable to execute HTTP request: Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: ConnectTimeoutException[Connect to redacted.s3.amazonaws.com:443 [redacted.s3.amazonaws.com/52.216.176.27] failed: connect timed out]; nested: SocketTimeoutException[connect timed out]; ",
      "shard_id": 19,
      "status": "INTERNAL_SERVER_ERROR"
    }
  ],
  "include_global_state": true,
  "indices": [
    "redacted",
    "redacted",
    "redacted",
    "redacted",
    "redacted",
    "redacted",
    "redacted",
    "redacted",
    "redacted",
    "redacted"
  ],
  "shards": {
    "failed": 6,
    "successful": 194,
    "total": 200
  },
  "snapshot": "redacted_2019-02-02t00:00:03z",
  "start_time": "2019-02-02T00:00:03.479Z",
  "start_time_in_millis": 1549065603479,
  "state": "PARTIAL",
  "uuid": "mbCeCSWuTXWkauPLgNq3Hg",
  "version": "6.3.0",
  "version_id": 6030099
}

```

**Note:** The 6 shards that failed to snapshot were all on the same host `lxtigl9JRvm1dLX0RurmUg`. Every subsequent attempt to create a snapshot results in the exact same error, with only this single node failing with a timeout to S3.

Restarting the Elasticsearch process on this host, waiting for the cluster to go green, and then snapshotting again is successful.

This is pretty serious problem for us, we need to be able to reliably take snapshots every 24 hours. If there is more information you need us to provide in order to get this triaged, please let us know.

---

<div class="post-metadata">

**Author:** ![CathalCoffey](https://avatars.discourse-cdn.com/v4/letter/c/3e96dc/32.png) [@CathalCoffey](https://discuss.elastic.co/u/CathalCoffey)\
**Post date:** [February 2, 2019, 8:29pm UTC](https://discuss.elastic.co/t/s3-snapshotting-timed-out-requiring-elasticsearch-process-restart-6-3-0/166205/3 "2019-02-02T20:29:07Z")

</div>

> The Java virtual machine (JVM) caches DNS name lookups. When the JVM resolves a hostname to an IP address, it caches the IP address for a specified period of time, known as the _time-to-live_ (TTL).
> 
> Because AWS resources use DNS name entries that occasionally change, we recommend that you configure your JVM with a TTL value of no more than 60 seconds. This ensures that when a resource's IP address changes, your application will be able to receive and use the resource's new IP address by requerying the DNS.
> 
> On some Java configurations, the JVM default TTL is set so that it will _never_ refresh DNS entries until the JVM is restarted. Thus, if the IP address for an AWS resource changes while your application is still running, it won't be able to use that resource until you _manually restart_ the JVM and the cached IP information is refreshed. In this case, it's crucial to set the JVM's TTL so that it will periodically refresh its cached IP information.

This ^ seems likely to be the root cause:

> **[Setting the JVM TTL for DNS Name Lookups - AWS SDK for Java 1.x](https://docs.aws.amazon.com/sdk-for-java/v1/developer-guide/java-dg-jvm-ttl.html)**
>
> How to set Java virtual machine (JVM) for DNS name lookups using the AWS SDK for Java.

Should you guys be setting this TTL somewhere in this repo? [https://github.com/elastic/elasticsearch/tree/master/plugins/repository-s3](https://github.com/elastic/elasticsearch/tree/master/plugins/repository-s3)

---

<div class="post-metadata">

**Author:** ![CathalCoffey](https://avatars.discourse-cdn.com/v4/letter/c/3e96dc/32.png) [@CathalCoffey](https://discuss.elastic.co/u/CathalCoffey)\
**Post date:** [February 2, 2019, 8:42pm UTC](https://discuss.elastic.co/t/s3-snapshotting-timed-out-requiring-elasticsearch-process-restart-6-3-0/166205/4 "2019-02-02T20:42:34Z")

</div>

Just in case it's useful, according to the Elasticsearch log, we are **NOT** passing any JVM arg that could be overriding any **ttl** defaults.

> [2019-02-02T14:33:52,392][INFO][o.e.n.Node] [es-master-3.localdomain] version[6.3.0], pid[27981], build[default/rpm/424e937/2018-06-11T23:38:03.357887Z], OS[Linux/4.14.88-72.76.amzn1.x86\_64/amd64], JVM[Oracle Corporation/Java HotSpot(TM) 64-Bit Server VM/1.8.0\_171/25.171-b11]

> [2019-02-02T14:33:52,392][INFO][o.e.n.Node] [es-master-3.localdomain] JVM arguments [-Xms7789m, -Xmx7789m, -XX:+UseConcMarkSweepGC, -XX:CMSInitiatingOccupancyFraction=75, -XX:+UseCMSInitiatingOccupancyOnly, -XX:+AlwaysPreTouch, -Xss1m, -Djava.awt.headless=true, -Dfile.encoding=UTF-8, -Djna.nosys=true, -XX:-OmitStackTraceInFastThrow, -Dio.netty.noUnsafe=true, -Dio.netty.noKeySetOptimization=true, -Dio.netty.recycler.maxCapacityPerThread=0, -Dlog4j.shutdownHookEnabled=false, -Dlog4j2.disable.jmx=true, -XX:+PrintGCDetails, -XX:+PrintGCDateStamps, -XX:+PrintTenuringDistribution, -XX:+PrintGCApplicationStoppedTime, -Xloggc:logs/gc.log, -XX:+UseGCLogFileRotation, -XX:NumberOfGCLogFiles=32, -XX:GCLogFileSize=64m, -Des.index.memory.max\_index\_buffer\_size=10240mb, -Des.path.home=/usr/share/elasticsearch, -Des.path.conf=/etc/elasticsearch, -Des.distribution.flavor=default, -Des.distribution.type=rpm]

---

<div class="post-metadata">

**Author:** ![jtyoung](https://avatars.discourse-cdn.com/v4/letter/j/f19dbf/32.png) [@jtyoung](https://discuss.elastic.co/u/jtyoung)\
**Post date:** [February 18, 2019, 3:54pm UTC](https://discuss.elastic.co/t/s3-snapshotting-timed-out-requiring-elasticsearch-process-restart-6-3-0/166205/5 "2019-02-18T15:54:07Z")

</div>

I am experiencing the same issue here, on a 6.5.0 cluster. As mentioned above, restarting the node that gave the errors fixed the issue for the time being.

> Just in case it's useful, according to the Elasticsearch log, we are **NOT** passing any JVM arg that could be overriding any **ttl** defaults.

Same here.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [February 18, 2019, 4:22pm UTC](https://discuss.elastic.co/t/s3-snapshotting-timed-out-requiring-elasticsearch-process-restart-6-3-0/166205/6 "2019-02-18T16:22:36Z")

</div>

It does indeed sound like a DNS caching issue. The [manual for 6.5 and earlier](https://www.elastic.co/guide/en/elasticsearch/reference/6.5/networkaddress-cache-ttl.html) says:

> Elasticsearch runs with a security manager in place. With a security manager in place, the JVM defaults to caching positive hostname resolutions indefinitely. If your Elasticsearch nodes rely on DNS in an environment where DNS resolutions vary with time (e.g., for node-to-node discovery) then you might want to modify the default JVM behavior. This can be modified by adding[`networkaddress.cache.ttl=<timeout>`](http://docs.oracle.com/javase/8/docs/technotes/guides/net/properties.html) to your [Java security policy](http://docs.oracle.com/javase/8/docs/technotes/guides/security/PolicyFiles.html). Any hosts that fail to resolve will be logged. Note also that with the Java security manager in place, the JVM defaults to caching negative hostname resolutions for ten seconds. This can be modified by adding[`networkaddress.cache.negative.ttl=<timeout>`](http://docs.oracle.com/javase/8/docs/technotes/guides/net/properties.html) to your [Java security policy](http://docs.oracle.com/javase/8/docs/technotes/guides/security/PolicyFiles.html).

This was a deliberate decision, made some time ago, but we have reflected on its consequences more recently and changed our mind so that [in 6.6 by default we now do not cache DNS resolutions forever](https://www.elastic.co/guide/en/elasticsearch/reference/6.6/networkaddress-cache-ttl.html):

> Elasticsearch runs with a security manager in place. With a security manager in place, the JVM defaults to caching positive hostname resolutions indefinitely and defaults to caching negative hostname resolutions for ten seconds. Elasticsearch overrides this behavior with default values to cache positive lookups for sixty seconds, and to cache negative lookups for ten seconds. These values should be suitable for most environments, including environments where DNS resolutions vary with time. If not, you can edit the values `es.networkaddress.cache.ttl` and `es.networkaddress.cache.negative.ttl` in the [JVM options](https://www.elastic.co/guide/en/elasticsearch/reference/6.6/jvm-options.html). Note that the values[`networkaddress.cache.ttl=<timeout>`](http://docs.oracle.com/javase/8/docs/technotes/guides/net/properties.html) and [`networkaddress.cache.negative.ttl=<timeout>`](http://docs.oracle.com/javase/8/docs/technotes/guides/net/properties.html) in the[Java security policy](http://docs.oracle.com/javase/8/docs/technotes/guides/security/PolicyFiles.html) are ignored by Elasticsearch unless you remove the settings for `es.networkaddress.cache.ttl` and `es.networkaddress.cache.negative.ttl` .

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 18, 2019, 4:22pm UTC](https://discuss.elastic.co/t/s3-snapshotting-timed-out-requiring-elasticsearch-process-restart-6-3-0/166205/7 "2019-03-18T16:22:45Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
