# Snapshot problems with Amazon S3

**URL:** <https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765>\
**Category:** Elasticsearch\
**Created:** [December 14, 2017, 12:22pm UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765 "2017-12-14T12:22:09Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![baltendo](https://avatars.discourse-cdn.com/v4/letter/b/9de053/32.png) [@baltendo](https://discuss.elastic.co/u/baltendo)\
**Post date:** [December 14, 2017, 12:22pm UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/1 "2017-12-14T12:22:09Z")

</div>

Hi!

Elasticsearch: 5.4.2  
Cloud: AWS  
OS: Amazon Linux

Since 17th of November we get failures when our daily backup snapshots are created:

```auto
"failures": [
        {
          "index": "user-reviews",
          "index_uuid": "user-reviews",
          "shard_id": 1,
          "reason": "IndexShardSnapshotFailedException[com.amazonaws.AmazonClientException: Unable to execute HTTP request: connect timed out]; nested: AmazonClientException[Unable to execute HTTP request: connect timed out]; nested: SocketTimeoutException[connect timed out]; ",
          "node_id": "Ao5nMnDNSrmNarISSowGgA",
          "status": "INTERNAL_SERVER_ERROR"
        },
        ...
]

```

The indices which are throwing those failures are changing, so we don't see any pattern there.

We then tried to create a new snapshot repository on S3 but we got this error:

```auto
{
  "error": {
    "root_cause": [
      {
        "type": "repository_verification_exception",
        "reason": "[repo-3] path is not accessible on master node"
      }
    ],
    "type": "repository_verification_exception",
    "reason": "[repo-3] path is not accessible on master node",
    "caused_by": {
      "type": "i_o_exception",
      "reason": "Unable to upload object tests-2lCYsX97Q8-qz15qPdeS0Q/master.dat-temp",
      "caused_by": {
        "type": "amazon_s3_exception",
        "reason": "amazon_s3_exception: The request signature we calculated does not match the signature you provided. Check your key and signing method. (Service: Amazon S3; Status Code: 403; Error Code: SignatureDoesNotMatch; Request ID: 4398DBFF11FFA4E7)"
      }
    }
  },
  "status": 500
}

```

This leads us to the assumption that something has been changed related to Amazon S3.

We saw that the aws sdk is upgraded in version 5.6.5, could this be a solution to our problem?

Kind Regards,  
Bernhard

---

<div class="post-metadata">

**Author:** ![mujtabahussain](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mujtabahussain/32/17514_2.png) [@mujtabahussain](https://discuss.elastic.co/u/mujtabahussain)\
**Post date:** [December 15, 2017, 5:31am UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/2 "2017-12-15T05:31:15Z")

</div>

Hey!

- Can you PUT a file on the S3 bucket from the CLI from any of the nodes?
- Have you double checked what IAM permissions you have given to the nodes to be able to access S3?

Also, this may not be related, but last time I had this error

> [@](#):
>
> The request signature we calculated does not match the signature you provided

was due to NTP issues. ☹

But confirm the first two things 🙂

---

<div class="post-metadata">

**Author:** ![baltendo](https://avatars.discourse-cdn.com/v4/letter/b/9de053/32.png) [@baltendo](https://discuss.elastic.co/u/baltendo)\
**Post date:** [December 15, 2017, 8:41am UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/3 "2017-12-15T08:41:07Z")

</div>

Hi!

I tested uploading a file with AWS CLI 1.11.83 and 1.10.67 and it is working fine.  
I put the file in the same bucket as Elasticsearch should put the snapshots, which means the IAM permissions are fine as well.

So I guess the first to points are confirmed?

Kind Regards,  
Bernhard

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 15, 2017, 8:54am UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/4 "2017-12-15T08:54:46Z")

</div>

Can you also remove the file you uploaded from the CLI?

---

<div class="post-metadata">

**Author:** ![baltendo](https://avatars.discourse-cdn.com/v4/letter/b/9de053/32.png) [@baltendo](https://discuss.elastic.co/u/baltendo)\
**Post date:** [December 15, 2017, 9:09am UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/5 "2017-12-15T09:09:53Z")

</div>

Hi!

Removing the file using the CLI also works.

I just tried again to create a snapshot repository but from a test cluster with a single node.  
In scenario 1 the cluster is on 5.4.2 like our production cluster.  
In scenario 2 the cluster is on 5.6.5 with upgraded AWS libraries.  
Here are the errors:

```auto
// 5.4.2
{
	"error": {
		"root_cause": [{
			"type": "repository_verification_exception",
			"reason": "[cluster] path is not accessible on master node"
		}],
		"type": "repository_verification_exception",
		"reason": "[cluster] path is not accessible on master node",
		"caused_by": {
			"type": "i_o_exception",
			"reason": "Unable to upload object tests-WYj-gzdiQTqWXnPwxeNllQ/master.dat-temp",
			"caused_by": {
				"type": "amazon_s3_exception",
				"reason": "The request signature we calculated does not match the signature you provided. Check your key and signing method. (Service: Amazon S3; Status Code: 403; Error Code: SignatureDoesNotMatch; Request ID: 1BDDD42316F78C63)"
			}
		}
	},
	"status": 500
}

```

```auto
// 5.6.5
{
	"error": {
		"root_cause": [{
			"type": "repository_exception",
			"reason": "[cluster] failed to create repository"
		}],
		"type": "repository_exception",
		"reason": "[cluster] failed to create repository",
		"caused_by": {
			"type": "amazon_s3_exception",
			"reason": "Method Not Allowed (Service: Amazon S3; Status Code: 405; Error Code: 405 Method Not Allowed; Request ID: 3FAEFF31D2A9A406; S3 Extended Request ID: GXfAJDL54/WicbXdIJcMJI7RSl17eGc6VexJnJpvavlGSS6ByGLOqi1FnhiD1cMn3oxplCkEkcE=)"
		}
	},
	"status": 500
}

```

Kind Regards,  
Bernhard

---

<div class="post-metadata">

**Author:** ![mujtabahussain](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mujtabahussain/32/17514_2.png) [@mujtabahussain](https://discuss.elastic.co/u/mujtabahussain)\
**Post date:** [December 17, 2017, 11:55pm UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/6 "2017-12-17T23:55:58Z")

</div>

Could you show us the IAM permissions you have allowed the cluster for S3?

---

<div class="post-metadata">

**Author:** ![baltendo](https://avatars.discourse-cdn.com/v4/letter/b/9de053/32.png) [@baltendo](https://discuss.elastic.co/u/baltendo)\
**Post date:** [December 18, 2017, 7:23am UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/7 "2017-12-18T07:23:16Z")

</div>

This is our IAM policy:

```auto
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Action": "s3:*",
            "Effect": "Allow",
            "Resource": [
                "arn:aws:s3:::elasticsearch-5-snapshots",
                "arn:aws:s3:::elasticsearch-5-snapshots/*",
                "arn:aws:s3:::elasticsearch-5-archive",
                "arn:aws:s3:::elasticsearch-5-archive/*"
            ]
        },
        {
            "Effect": "Allow",
            "Action": "EC2:Describe*",
            "Resource": "*"
        },
        {
            "Effect": "Allow",
            "Action": "cloudwatch:DeleteAlarms",
            "Resource": "*"
        },
        {
            "Effect": "Allow",
            "Action": "cloudwatch:PutMetricAlarm",
            "Resource": "*"
        }
    ]
}

```

---

<div class="post-metadata">

**Author:** ![mujtabahussain](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mujtabahussain/32/17514_2.png) [@mujtabahussain](https://discuss.elastic.co/u/mujtabahussain)\
**Post date:** [December 18, 2017, 11:01pm UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/8 "2017-12-18T23:01:37Z")

</div>

I highly recommend raising an AWS support issue as well via the console. !

---

<div class="post-metadata">

**Author:** ![baltendo](https://avatars.discourse-cdn.com/v4/letter/b/9de053/32.png) [@baltendo](https://discuss.elastic.co/u/baltendo)\
**Post date:** [December 21, 2017, 1:29pm UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/9 "2017-12-21T13:29:31Z")

</div>

Hi!

I noticed that one particular node was mostly causing the snapshot failures.  
Today I took the node out of the cluster, reinstalled ES 5.4.2 and rebooted the machine.  
I started another snapshot and it was successful!

In case this happens again I will try just a reboot first to see if this already solves the issue.

Thanks for your help!

Kind Regards,  
Bernhard

---

<div class="post-metadata">

**Author:** ![mujtabahussain](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mujtabahussain/32/17514_2.png) [@mujtabahussain](https://discuss.elastic.co/u/mujtabahussain)\
**Post date:** [December 21, 2017, 11:28pm UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/10 "2017-12-21T23:28:27Z")

</div>

This is my suspicion about what was happening.

AWS SDK's sign any requests that emerge from one resource to another, so the destination can ensure that the request is coming from one of the valid resources. One of the things they use to sign the request is the time at which the request was made. This is where this comes in:

> [@](#):
>
> If you use the AWS CLI or an AWS SDK to make requests from your instance, these tools sign requests on your behalf. If your instance's date and time are not set correctly, the request can be rejected if the date in the signature does not match the date of the request.

taken from [here](https://aws.amazon.com/premiumsupport/knowledge-center/system-clock-drift-ubuntu/).

So if you follow the setup of [this document](https://aws.amazon.com/premiumsupport/knowledge-center/system-clock-drift-ubuntu/), hopefully you can atleast negate this issue.

This issue might also explain why it was happening only on one instance.

Again, this is my suspicion. If this issue re-emerges, try this first before rebooting and let us know. 🙂

Best of luck.

---

<div class="post-metadata">

**Author:** ![baltendo](https://avatars.discourse-cdn.com/v4/letter/b/9de053/32.png) [@baltendo](https://discuss.elastic.co/u/baltendo)\
**Post date:** [January 10, 2018, 8:03am UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/11 "2018-01-10T08:03:46Z")

</div>

The issue appears again and I wanted to follow your advice, but the links in your last answer are returning a 404.

---

<div class="post-metadata">

**Author:** ![mujtabahussain](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mujtabahussain/32/17514_2.png) [@mujtabahussain](https://discuss.elastic.co/u/mujtabahussain)\
**Post date:** [January 11, 2018, 11:52pm UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/12 "2018-01-11T23:52:35Z")

</div>

Yeah! Those pages from AWS Support Docs seem to have disappeared. ☹

Google `system clock drift AWS` and you are on your way 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 8, 2018, 11:52pm UTC](https://discuss.elastic.co/t/snapshot-problems-with-amazon-s3/111765/13 "2018-02-08T23:52:39Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
