# Removing second S3 repository causes "Connection Pool Shutdown" in 8.13.2

**URL:** <https://discuss.elastic.co/t/removing-second-s3-repository-causes-connection-pool-shutdown-in-8-13-2/364797>\
**Category:** Elasticsearch\
**Tags:** snapshot-and-restore\
**Created:** [August 12, 2024, 7:47pm UTC](https://discuss.elastic.co/t/removing-second-s3-repository-causes-connection-pool-shutdown-in-8-13-2/364797 "2024-08-12T19:47:29Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Doc\_Kaos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/doc_kaos/32/53671_2.png) [@Doc\_Kaos](https://discuss.elastic.co/u/Doc_Kaos)\
**Post date:** [August 12, 2024, 7:47pm UTC](https://discuss.elastic.co/t/removing-second-s3-repository-causes-connection-pool-shutdown-in-8-13-2/364797/1 "2024-08-12T19:47:30Z")

</div>

We upgraded to 8.13.2 after having issues in 8.10 with [Elasticsearch does not refresh AWS Web Identity Token file when changed on disk · Issue #101828 · elastic/elasticsearch · GitHub](https://github.com/elastic/elasticsearch/issues/101828)

We just had a cluster running 8.13.2 lose access to the S3 repository. Restarting all of the Master pods in K8s allowed us to _view_ the snapshots again and run a verification that shows all Data nodes cannot talk to the S3 repo. The nodes were up for 46 days before snapshots started failing.

We are around 20 ES clusters and so far this is the only one we've seen affected.

The Verify response after restarting all masters is:

```auto
{
  "name": "ResponseError",
  "message": "repository_verification_exception\n\tRoot causes:\n\t\trepository_verification_exception: [s3backup] [[3OY76lDlQEeXLa7ewplPgA, 'org.elasticsearch.transport.RemoteTransportException: [es-data-33][10.10.10.10:9300][internal:admin/repository/verify]'], [Same for the every data node ...

```

---

<div class="post-metadata">

**Author:** ![Doc\_Kaos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/doc_kaos/32/53671_2.png) [@Doc\_Kaos](https://discuss.elastic.co/u/Doc_Kaos)\
**Post date:** [August 20, 2024, 2:09pm UTC](https://discuss.elastic.co/t/removing-second-s3-repository-causes-connection-pool-shutdown-in-8-13-2/364797/2 "2024-08-20T14:09:50Z")

</div>

We discovered that having the Masters able to talk to the S3 Repo, but not the data nodes caused high load on the Masters for some reason. From typical 8% CPU, they ran at 75-100% CPU constantly.  
We restarted all data nodes and still were seeing high Masters CPU.  
Final resolution was to delete the S3 repo and recreate it ... suddenly everything dropped to expected levels again.  
Could be somethiing to do with:  
Elasticsearch Exporter collecting snapshot statistics (it was failing and OOM'ing until the repo was deleted)  
Bad state when Masters can see S3 repo, but data nodes can't

---

<div class="post-metadata">

**Author:** ![Doc\_Kaos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/doc_kaos/32/53671_2.png) [@Doc\_Kaos](https://discuss.elastic.co/u/Doc_Kaos)\
**Post date:** [August 20, 2024, 4:49pm UTC](https://discuss.elastic.co/t/removing-second-s3-repository-causes-connection-pool-shutdown-in-8-13-2/364797/3 "2024-08-20T16:49:46Z")

</div>

Halfway through a new snapshot, we again started receiving errors. After a chat with AWS we discovered the EKS pod began using the Instance Role instead of the Pod Role

---

<div class="post-metadata">

**Author:** ![Doc\_Kaos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/doc_kaos/32/53671_2.png) [@Doc\_Kaos](https://discuss.elastic.co/u/Doc_Kaos)\
**Post date:** [August 20, 2024, 6:51pm UTC](https://discuss.elastic.co/t/removing-second-s3-repository-causes-connection-pool-shutdown-in-8-13-2/364797/4 "2024-08-20T18:51:11Z")

</div>

This appears to be something to do with the AWS token refreshing, perhaps during a snapshot, that causes the S3 Client to shut down. It's not closed by ES so all future calls fail:

```auto
com.amazonaws.AmazonClientException: java.lang.IllegalStateException: Connection pool shut down\n\tat com.amazonaws.auth.RefreshableTask.refreshValue(RefreshableTask.java:303)\n\tat com.amazonaws.auth.RefreshableTask.blockingRefresh(RefreshableTask.java:251)\n\tat com.amazonaws.auth.RefreshableTask.getValue(RefreshableTask.java:192)\n\tat com.amazonaws.auth.STSAssumeRoleWithWebIdentitySessionCredentialsProvider.getCredentials(STSAssumeRoleWithWebIdentitySessionCredentialsProvider.java:130)

```

We doubled the memory of the masters, but this can still happen.

---

<div class="post-metadata">

**Author:** ![Doc\_Kaos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/doc_kaos/32/53671_2.png) [@Doc\_Kaos](https://discuss.elastic.co/u/Doc_Kaos)\
**Post date:** [August 21, 2024, 1:09pm UTC](https://discuss.elastic.co/t/removing-second-s3-repository-causes-connection-pool-shutdown-in-8-13-2/364797/5 "2024-08-21T13:09:36Z")

</div>

It appears this is because this cluster had a second S3 Repository added, then removed. We haven't found a way to recover from this yet, but see that 8.15.0 has a release note that potentially this was fixed.

---

<div class="post-metadata">

**Author:** ![Doc\_Kaos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/doc_kaos/32/53671_2.png) [@Doc\_Kaos](https://discuss.elastic.co/u/Doc_Kaos)\
**Post date:** [August 23, 2024, 1:17pm UTC](https://discuss.elastic.co/t/removing-second-s3-repository-causes-connection-pool-shutdown-in-8-13-2/364797/6 "2024-08-23T13:17:58Z")

</div>

Removing the second S3 Repository, then restarting all masters ... then restarting all data nodes appears to have solved the problem. 🤞

---

<div class="post-metadata">

**Author:** ![Doc\_Kaos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/doc_kaos/32/53671_2.png) [@Doc\_Kaos](https://discuss.elastic.co/u/Doc_Kaos)\
**Post date:** [August 26, 2024, 2:03pm UTC](https://discuss.elastic.co/t/removing-second-s3-repository-causes-connection-pool-shutdown-in-8-13-2/364797/7 "2024-08-26T14:03:06Z")

</div>

Seeing other reporting same issue: [Removing one of the s3 snapshot repository causing connection pool shutdown](https://discuss.elastic.co/t/removing-one-of-the-s3-snapshot-repository-causing-connection-pool-shutdown/329857)  
[Repository\_verification\_exception](https://discuss.elastic.co/t/repository-verification-exception/363174)

---

<div class="post-metadata">

**Author:** ![Doc\_Kaos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/doc_kaos/32/53671_2.png) [@Doc\_Kaos](https://discuss.elastic.co/u/Doc_Kaos)\
**Post date:** [August 26, 2024, 2:22pm UTC](https://discuss.elastic.co/t/removing-second-s3-repository-causes-connection-pool-shutdown-in-8-13-2/364797/8 "2024-08-26T14:22:42Z")

</div>

It appears that the "snapshot retention" job is the one that causes the failure, not the snapshot itself

---

<div class="post-metadata">

**Author:** ![Doc\_Kaos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/doc_kaos/32/53671_2.png) [@Doc\_Kaos](https://discuss.elastic.co/u/Doc_Kaos)\
**Post date:** [August 29, 2024, 2:08pm UTC](https://discuss.elastic.co/t/removing-second-s3-repository-causes-connection-pool-shutdown-in-8-13-2/364797/9 "2024-08-29T14:08:45Z")

</div>

Resolved (for a few days now at least) by removing all but one S3 repo. Restarting all nodes in the cluster, then immediately removing old snapshots up to our current retention.
