# Elasticsearch 7.17.28 Critical Bugs

**URL:** <https://discuss.elastic.co/t/elasticsearch-7-17-28-critical-bugs/387493>\
**Category:** Elasticsearch\
**Tags:** snapshot-and-restore\
**Created:** [July 3, 2026, 7:40am UTC](https://discuss.elastic.co/t/elasticsearch-7-17-28-critical-bugs/387493 "2026-07-03T07:40:20Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Abdullah\_Shah](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/abdullah_shah/32/147790_2.png) [@Abdullah\_Shah](https://discuss.elastic.co/u/Abdullah_Shah)\
**Post date:** [July 3, 2026, 7:40am UTC](https://discuss.elastic.co/t/elasticsearch-7-17-28-critical-bugs/387493/1 "2026-07-03T07:40:20Z")

</div>

Anyone still using Elasticsearch 7.17.28  
I'm facing some issues and would like to see if it's a elasticsearch issue or something else

Issues like

1. Snapshot repo permission issue

Desc:  
Elasticsearch loses repo access and gains it automatically after some time causing missed snapshots

1. Client side connection dropping needing restart which further causes downtimes

if anyone faced similar or any elasticsearch side issue in this version (7.17.28)

Do share please

---

<div class="post-metadata">

**Author:** ![Rafa\_Silva](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rafa_silva/32/147814_2.png) [@Rafa\_Silva](https://discuss.elastic.co/u/Rafa_Silva)\
**Post date:** [July 7, 2026, 3:39am UTC](https://discuss.elastic.co/t/elasticsearch-7-17-28-critical-bugs/387493/2 "2026-07-07T03:39:05Z")

</div>

Hi Abdullah\_Shah, welcome to the Elastic community!

Based on the information shared so far, I would be careful before classifying this as a confirmed Elasticsearch 7.17.28 bug. Both symptoms can happen because of Elasticsearch, but they can also be caused by repository backend, permissions, storage, network, client configuration, or cluster instability.

For the snapshot repository issue, I would first validate the repository itself and the path from every relevant Elasticsearch node to that repository. Elasticsearch verifies that a repository is available and functional on the master and data nodes, so an intermittent failure usually needs to be checked from the repository/backend side as well, not only from the Elasticsearch version side.

Official reference:

> **[Manage snapshot repositories in self-managed deployments | Elastic Docs](https://www.elastic.co/docs/deploy-manage/tools/snapshot-and-restore/self-managed)**
>
> This guide shows you how to register a snapshot repository on a self-managed deployment. A snapshot repository is an off-cluster storage location for...

A few important things to confirm:

1. What repository type are you using? fs, s3, azure, gcs, hdfs, etc.
2. Is this self-managed, ECK, ECE, or Elastic Cloud?
3. Are multiple clusters using the same repository? If yes, only one cluster should have write access.
4. Do all master/data nodes have the same access to the repository?
5. Are there any IAM, service account, temporary credential, mount, NFS, firewall, proxy, or object storage endpoint changes around the failure time?
6. What exact error appears in the Elasticsearch logs when the snapshot fails?

I would run these checks when the repository is healthy and again when it is failing:

GET \_snapshot  
GET \_snapshot/\<repo\_name\>  
POST \_snapshot/\<repo\_name\>/\_verify  
GET \_slm/status  
GET \_slm/policy/\*?human=true  
GET \_slm/stats?human=true

For repeated SLM failures, the Elastic troubleshooting guide recommends checking the affected SLM policy and then checking the elected master node logs during the snapshot execution window.

Official reference:

> **[Fix repeated snapshot policy failures | Elastic Docs](https://www.elastic.co/docs/troubleshoot/elasticsearch/repeated-snapshot-failures)**
>
> Repeated snapshot failures are usually an indicator of a problem with your deployment. Continuous failures of automated snapshots can leave a deployment...

For the client-side connection issue, I would separate two scenarios:

1. Application/client HTTP connections to Elasticsearch are dropping.
2. Elasticsearch nodes are leaving/rejoining the cluster internally.

If this is the application/client side, please share the client language and version, the full exception/stack trace, whether there is a load balancer/proxy/firewall between the app and Elasticsearch, and the client timeout/retry configuration.

If this is Elasticsearch node-to-node instability, the elected master logs are the best starting point. Look for messages such as NodeLeftExecutor and NodeJoinExecutor, and check whether the reason is disconnected, lagging, followers check retry count exceeded, or joining after restart.

Official reference:

> **[Troubleshoot an unstable cluster | Elastic Docs](https://www.elastic.co/docs/troubleshoot/elasticsearch/troubleshooting-unstable-cluster)**
>
> Normally, a node will only leave a cluster if deliberately shut down. If a node leaves the cluster unexpectedly, it’s important to address the cause...

I would also check cluster pressure around the same time, especially JVM memory pressure, GC overhead, CPU, disk latency, rejected thread pools, and node restarts. A client may experience connection drops or timeouts when the cluster or one of the target nodes is overloaded.

Useful checks:

GET \_cluster/health?pretty  
GET \_cat/nodes?v&h=name,ip,roles,master,heap.percent,ram.percent,cpu,load\_1m  
GET \_nodes/stats?filter\_path=nodes._.name,nodes._.jvm.mem.pools.old,nodes._.thread\_pool._.rejected

One additional point: 7.17.x is already past its maintenance and support window according to Elastic’s version policy, so even if the immediate root cause is repository/network/client related, I would still include an upgrade plan in the remediation path.

Official reference:

> **[Elastic Product End of Life Dates](https://www.elastic.co/support/eol)**
>
> End of life schedule for Elastic product releases, including Elasticsearch, Kibana, Logstash, Beats, and more....

So, based on the current description, my first hypothesis would be:

- Snapshot issue: intermittent repository/backend/permission/network/access problem.
- Client connection issue: client/network/load balancer/proxy configuration or cluster pressure/instability.

But to confirm that, the exact repository type, deployment type, Elasticsearch logs, SLM policy output, client error stack trace, and cluster health around the failure time would be needed.

---

<div class="post-metadata">

**Author:** ![Abdullah\_Shah](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/abdullah_shah/32/147790_2.png) [@Abdullah\_Shah](https://discuss.elastic.co/u/Abdullah_Shah)\
**Post date:** [July 8, 2026, 11:34am UTC](https://discuss.elastic.co/t/elasticsearch-7-17-28-critical-bugs/387493/3 "2026-07-08T11:34:37Z")

</div>

Thank you Rafa\_SIlva for your detailed help.

and Yes repo is nfs and write-only to one cluster

Also, it is self managed

i further investigated the issue and it was found to be a node restart, an OS issue which was later solved by the team

But recently i faced another similar issue where snapshot failed as partial because one node shard failed due to permission issue of repo

In both scenarios one common factor was change of master  
that happened few min-hours before these problems

My question is

Is it necessary to refresh permissions of elasticsearch on backup repo when nodes rejoin the cluster  
because this file permission error solves on its own when we just open the vm to check the permissions

never had to regrant the permissions so if they never went away why elasticsearch refuses and throws the error

---

<div class="post-metadata">

**Author:** ![Rafa\_Silva](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rafa_silva/32/147814_2.png) [@Rafa\_Silva](https://discuss.elastic.co/u/Rafa_Silva)\
**Post date:** [July 9, 2026, 12:09am UTC](https://discuss.elastic.co/t/elasticsearch-7-17-28-critical-bugs/387493/4 "2026-07-09T00:09:22Z")

</div>

Hi Abdullah\_Shah,

No, you should not need to refresh Elasticsearch permissions when a node rejoins the cluster.

For an `fs` repository over NFS, Elasticsearch depends on the filesystem access available to the OS user running Elasticsearch. If the issue disappears when someone opens/checks the VM, I would first suspect the OS/NFS layer: mount state, stale NFS handle, UID/GID mapping, SELinux, or Elasticsearch starting before the NFS mount is ready.

The repository path must be mounted in the same location on all master/data nodes and allowed in `path.repo`:

> **[Shared file system repository | Elastic Docs](https://www.elastic.co/docs/deploy-manage/tools/snapshot-and-restore/shared-file-system-repository)**
>
> Use a shared file system repository to store snapshots on a shared file system. To register a shared file system repository, first mount the file system...

The master change is a useful clue. After a master change, a different master-eligible node may coordinate repository operations. If that node has a different NFS or permission state, snapshot verification can fail.

I would compare the affected node and the new master with simple OS checks, then run:

POST \_snapshot/\<repo\_name\>/\_verify

If `_verify` fails only sometimes, I would treat this as intermittent repository accessibility from one or more nodes, not as a permission refresh needed inside Elasticsearch.

---

<div class="post-metadata">

**Author:** ![Abdullah\_Shah](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/abdullah_shah/32/147790_2.png) [@Abdullah\_Shah](https://discuss.elastic.co/u/Abdullah_Shah)\
**Post date:** [July 9, 2026, 8:40am UTC](https://discuss.elastic.co/t/elasticsearch-7-17-28-critical-bugs/387493/5 "2026-07-09T08:40:31Z")

</div>

Yeah i do also think it's a stale NFS issue.

While i investigate it's root cause further  
I think the best solution right is to either make nfs pre boot check for Elasticsearch or just refresh permissions when doing maintenance on ecs nodes.

Well anyways,

Thank you for your time and valuable knowledge.

Stay blessed.
