# "AlreadyClosedException" with "Too Many Open Files" - breaking indexing

**URL:** <https://discuss.elastic.co/t/alreadyclosedexception-with-too-many-open-files-breaking-indexing/388808>\
**Category:** Elasticsearch\
**Created:** [July 27, 2026, 8:35am UTC](https://discuss.elastic.co/t/alreadyclosedexception-with-too-many-open-files-breaking-indexing/388808 "2026-07-27T08:35:45Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![astrodi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/astrodi/32/117732_2.png) [@astrodi](https://discuss.elastic.co/u/astrodi)\
**Post date:** [July 27, 2026, 8:35am UTC](https://discuss.elastic.co/t/alreadyclosedexception-with-too-many-open-files-breaking-indexing/388808/1 "2026-07-27T08:35:46Z")

</div>

Hello there 👋

I'm experiencing allocation issues on cluster with 22 data nodes (K8S cluster, each node holds 29-41 shards).

In total, there is 122M documents, 13TB of storage consumed and 456 pri. shards + 328 replica shards.

Version of Elasticsearch is 8.19.4

Symptoms:

- Elasticsearch cluster falls into yellow state (UNASSIGNED replica shards) during indexing phase
- number of UNASSIGNED replica shards increases over time
- issue is repetitive (when restart problematic nodes -\> cluster green, but next indexing causing same trouble)
- issue visible after added more indices into cluster (more shards)

```auto
- [indices:data/write/bulk[s][r]] Caused by: org.apache.lucene.store.AlreadyClosedException: this ReferenceManager is closed

- Too many open files in system

- failed to create shard, failure org.elasticsearch.ElasticsearchTimeoutException: timed out while waiting to acquire shard lock for ...

```

Monitored:

- Open file count on ES side: `"open_file_descriptors": 2461, "max_file_descriptors": 1048576` (checked on all nodes, looks OK)

- Swap: `swap.used = 0`

- Heap usage = 60-70%, RAM usage = 75-99%

- monitor `hot_threads` (stuck write operations, shard waiting indefinitely for closing, NOT CPU problem)

Steps I tried to resolve the issue:

- `POST /_cluster/reroute?retry_failed&metric=none` is NOT solving the issue (ends up on 5x retry and back unassigned)
- set `node_concurrent_recoveries = 2` (was 10)
- increase `refresh_interval = 30s` (was 10s)
- disable `slow_logs`
- removed unnecessary indices and decreased shard-count
- looks, that issue doesn't happen (so frequently) on non-replica indices

None of steps I tried helped to prevent the issue.

The only short-term solution was restart ES data nodes - but it is not resolving root-cause of the problem.

I got advice that this might be resources issue on OS level, but don't know where to focus & identify which resource to increase.

Thank you in advance for advises, what to check 👍

Dominik

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [July 27, 2026, 10:01am UTC](https://discuss.elastic.co/t/alreadyclosedexception-with-too-many-open-files-breaking-indexing/388808/2 "2026-07-27T10:01:29Z")

</div>

> [@astrodi](#):
>
> `Too many open files in system`

This error (with the `in system` suffix) is how glibc renders the `ENFILE` error, which means you've hit some kind of system-wide limit on open files rather than anything specific to a single process. Compare:

> <https://github.com/bminor/glibc/blob/04e750e75b73957cf1c791535a3f4319534a52fc/manual/errno.texi#L300-L306>

with `EMFILE` that doesn't have the "in system" suffix:

> <https://github.com/bminor/glibc/blob/04e750e75b73957cf1c791535a3f4319534a52fc/manual/errno.texi#L288-L298>

This aligns with `"open_file_descriptors": 2461` which is a totally reasonable number of files for the Elasticsearch process to hold open.

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [July 27, 2026, 10:46am UTC](https://discuss.elastic.co/t/alreadyclosedexception-with-too-many-open-files-breaking-indexing/388808/3 "2026-07-27T10:46:29Z")

</div>

> [@astrodi](#):
>
> K8S cluster

Complicates things a little bit ...

> [@astrodi](#):
>
> `- Too many open files in system`

As @DavidTurner said, that’s a system error. Do you have full access to the Kubernetes environment, or is that looked after by another team?

For example, can _you_ inspect `/proc/sys/fs/file-nr` and `/proc/sys/fs/file-max` on the Kubernetes nodes? Can you run `kubectl exec` commands? Do you know which container runtime your Kubernetes environment is using?

Troubleshooting this would, IMO, be a Unix/Linux/Kubernetes administration job rather than an “Elasticsearch team” job — noting that, in a ton of environments, but far from all, those are exactly the same people.

---

<div class="post-metadata">

**Author:** ![astrodi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/astrodi/32/117732_2.png) [@astrodi](https://discuss.elastic.co/u/astrodi)\
**Post date:** [July 27, 2026, 10:55am UTC](https://discuss.elastic.co/t/alreadyclosedexception-with-too-many-open-files-breaking-indexing/388808/4 "2026-07-27T10:55:54Z")

</div>

@RainTown - yes thanks for that, I'm already there - `/proc/sys/fs/file-nr` seems to be problem. Now I'm looking for the way, how to increase it via our Terra-form scripts & with DevOps specialist

![image](https://us1.discourse-cdn.com/elastic/original/3X/5/c/5cf207c24da2be4f9bdec14ee3699b41ba19c882.png)

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [July 27, 2026, 11:45am UTC](https://discuss.elastic.co/t/alreadyclosedexception-with-too-many-open-files-breaking-indexing/388808/5 "2026-07-27T11:45:20Z")

</div>

> [@astrodi](#):
>
> DevOps specialist

Glad that you seem on right track, a "DevOps Specialist” _sounds_ like the right guy/gal. But minor rant, and nothing to do with you @astrodi, but ....  
I remember when IT had about ten job titles. The more we fragment IT into hyper-specialised titles/roles, the more people we need to fix a problem that one competent sysadmin could understand and likely solve in about 10 minutes. Sadly.

---

<div class="post-metadata">

**Author:** ![astrodi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/astrodi/32/117732_2.png) [@astrodi](https://discuss.elastic.co/u/astrodi)\
**Post date:** [July 27, 2026, 4:52pm UTC](https://discuss.elastic.co/t/alreadyclosedexception-with-too-many-open-files-breaking-indexing/388808/6 "2026-07-27T16:52:03Z")

</div>

I was able to change it myself (`helm/values-dev.yaml` and `helm/values-prod.yaml`)

This was the code required:

```auto
sysctl:
  - name: sysctl
    image: *shell_image
    imagePullPolicy: Always
    command:
      - /bin/bash
      - -ec
      - |
        sysctl -w vm.max_map_count=262144 && sysctl -w fs.file-max=2097152
        sysctl vm.max_map_count=262144 && sysctl fs.file-max=2097152

    securityContext:
      privileged: true
      runAsUser: 0
      runAsGroup: 0

```

So far tested in DEV, will continue with PROD cluster

`kubectl exec -n ire-elasticsearch ire-elasticsearch-data-0 -- cat /proc/sys/fs/file-nr`

![image](https://us1.discourse-cdn.com/elastic/original/3X/d/2/d22b85a667e83ea7baa5adfe582dfcfb28c2d048.png)

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [July 27, 2026, 5:17pm UTC](https://discuss.elastic.co/t/alreadyclosedexception-with-too-many-open-files-breaking-indexing/388808/7 "2026-07-27T17:17:09Z")

</div>

Brilliant, well done. There's likely other ways, but I hope that sorts it out. Sorry for minor rant above.

> [@astrodi](#):
>
> ```auto
> sysctl -w vm.max_map_count=262144 && sysctl -w fs.file-max=2097152
> sysctl vm.max_map_count=262144 && sysctl fs.file-max=2097152
> 
> ```

btw, not sure what the second line is adding, looks redundant to me.
