# Elasticsearch 6.2.2 nodes crash after reaching ulimit setting

**URL:** <https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172>\
**Category:** Elasticsearch\
**Created:** [March 15, 2018, 8:00pm UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172 "2018-03-15T20:00:47Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![iamredlus](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/iamredlus/32/15504_2.png) [@iamredlus](https://discuss.elastic.co/u/iamredlus)\
**Post date:** [March 15, 2018, 8:00pm UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/1 "2018-03-15T20:00:47Z")

</div>

Hi,

We're experiencing a critical production issue in elasticsearch 6.2.2 related to open\_file\_descriptors. The cluster is an exact replica (as much as possible) of a 5.2.2 cluster, and documents are indexed into both clusters in parallel.  
While the indexing performance of the new cluster seems to be at least as good as the 5.2.2 cluster, the new nodes open\_file\_descriptors is reaching record-breaking levels (especially when compared to v5.2.2).

All machines have ulimit of 65536, as recommended by the official documentation.  
All nodes on the v5.2.2 cluster have up to 4,500 open\_file\_descripors, while the new v6.2.2 nodes are divided: some have up to 4,500 open\_file\_descriptors, while others consistently open more and more file descriptors until reaching the limit and crashing with java.nio.file.FileSystemException: Too many open files -

> [WARN][o.e.c.a.s.ShardStateAction] [prod-elasticsearch-master-002] [newlogs\_20180315-01][0] received shard failed for shard id [[newlogs\_20180315-01][0]], allocation id [G8NGOPNHRNuqNKYKzfiPcg], primary term [0], message [shard failure, reason [already closed by tragic event on the translog]], failure [FileSystemException[/mnt/nodes/0/indices/nIgarkzwRwe0DmT-nmLhvg/0/translog/translog.ckp: Too many open files]]  
> java.nio.file.FileSystemException: /mnt/nodes/0/indices/nIgarkzwRwe0DmT-nmLhvg/0/translog/translog.ckp: Too many open files

After this exception, some of the nodes throw many exceptions and reduce the number of open file descriptors. Other times they just crash. The issue repeats itself with interleaving nodes.

I'd be happy to provide additional details, whatever is needed.

Thanks!

---

<div class="post-metadata">

**Author:** ![nhat](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nhat/32/22170_2.png) [@nhat](https://discuss.elastic.co/u/nhat)\
**Post date:** [March 16, 2018, 3:15pm UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/2 "2018-03-16T15:15:05Z")

</div>

@iamredlus It will be super helpful for us to diagnose the issue if you can provide the shard level \_stats. You can get it via `GET /_stats?include_segment_file_sizes&level=shards`. Thank you.

---

<div class="post-metadata">

**Author:** ![Ariel\_Assaraf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ariel_assaraf/32/22545_2.png) [@Ariel\_Assaraf](https://discuss.elastic.co/u/Ariel_Assaraf)\
**Post date:** [March 16, 2018, 3:35pm UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/3 "2018-03-16T15:35:46Z")

</div>

Hey @nhat, any news reg the files @iamredlus sent you?

---

<div class="post-metadata">

**Author:** ![nhat](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nhat/32/22170_2.png) [@nhat](https://discuss.elastic.co/u/nhat)\
**Post date:** [March 16, 2018, 3:53pm UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/4 "2018-03-16T15:53:33Z")

</div>

@Ariel_Assaraf I haven't received anything yet.

---

<div class="post-metadata">

**Author:** ![farin99](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/farin99/32/27873_2.png) [@farin99](https://discuss.elastic.co/u/farin99)\
**Post date:** [March 16, 2018, 4:12pm UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/5 "2018-03-16T16:12:29Z")

</div>

Just forward you the email with the attachments.  
Please confirm you received it 🙂  
Thanks!

---

<div class="post-metadata">

**Author:** ![nhat](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nhat/32/22170_2.png) [@nhat](https://discuss.elastic.co/u/nhat)\
**Post date:** [March 16, 2018, 4:30pm UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/6 "2018-03-16T16:30:31Z")

</div>

Not yet. My email is first name dot last name at [elastic.co](http://elastic.co)

---

<div class="post-metadata">

**Author:** ![nhat](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nhat/32/22170_2.png) [@nhat](https://discuss.elastic.co/u/nhat)\
**Post date:** [March 17, 2018, 10:32pm UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/7 "2018-03-17T22:32:16Z")

</div>

For future readers:

The root cause is that one replica in the user's cluster got in an infinite flushing loop. We helped the user to resolve the issue by rebuilding replica.

---

<div class="post-metadata">

**Author:** ![LLin](https://avatars.discourse-cdn.com/v4/letter/l/e68b1a/32.png) [@LLin](https://discuss.elastic.co/u/LLin)\
**Post date:** [March 20, 2018, 8:36am UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/8 "2018-03-20T08:36:37Z")

</div>

Hello,

I'm facing a similar issue with Elasticsearch after upgrading to 6.2.2 from 5.6.4.  
The number of open files goes to a number that is not reasonable and the cluster node crashes.

The culprit seems to be that a very large number of .tlog files is created for some indices:

for example:

> java 53621 elasticsearch \*767r REG 253,10 43 1086962675 /data/elasticsearch/nodes/0/indices/91x35hspTdSfPy84E4cUwQ/6/translog/translog-100105.tlog

This index has around 120k .tlog files and it's a primary shard.  
Currently, the only way I've found to get rid of all the files is to use Cluster Reroute to move the shard to a different server

---

<div class="post-metadata">

**Author:** ![nhat](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nhat/32/22170_2.png) [@nhat](https://discuss.elastic.co/u/nhat)\
**Post date:** [March 20, 2018, 12:06pm UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/9 "2018-03-20T12:06:09Z")

</div>

@LLin  
Would you please share the shard-level stats of that index (/{index}/\_stats?level=shards). You can email me at firstname dot lastname at [elastic.co](http://elastic.co). Thank you!

---

<div class="post-metadata">

**Author:** ![iamredlus](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/iamredlus/32/15504_2.png) [@iamredlus](https://discuss.elastic.co/u/iamredlus)\
**Post date:** [March 20, 2018, 2:47pm UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/10 "2018-03-20T14:47:21Z")

</div>

This seems to be a deeper issue than just rebuilding the replicas -

> <https://github.com/elastic/elasticsearch/issues/29097>

A pull request with a fix is being discussed here:

> <https://github.com/elastic/elasticsearch/pull/29125>

I hope this will be merged soon and we can test this in our production as well.

---

<div class="post-metadata">

**Author:** ![imehl](https://avatars.discourse-cdn.com/v4/letter/i/8797f3/32.png) [@imehl](https://discuss.elastic.co/u/imehl)\
**Post date:** [March 22, 2018, 7:17am UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/11 "2018-03-22T07:17:34Z")

</div>

Hi,

I have the same issues after upgrading from 5.6.4 to 6.2.2.

@nhat How can I stabilize my cluster based on the shards \_stats until an offical fix is ready?

Edit: If I look at /proc/ES-PID/fd the most files (of over 100.000) are ...indices/AB564m6dTgOWBf7gEqvBiw/translog... -\> Now I know the problematic index.. What should I do with this index?

Edit 2: I removed the replica of this index and the file descriptor count drops from 130.000 to 8.000 🙂

Thanks

---

<div class="post-metadata">

**Author:** ![iamredlus](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/iamredlus/32/15504_2.png) [@iamredlus](https://discuss.elastic.co/u/iamredlus)\
**Post date:** [April 9, 2018, 9:35am UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/13 "2018-04-09T09:35:24Z")

</div>

Commit #29125 fixes this issue. Still unavailable on the latest elasticsearch version, but can be built from the branch:

> <https://github.com/elastic/elasticsearch/issues/29097#issuecomment-379690473>

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 7, 2018, 9:35am UTC](https://discuss.elastic.co/t/elasticsearch-6-2-2-nodes-crash-after-reaching-ulimit-setting/124172/14 "2018-05-07T09:35:35Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
