# Using s3 input plugin logstash file\_descriptors aren't closed leading to crash of system

**URL:** <https://discuss.elastic.co/t/using-s3-input-plugin-logstash-file-descriptors-arent-closed-leading-to-crash-of-system/254270>\
**Category:** Logstash\
**Tags:** docker\
**Created:** [November 4, 2020, 12:52pm UTC](https://discuss.elastic.co/t/using-s3-input-plugin-logstash-file-descriptors-arent-closed-leading-to-crash-of-system/254270 "2020-11-04T12:52:25Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![petersedivec](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/petersedivec/32/81934_2.png) [@petersedivec](https://discuss.elastic.co/u/petersedivec)\
**Post date:** [November 4, 2020, 12:52pm UTC](https://discuss.elastic.co/t/using-s3-input-plugin-logstash-file-descriptors-arent-closed-leading-to-crash-of-system/254270/1 "2020-11-04T12:52:26Z")

</div>

Been trying to figure out our problem for the past week and haven't figured out a solution other than restart logstash multiple times when it runs out of resources. Looking for any help or suggestions on what we can try, as of right now i'm doing ingests in chunks and

## Problem in a nutshell

The s3 plugin appears to continue to open file descriptors until it reaches the limit of open file descriptors at which point logstash becomes totally unresponsive and effectively is hung. If the file limits count is set to something small like 4, 8 or 16K files that's what happens, when I set it to a higher count (e.g. 32K or 64K files) it appears that it reaches a certain limit (around 25K file descriptors open and then completely chokes from lack of memory where the jvm heap is at roughly 95%+ of the 8GB allocated. Either way, it seems the root problem is from too many open file descriptors.

## Related threads

There are a handful of topics, issues, describing similar problems, none are smoking guns and most seem unresolved

> [@Logstash file descriptors increasing rapidly](https://discuss.elastic.co/t/logstash-file-descriptors-increasing-rapidly/155947):
>
> Hi, We are using Elastic stack 6.2.4. We have 12 logstash nodes in prod. Each node was working perfectly. Somehow each node started consuming too much file descriptors. 65000 is the max limit of file descriptors set up on each node. Restarting a node resolves our problem temporarily but after some time node starts consuming file descriptors and it takes only 15 mins for a node to reach to its max limit. We are using beats, tcp and http input plugins. We tried to disable it one by one to see wh…

[Logstash crashing with exception IOError: Too many open files · Issue #4815 · elastic/logstash · GitHub](https://github.com/elastic/logstash/issues/4815) (OPEN since 2018)  
[File descriptors are leaked when using HTTP · Issue #1604 · elastic/logstash · GitHub](https://github.com/elastic/logstash/issues/1604) (CLOSED)  
[https://stackoverflow.com/questions/30593849/logstash-close-file-descriptors](https://stackoverflow.com/questions/30593849/logstash-close-file-descriptors)

## Configuration & Pipeline

Been mainly running this in a docker container with Xmx and Xms set to 8GB. Have tried running logstash directly on our linux machine (no docker) and the results are the same.

Here are the config files, pretty simple overall

Pipeline in a nutshell, we have several pipelines reading data from s3. Each s3 input points to a separate bucket that has multiple subfolders (prefixes), each subfolder has ~200 files/day and we're trying at the moment to ingest about 90 days worth of data so about 18K files per pipeline.  
logstash.yml - nothing special configured, basically using defaults  
pipelines.yml - example of pipeline

- pipeline.id: sample  
path.config: "sample-pipeline.conf"  
queue.type: persisted

sample-pipeline.conf

```auto
input { s3 {} }
filter { mutate, kv, grok, drop, fingerprint }
output { elasticsearch {} }

```

## Attempted Fixes

I've played with the open file limits, allocated JVM RAM, and other configuration options, but at the end of the day I can't figure out how to get file descriptors to close which leads to either reaching the limit or heap to max out and GC takes over and dominates all free resources. I've attempted a few things on the input side, like setting option delete option so after the file is read from s3 it is deleted. Still, no luck and so we're stuck at the solution of just running batches of files, shutdown logstash and restart.

---

<div class="post-metadata">

**Author:** ![petersedivec](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/petersedivec/32/81934_2.png) [@petersedivec](https://discuss.elastic.co/u/petersedivec)\
**Post date:** [November 5, 2020, 7:33am UTC](https://discuss.elastic.co/t/using-s3-input-plugin-logstash-file-descriptors-arent-closed-leading-to-crash-of-system/254270/2 "2020-11-05T07:33:13Z")

</div>

Was thinking about posting this as a bug on github but figured I'd attempt getting support here first. our workaround of restarting logstash a few times a day is working for now but not ideal and definitely feels like this should be a bug report. Not sure if it's an S3 plugin issue or a wider issue

---

<div class="post-metadata">

**Author:** ![petersedivec](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/petersedivec/32/81934_2.png) [@petersedivec](https://discuss.elastic.co/u/petersedivec)\
**Post date:** [November 13, 2020, 6:02am UTC](https://discuss.elastic.co/t/using-s3-input-plugin-logstash-file-descriptors-arent-closed-leading-to-crash-of-system/254270/3 "2020-11-13T06:02:32Z")

</div>

ok, so pretty certain this is an issue and looking to open an issue on github related to this thread. It continues to happen to us with our dataset, have bumped the JVM heap for logstash to 16GB and that allows us to push a little more data but eventually things still crap out with too many open file descriptors.

- With 8GB heap in logstash things run well up to about 25K open file descriptors at which point JVM heap is at 99% and GC consumes all the CPU effectively shutting our pipeline down.
- At 16GB heap we get up to 50K open file descriptors before the same issue happens.

Am curious if anyone can explain the 8GB recommendation for JVM Heap from elastic, why not allocate more? Why is the recommendation 4-8GB?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [November 13, 2020, 3:07pm UTC](https://discuss.elastic.co/t/using-s3-input-plugin-logstash-file-descriptors-arent-closed-leading-to-crash-of-system/254270/4 "2020-11-13T15:07:41Z")

</div>

> [@petersedivec](#):
>
> why not allocate more?

Increasing the heap size comes with costs. Not just the memory usage, but an increase in the number of GC roots that remain in memory, which drives up the cost of running GC.

Depending on the pattern of memory usage, you may see an increase in GC costs as heap allocation increases (not as a percentage, but as an absolute value). In some case those increases can be significant.

---

<div class="post-metadata">

**Author:** ![petersedivec](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/petersedivec/32/81934_2.png) [@petersedivec](https://discuss.elastic.co/u/petersedivec)\
**Post date:** [November 16, 2020, 8:03am UTC](https://discuss.elastic.co/t/using-s3-input-plugin-logstash-file-descriptors-arent-closed-leading-to-crash-of-system/254270/5 "2020-11-16T08:03:22Z")

</div>

@Badger - thanks, appreciate the feedback. I've been reading up and found this article very useful [https://www.elastic.co/blog/a-heap-of-trouble](https://www.elastic.co/blog/a-heap-of-trouble). I'm starting to understand the inner workings of elastic and java applications a bit more.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 14, 2020, 8:03am UTC](https://discuss.elastic.co/t/using-s3-input-plugin-logstash-file-descriptors-arent-closed-leading-to-crash-of-system/254270/6 "2020-12-14T08:03:32Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
