# Logstash failing with "ERROR::EMFILE- Too many open files"

**URL:** <https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715>\
**Category:** Logstash\
**Created:** [July 16, 2019, 11:09am UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715 "2019-07-16T11:09:24Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![Laddu\_ps](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/laddu_ps/32/50234_2.png) [@Laddu\_ps](https://discuss.elastic.co/u/Laddu_ps)\
**Post date:** [July 16, 2019, 11:09am UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/1 "2019-07-16T11:09:25Z")

</div>

Hi. I'm having issues running logstash withi large sets of data. I'm trying to push data to Elasticsearch via Logstash Where both the applications are hosted on CENTOS servers. I'm trying to send 1000 files per minute (each of size 2KB) for 1 hour. But I'm facing the following issue after sending 3992 (exactly) files(ie., in four minutes).

`[2019-07-16T06:19:36,398][DEBUG][logstash.instrument.periodicpoller.cgroup] Error, cannot retrieve cgroups information {:exception=>"Errno::EMFILE", :message=>"Too many open files - Too many open files"`

I've tried increasing the system limits too. I've increased the pipeline workers to 6 as well as the batch size to 1000.I Don't have any other idea of what else to change or whether I've made something wrong in the configuration. Thanks in advance.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 16, 2019, 12:10pm UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/2 "2019-07-16T12:10:02Z")

</div>

What does your input configuration look like?

---

<div class="post-metadata">

**Author:** ![Laddu\_ps](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/laddu_ps/32/50234_2.png) [@Laddu\_ps](https://discuss.elastic.co/u/Laddu_ps)\
**Post date:** [July 16, 2019, 12:22pm UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/3 "2019-07-16T12:22:34Z")

</div>

```
input {
  file {
    path => ["/home/stats/archive/*.csv"]
    start_position => "beginning"
    max_open_files => 40000
# sincedb_path => "/tmp/.productsSince15.db"
  }
}
filter {
  csv {
    separator => ","
    skip_header => true
    columns => ["mitigation_id","time","total_in_packets","total_out_packets","total_dropped_packets","total_in_bytes","total_out_bytes","total_dropped_bytes","invalid_pkts_in_packets","invalid_pkts_out_packets","invalid_pkts_dropped_packets","invalid_pkts_in_bytes","invalid_pkts_out_bytes","invalid_pkts_dropped_bytes","acl_in_packets","acl_out_packets","acl_dropped_packets","acl_in_bytes","acl_out_bytes","acl_dropped_bytes","zombie_detection_in_packets","zombie_detection_out_packets","zombie_detection_dropped_packets","zombie_detection_in_bytes","zombie_detection_out_bytes","zombie_detection_dropped_bytes","geo_ip_in_packets","geo_ip_out_packets","geo_ip_dropped_packets","geo_ip_in_bytes","geo_ip_out_bytes","geo_ip_dropped_bytes","tcpsyn_in_packets","tcpsyn_out_packets","tcpsyn_dropped_packets","tcpsyn_in_bytes","tcpsyn_out_bytes","tcpsyn_dropped_bytes","payload_in_packets","payload_out_packets","payload_dropped_packets","payload_in_bytes","payload_out_bytes","payload_dropped_bytes", "http_mlfd_in_packets","http_mlfd_out_packets","http_mlfd_dropped_packets","http_mlfd_in_bytes","http_mlfd_out_bytes","http_mlfd_dropped_bytes","dns_mlfd_in_packets","dns_mlfd_out_packets","dns_mlfd_dropped_packets","dns_mlfd_in_bytes","dns_mlfd_out_bytes","dns_mlfd_dropped_bytes","traffic_management_in_packets","traffic_management_out_packets","traffic_management_dropped_packets","traffic_management_in_bytes","traffic_management_out_bytes","traffic_management_dropped_bytes"]
  }
  date {
    match => ["time", "YYYY-MM-dd HH:mm:ss"]
    target => "timestamp"
   } 
 
mutate {
    remove_field => ["message"]
}

mutate {
  convert => { 
               "mitigation_id" => "integer_eu"
               "total_in_packets" => "integer_eu"
               "total_out_packets" => "integer_eu"
               "total_dropped_packets" => "integer_eu"
               "total_in_bytes" => "integer_eu"
               "total_out_bytes" => "integer_eu"
               "total_dropped_bytes" => "integer_eu"
               "invalid_pkts_in_packets" => "integer_eu"
               "invalid_pkts_out_packets" => "integer_eu"
               "invalid_pkts_dropped_packets" => "integer_eu"
               "invalid_pkts_in_bytes" => "integer_eu"
               "invalid_pkts_out_bytes" => "integer_eu"
               "invalid_pkts_dropped_bytes" => "integer_eu"
               "acl_in_packets" => "integer_eu"
               "acl_out_packets" => "integer_eu"
               "acl_dropped_packets" => "integer_eu"
               "acl_in_bytes" => "integer_eu"
               "acl_out_bytes" => "integer_eu"
               "acl_dropped_bytes" => "integer_eu"
               "zombie_detection_in_packets" => "integer_eu"
               "zombie_detection_out_packets" => "integer_eu"
               "zombie_detection_dropped_packets" => "integer_eu"
               "zombie_detection_in_bytes" => "integer_eu"
               "zombie_detection_out_bytes" => "integer_eu"
               "zombie_detection_dropped_bytes" => "integer_eu"
               "geo_ip_in_packets" => "integer_eu"
               "geo_ip_out_packets" => "integer_eu"
               "geo_ip_dropped_packets" => "integer_eu"
               "geo_ip_in_bytes" => "integer_eu"
               "geo_ip_out_bytes" => "integer_eu"
               "geo_ip_dropped_bytes" => "integer_eu"
               "tcpsyn_in_packets" => "integer_eu"
               "tcpsyn_out_packets" => "integer_eu"
               "tcpsyn_dropped_packets" => "integer_eu"
               "tcpsyn_in_bytes" => "integer_eu"
               "tcpsyn_out_bytes" => "integer_eu"
               "tcpsyn_dropped_bytes" => "integer_eu"
               "payload_in_packets" => "integer_eu"
               "payload_out_packets" => "integer_eu"
               "payload_dropped_packets" => "integer_eu"
               "payload_in_bytes" => "integer_eu"
               "payload_out_bytes" => "integer_eu"
               "payload_dropped_bytes" => "integer_eu"
               "http_mlfd_in_packets" => "integer_eu"
               "http_mlfd_out_packets" => "integer_eu"
               "http_mlfd_dropped_packets" => "integer_eu"
               "http_mlfd_in_bytes" => "integer_eu"
               "http_mlfd_out_bytes" => "integer_eu"
               "http_mlfd_dropped_bytes" => "integer_eu"
               "dns_mlfd_in_packets" => "integer_eu"
               "dns_mlfd_out_packets" => "integer_eu"
               "dns_mlfd_dropped_packets" => "integer_eu"
               "dns_mlfd_in_bytes" => "integer_eu"
               "dns_mlfd_out_bytes" => "integer_eu"
               "dns_mlfd_dropped_bytes" => "integer_eu"
               "traffic_management_in_packets" => "integer_eu"
               "traffic_management_out_packets" => "integer_eu"
               "traffic_management_dropped_packets" => "integer_eu"
               "traffic_management_in_bytes" => "integer_eu"
               "traffic_management_out_bytes" => "integer_eu"
               "traffic_management_dropped_bytes" => "integer_eu"   
  }
}
}
output {
  elasticsearch {
    hosts => "http://localhost:9200"
    index => "typestats-%{+YYYY-MM-dd_HH_mm}"
  }
}
```

---

<div class="post-metadata">

**Author:** ![Laddu\_ps](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/laddu_ps/32/50234_2.png) [@Laddu\_ps](https://discuss.elastic.co/u/Laddu_ps)\
**Post date:** [July 16, 2019, 12:23pm UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/4 "2019-07-16T12:23:40Z")

</div>

I'm trying to index .csv files in elasticsearch via Logstash.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 16, 2019, 12:31pm UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/5 "2019-07-16T12:31:11Z")

</div>

It sounds like you have not increased the open file limit for the process, so I would look into that. Also, you might want to significantly lower the value of [close\_older](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html#plugins-inputs-file-close_older) on the file input, and also consider switching to read mode instead of tail mode.

---

<div class="post-metadata">

**Author:** ![Laddu\_ps](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/laddu_ps/32/50234_2.png) [@Laddu\_ps](https://discuss.elastic.co/u/Laddu_ps)\
**Post date:** [July 18, 2019, 6:57pm UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/6 "2019-07-18T18:57:13Z")

</div>

Hi, Sorry for the late reply. Thanks for the heads-up! I did the changes like you said

- changing the mode to read mode
- increasing the max open files limit (10000)
- decreasing the close\_older limit (3 minutes)  
This setup was working fine for 4-5 hours. But again the Logstash stopped processing the incoming data after that. While there were 2,88,000 files and more, Logstash was able to process only 1,65,900 files and got stuck at it.  
This is the debug logs

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/9/4/94d510a22e4f12672c014b15d88fab052ef08fb8.png)

Is there anything I'm missing out?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 18, 2019, 7:10pm UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/7 "2019-07-18T19:10:26Z")

</div>

In tail mode I can see why you would want to tail a large number of files, but in read mode logstash is going to open a file, read the whole file, and process the whole file through the pipeline. That's basically going to keep 1 CPU busy. I would not set the open file limit much above the number of CPUs that you have. If you increase it further then if anything I would expect throughput to get worse.

Are you deleting files and then adding more files? That can result in [inode reuse](https://discuss.elastic.co/t/logstash-configuration-returns-different-results/190471/11).

---

<div class="post-metadata">

**Author:** ![Laddu\_ps](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/laddu_ps/32/50234_2.png) [@Laddu\_ps](https://discuss.elastic.co/u/Laddu_ps)\
**Post date:** [July 18, 2019, 7:23pm UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/8 "2019-07-18T19:23:18Z")

</div>

There is still so many inodes unused.  
Is it still necessary to delete/remove the files for inode reuse?  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/b/e/beec71cc1b0f95f2790ec53f5da99c108d41078c.png)

---

<div class="post-metadata">

**Author:** ![Laddu\_ps](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/laddu_ps/32/50234_2.png) [@Laddu\_ps](https://discuss.elastic.co/u/Laddu_ps)\
**Post date:** [July 18, 2019, 7:24pm UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/9 "2019-07-18T19:24:26Z")

</div>

Actually, I'm not deleting any files as of now.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 18, 2019, 7:33pm UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/10 "2019-07-18T19:33:17Z")

</div>

Then the next step would be '--log.level trace' (not debug) which will cause the filewatch code to log voluminously. There will be millions of lines of logs, but you just need to find the lines related to one file that does not get read.

---

<div class="post-metadata">

**Author:** ![Laddu\_ps](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/laddu_ps/32/50234_2.png) [@Laddu\_ps](https://discuss.elastic.co/u/Laddu_ps)\
**Post date:** [July 18, 2019, 7:35pm UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/11 "2019-07-18T19:35:41Z")

</div>

Will try that too and tell you how it works out. Thank you for your valuable time 😃😃

---

<div class="post-metadata">

**Author:** ![Laddu\_ps](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/laddu_ps/32/50234_2.png) [@Laddu\_ps](https://discuss.elastic.co/u/Laddu_ps)\
**Post date:** [July 23, 2019, 7:31am UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/12 "2019-07-23T07:31:58Z")

</div>

I tried debugging the issue by enabling log level as "trace".

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/0/1/0192a1e0ac0567f5598d64783749c8851a13f020.png)  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/d/7/d700c88a41cc1e930017ff2a8bd4ef334a410cde.png)  
Is there any issue with SinceDBcollection?

The logs are shipped with so much delay. Initially 0.15 million log files are shipped without any issue, but whatever comes after that are taking so much time.

Totally 5 million files were created in 3 days but only 0.7 million files got shipped till now. The logstash didn't stop but it processed with much delay.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 23, 2019, 1:16pm UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/13 "2019-07-23T13:16:02Z")

</div>

The fact that you are seeing bytes\_read = bytes\_unread = 0 with a non-zero file size definitely suggests inode re-use to me.

Pick an inode, such as 1812511751, and search for every reference to it. I suspect you will find it being used for two different files.

---

<div class="post-metadata">

**Author:** ![Laddu\_ps](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/laddu_ps/32/50234_2.png) [@Laddu\_ps](https://discuss.elastic.co/u/Laddu_ps)\
**Post date:** [July 24, 2019, 9:52am UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/14 "2019-07-24T09:52:43Z")

</div>

If it's inode re-use, what is the suggested configuration?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 24, 2019, 12:06pm UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/15 "2019-07-24T12:06:28Z")

</div>

I do not have a good answer for that.

---

<div class="post-metadata">

**Author:** ![Laddu\_ps](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/laddu_ps/32/50234_2.png) [@Laddu\_ps](https://discuss.elastic.co/u/Laddu_ps)\
**Post date:** [August 8, 2019, 5:39am UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/16 "2019-08-08T05:39:31Z")

</div>

Thanks anyway!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 5, 2019, 5:39am UTC](https://discuss.elastic.co/t/logstash-failing-with-error-emfile-too-many-open-files/190715/17 "2019-09-05T05:39:32Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
