# Logstash file input periodically skips over files

**URL:** <https://discuss.elastic.co/t/logstash-file-input-periodically-skips-over-files/370583>\
**Category:** Logstash\
**Created:** [November 14, 2024, 9:56pm UTC](https://discuss.elastic.co/t/logstash-file-input-periodically-skips-over-files/370583 "2024-11-14T21:56:00Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![ddiguru](https://avatars.discourse-cdn.com/v4/letter/d/7c8e57/32.png) [@ddiguru](https://discuss.elastic.co/u/ddiguru)\
**Post date:** [November 14, 2024, 9:56pm UTC](https://discuss.elastic.co/t/logstash-file-input-periodically-skips-over-files/370583/1 "2024-11-14T21:56:00Z")

</div>

I am attempting to use the `file` input module to "watch" a directory for inbound files. Gzipped log files are shipped to this directory from various remote hosts, and the `file` input picks them up, decompresses or processes them "as is" and sends data to some upstream collector - graylog, elasticsearch, etc..

ubuntu 24.0.1  
logstash 8.15.3

It is mostly working but every now and then it will skip a file or set of files. I am pretty sure my config is valid b/c it works MOST of the time.

```auto
input {
    file {
        path => "/var/lib/logstash/data/*gz"
        start_position => "beginning"
        mode => "read"
        file_completed_log_path => "/var/lib/logstash/consumed.log"
        file_completed_action => "log_and_delete"
    }
}

```

Anyone have ideas how best to troubleshoot this and/or know if i'm running into any known issues?

Thanks

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [November 14, 2024, 10:02pm UTC](https://discuss.elastic.co/t/logstash-file-input-periodically-skips-over-files/370583/2 "2024-11-14T22:02:07Z")

</div>

That could be inode re-use. A low value for [sincedb\_clean\_after](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html#plugins-inputs-file-sincedb_clean_after) might help.

---

<div class="post-metadata">

**Author:** ![ddiguru](https://avatars.discourse-cdn.com/v4/letter/d/7c8e57/32.png) [@ddiguru](https://discuss.elastic.co/u/ddiguru)\
**Post date:** [November 14, 2024, 10:12pm UTC](https://discuss.elastic.co/t/logstash-file-input-periodically-skips-over-files/370583/3 "2024-11-14T22:12:55Z")

</div>

Is that possible given i'm getting files predictably every 10m from various sources?

How "low" do you suggest?... i think the def. is 2w.

---

<div class="post-metadata">

**Author:** ![ddiguru](https://avatars.discourse-cdn.com/v4/letter/d/7c8e57/32.png) [@ddiguru](https://discuss.elastic.co/u/ddiguru)\
**Post date:** [November 14, 2024, 10:15pm UTC](https://discuss.elastic.co/t/logstash-file-input-periodically-skips-over-files/370583/4 "2024-11-14T22:15:33Z")

</div>

maybe do 15 mins?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [November 14, 2024, 10:27pm UTC](https://discuss.elastic.co/t/logstash-file-input-periodically-skips-over-files/370583/5 "2024-11-14T22:27:41Z")

</div>

> [@ddiguru](#):
>
> Is that possible given i'm getting files predictably every 10m from various sources?

Yes. If you have logstash delete files after reading them then the inodes are freed up and on some filesystems that will put them into a cache to be reused. This can make re-use quite common.

I would base the value of sincedb\_clean\_after on the maximum time you ever expect a file to stay in /var/lib/logstash/data/

---

<div class="post-metadata">

**Author:** ![ddiguru](https://avatars.discourse-cdn.com/v4/letter/d/7c8e57/32.png) [@ddiguru](https://discuss.elastic.co/u/ddiguru)\
**Post date:** [November 14, 2024, 10:34pm UTC](https://discuss.elastic.co/t/logstash-file-input-periodically-skips-over-files/370583/6 "2024-11-14T22:34:02Z")

</div>

i wonder if i would be better of NOT `log_and_delete`, and allow them to hang 24h period and run a `find /var/lib/logstash/data -delete -mtime +1` sort of thing?

---

<div class="post-metadata">

**Author:** ![ddiguru](https://avatars.discourse-cdn.com/v4/letter/d/7c8e57/32.png) [@ddiguru](https://discuss.elastic.co/u/ddiguru)\
**Post date:** [November 14, 2024, 10:38pm UTC](https://discuss.elastic.co/t/logstash-file-input-periodically-skips-over-files/370583/7 "2024-11-14T22:38:03Z")

</div>

my original thinking is that:

- servers send logs every 10m at varying times
- logstash consumes them just as soon as it can
- applies groks, filters, transformations, etc...
- sends data to upstream graylog for persistence
- next() log file...

So, i am counting on logstash to consume them as quick as possible, then `log_and_delete` those it was able to process.

That's where i'm at... so, i just set `sincedb_clean_after` to 1h. i'll see what that brings me. Interesting challenge inode reuse is...

---

<div class="post-metadata">

**Author:** ![ddiguru](https://avatars.discourse-cdn.com/v4/letter/d/7c8e57/32.png) [@ddiguru](https://discuss.elastic.co/u/ddiguru)\
**Post date:** [November 15, 2024, 5:02pm UTC](https://discuss.elastic.co/t/logstash-file-input-periodically-skips-over-files/370583/8 "2024-11-15T17:02:54Z")

</div>

So, i have changed to `log` ONLY, i've set the `sincedb_clean_after` to .005 which is super low. I have 28 log files that have been shipped to the processing folder and i have 28 entries in the consumed.log file.

I will monitor this for a while, but that seems to be a better way to process logs w/o the risk of freq. inode re-use. Thanks @Badger for the tip on where to investigate.
