# Filebeat Not harvesting, file didn't change - do not use modification time

**URL:** <https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834>\
**Category:** Beats\
**Tags:** filebeat\
**Created:** [May 11, 2016, 11:26pm UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834 "2016-05-11T23:26:39Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![wmcdonald](https://avatars.discourse-cdn.com/v4/letter/w/f6c823/32.png) [@wmcdonald](https://discuss.elastic.co/u/wmcdonald)\
**Post date:** [May 11, 2016, 11:26pm UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834/1 "2016-05-11T23:26:39Z")

</div>

My understanding is that filebeat will look at the modification timestamp provided by the OS to determine if the file has been modified and then the harvester will try and read from where it left off, correct?

There is a [bug in the Windows 2008 R2](https://social.technet.microsoft.com/Forums/office/en-US/2b8baca2-9c1b-4d80-80ed-87a3d6b1336f/file-timestamp-not-updating-on-2008-but-does-on-2003?forum=winservergen) that manifests itself in not updating the modification time.

Wouldn't it be a better strategy to compare the offset from the last read to the filesize and see if it has grown, instead of only relying on the OS mod time? This could be an option like:

`modification_detection_strategy => ["filesize", "modtime"]`  
**this would indicate trying both strategies in the order listed.**

Since I am using filebeat to monitor log files that simply grow until they roll over, a filesize greater than the filesize of the last time it harvested should signify more new data to harvest.

I don't know JRuby, so if someone can create a patch and explain how to install it using simple OS commands - that would be very helpful.

---

<div class="post-metadata">

**Author:** ![ruflin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ruflin/32/3116_2.png) [@ruflin](https://discuss.elastic.co/u/ruflin)\
**Post date:** [May 12, 2016, 7:18am UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834/2 "2016-05-12T07:18:53Z")

</div>

Interesting, I wasn't aware of this bug. @steffens you could also be interested in this one.

If the harvester is already reading a file, the ModTime is not used and only the offset. Only to decide if a file should be picked up again, a decision is made on the ModTime. Currently the Prospector does not "open" the file to check for the offset, so this would mean a change in logic.

Filebeat is actually written in Golang and not JRuby.

Feel free to open a feature request for this here: [https://github.com/elastic/beats](https://github.com/elastic/beats)

---

<div class="post-metadata">

**Author:** ![logstash\_oz](https://avatars.discourse-cdn.com/v4/letter/l/2bfe46/32.png) [@logstash\_oz](https://discuss.elastic.co/u/logstash_oz)\
**Post date:** [May 19, 2016, 7:38am UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834/3 "2016-05-19T07:38:44Z")

</div>

I think it is not a bug but design in windows, to reduce disk I/O

[file-date-modified-property-are-not-updating-while-modifying-a-file-without-closing-i](https://blogs.technet.microsoft.com/asiasupp/2010/12/14/file-date-modified-property-are-not-updating-while-modifying-a-file-without-closing-it/)

The similar issue was seen in splunk as well, so I am sure filebeat may have that as well. I know splunk resolved the issue by utilizing the timestamp for last read log, instead of the modtime.  
Any suggestions on how to resolve this? We are losing events in our set up  
What role does ignore\_older, close\_older play with this? In our case it is set to 10m. And I see that even though the log file is getting written to the moddate is same as yesterday when the file rolled over.

Another question:  
Hi ruflin,

1. Can you please let me know the entry for Windows server:  
[https://www.elastic.co/guide/en/beats/filebeat/1.0.1/filebeat-configuration.html](https://www.elastic.co/guide/en/beats/filebeat/1.0.1/filebeat-configuration.html)

The above link shows the entries for unix machines

`paths: - "/var/log/*.log"`

We have logs being written to  
C:\pb\apache-tomcat-7.0.22\logs\abc.log  
getting rolled over every 50 MB, to abc.log.1, abc.log.2 , and so on

Thanks, kindly

---

<div class="post-metadata">

**Author:** ![ruflin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ruflin/32/3116_2.png) [@ruflin](https://discuss.elastic.co/u/ruflin)\
**Post date:** [May 20, 2016, 6:23am UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834/4 "2016-05-20T06:23:20Z")

</div>

The docs for ignore\_older and close\_older can be found here: [https://www.elastic.co/guide/en/beats/filebeat/1.2/configuration-filebeat-options.html#ignore-older](https://www.elastic.co/guide/en/beats/filebeat/1.2/configuration-filebeat-options.html#ignore-older) This is mainly about closing the file handler from the filebeat side which is not related if your app closes the file handler.

Thanks for the above link. Very interesting. If you mention timestamp, which timestamp do you mean? Based on the thread above, currently the only solution would be that your application that writes the log closes it from time to time.

The config entries for Windows are identical. Just use

```auto
paths:
- "C:\pb\apache-tomcat-7.0.22\logs\abc*"

```

---

<div class="post-metadata">

**Author:** ![logstash\_oz](https://avatars.discourse-cdn.com/v4/letter/l/2bfe46/32.png) [@logstash\_oz](https://discuss.elastic.co/u/logstash_oz)\
**Post date:** [May 20, 2016, 4:25pm UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834/5 "2016-05-20T16:25:26Z")

</div>

Timestamp of the last log read from the log file by splunk. So, even if the modtime is not getting updated, but the last log read was a second ago, splunk considers that the file is being actively written to.

```auto
The docs for ignore_older and close_older can be found here: https://www.elastic.co/guide/en/beats/filebeat/1.2/configuration-filebeat-options.html#ignore-older This is mainly about closing the file handler from the filebeat side which is not related if your app closes the file handler.

```

* * *

How do we check if our app does that? We have a webservice running on apache tomcat server.  
We implemented, a script to touch the file every 5 minutes.

touch -m filename

I will run a debug mode to compare the transactions sent in the debug logs and actual files.  
What role does logstash workers play, here?  
We have logstash workers as 3, will that send the logs 3 times? We are also load balancing between 4 logstash nodes, one of them is down, which should not be an issue. I am currently seeing multiple events in Kibana, wondering if that is because of the logsatsh workers in filebeat.yml?

Thanks!

---

<div class="post-metadata">

**Author:** ![logstash\_oz](https://avatars.discourse-cdn.com/v4/letter/l/2bfe46/32.png) [@logstash\_oz](https://discuss.elastic.co/u/logstash_oz)\
**Post date:** [May 20, 2016, 5:02pm UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834/6 "2016-05-20T17:02:30Z")

</div>

I do not know coding but if , you can code to store the timestamp of the last log read in a variable, and check against that everytime, instead of checking the filetimestamp, it would make more sense?

Thanks!

---

<div class="post-metadata">

**Author:** ![ruflin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ruflin/32/3116_2.png) [@ruflin](https://discuss.elastic.co/u/ruflin)\
**Post date:** [May 26, 2016, 7:18am UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834/7 "2016-05-26T07:18:53Z")

</div>

@logstash_oz Sorry for the late reply. For the timestamp: With this PR ([https://github.com/elastic/beats/pull/1703](https://github.com/elastic/beats/pull/1703)) I introduced a field `last_read` which stores the timestamp when the file was last read (not last modified). Currently this is not taken into account when starting a harvester, but it definitively gives the opportunity to do so. I'm thinking of doing some comparison of last\_read and modTime in the future.

About your LS questions: If you have 3 workers the events should still be sent only once. Filebeat has the at least once principle so it can happen in some cases, that a line is sent more then once. Is it just an edge case that you see events multiple times (for example when a node goes down) or happens that very often?

---

<div class="post-metadata">

**Author:** ![logstash\_oz](https://avatars.discourse-cdn.com/v4/letter/l/2bfe46/32.png) [@logstash\_oz](https://discuss.elastic.co/u/logstash_oz)\
**Post date:** [May 26, 2016, 7:15pm UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834/8 "2016-05-26T19:15:27Z")

</div>

Happens often. Thanks!

---

<div class="post-metadata">

**Author:** ![wmcdonald](https://avatars.discourse-cdn.com/v4/letter/w/f6c823/32.png) [@wmcdonald](https://discuss.elastic.co/u/wmcdonald)\
**Post date:** [May 27, 2016, 10:58pm UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834/9 "2016-05-27T22:58:12Z")

</div>

I decided to go with the strategy of using close\_older: 167h

Since the file 'rolls over' every 7 days, I figured that this would keep the harvester going long enough. The only trick is to not have the harvester locking the file when the application rolls the log over and tries to write to a new version of it (on Windows), and to make sure that the roll over is detected correctly. The main file is foo.log and rolls over to something like foo.2016-05-27.log

I still think that, given this problem with the OS, it would be a better design to check the file size and the last read value or have them as policies available.

---

<div class="post-metadata">

**Author:** ![ruflin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ruflin/32/3116_2.png) [@ruflin](https://discuss.elastic.co/u/ruflin)\
**Post date:** [May 30, 2016, 8:34am UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834/10 "2016-05-30T08:34:06Z")

</div>

@logstash_oz @wmcdonald We will definitively take this into account when improving filebeat to have different "configuration" on how to track files.

---

<div class="post-metadata">

**Author:** ![ruflin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ruflin/32/3116_2.png) [@ruflin](https://discuss.elastic.co/u/ruflin)\
**Post date:** [June 13, 2016, 6:57am UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834/11 "2016-06-13T06:57:59Z")

</div>

@logstash_oz @wmcdonald The newest version of filebeat (5.0.0-alpha3) now relies much less on the modification date for harvesting. Could you try out if this version resolves the above issues?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 9:51pm UTC](https://discuss.elastic.co/t/filebeat-not-harvesting-file-didnt-change-do-not-use-modification-time/49834/12 "2017-07-05T21:51:13Z")

</div>


