# Intentionally delaying log harvesting at startup

**URL:** https://discuss.elastic.co/t/intentionally-delaying-log-harvesting-at-startup/149855
**Category:** Beats
**Tags:** filebeat
**Created:** [September 25, 2018, 2:43pm UTC](https://discuss.elastic.co/t/intentionally-delaying-log-harvesting-at-startup/149855 "2018-09-25T14:43:06Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![YvorL](https://avatars.discourse-cdn.com/v4/letter/y/9fc348/32.png) [@YvorL](https://discuss.elastic.co/u/YvorL)
#### Post date: [September 25, 2018, 2:43pm UTC](https://discuss.elastic.co/t/intentionally-delaying-log-harvesting-at-startup/149855/1 "2018-09-25T14:43:06Z")

</div>

I have a use case where I need to delay the log harvesting. As I read my best option would be `scan_frequency` but I'm not sure if that applies for the first start and how it keeps track of this information (I suspect that in memory).  
The issue is that when I recreate a container from a snapshot, I can't stop Filebeat soon enough and I get couple thousand documents added to my cluster. In a way that's really good and fast, but in this case, I have to delay the start while I can remove the logs.  
Can I achieve my goal by modifying the default `scan_frequency` value (10s) to a higher one? Is there any other way without adding more services?

---

<div class="post-metadata">

### Author: ![pierhugues](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pierhugues/32/48383_2.png) [@pierhugues](https://discuss.elastic.co/u/pierhugues)
#### Post date: [September 26, 2018, 1:17pm UTC](https://discuss.elastic.co/t/intentionally-delaying-log-harvesting-at-startup/149855/2 "2018-09-26T13:17:51Z")

</div>

@YvorL _scan\_frequency_ is only effective after the first scan so I don't it will help.

Are you using [Filebeat's autodiscover](https://www.elastic.co/guide/en/beats/filebeat/current/configuration-autodiscover.html) for managing configuration when new container appear disapear?

---

<div class="post-metadata">

### Author: ![YvorL](https://avatars.discourse-cdn.com/v4/letter/y/9fc348/32.png) [@YvorL](https://discuss.elastic.co/u/YvorL)
#### Post date: [September 26, 2018, 4:04pm UTC](https://discuss.elastic.co/t/intentionally-delaying-log-harvesting-at-startup/149855/3 "2018-09-26T16:04:57Z")

</div>

No, I'm not using that. I'm running an instance of Filebeat in the container itself. When a new container is created from the snapshot all processes launch including Filebeat. Unfortunately, that's before the logs are removed. You can think about it as restoring a backup from a snapshot. While the registry file contains the information about the processed logs, I believe, that since the environment changed (e.g., hostname) it'll re-read the logs.

---

<div class="post-metadata">

### Author: ![ruflin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ruflin/32/3116_2.png) [@ruflin](https://discuss.elastic.co/u/ruflin)
#### Post date: [September 26, 2018, 9:28pm UTC](https://discuss.elastic.co/t/intentionally-delaying-log-harvesting-at-startup/149855/4 "2018-09-26T21:28:40Z")

</div>

I wonder if the `tail_files` option could help you? [https://www.elastic.co/guide/en/beats/filebeat/current/filebeat-input-log.html#\_literal\_tail\_files\_literal](https://www.elastic.co/guide/en/beats/filebeat/current/filebeat-input-log.html#_literal_tail_files_literal) Be aware that it can also mean that some events are skipped for new logs.

---

<div class="post-metadata">

### Author: ![YvorL](https://avatars.discourse-cdn.com/v4/letter/y/9fc348/32.png) [@YvorL](https://discuss.elastic.co/u/YvorL)
#### Post date: [September 27, 2018, 10:08am UTC](https://discuss.elastic.co/t/intentionally-delaying-log-harvesting-at-startup/149855/5 "2018-09-27T10:08:28Z")

</div>

Unfortunately, it seems it'd cause more issues than it'd solve. This would work if I can add the setting before the snapshot is taken and remove it immediately, but that's far from optimal and would slow down the process unnecessarily. 😞  
I need either having an option which would delay the start of scanning & harvesting on first run or one where I can set that the environment change (e.g., hostname) doesn't mean that the logs aren't already processed.

---

<div class="post-metadata">

### Author: ![ruflin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ruflin/32/3116_2.png) [@ruflin](https://discuss.elastic.co/u/ruflin)
#### Post date: [October 1, 2018, 6:57am UTC](https://discuss.elastic.co/t/intentionally-delaying-log-harvesting-at-startup/149855/6 "2018-10-01T06:57:28Z")

</div>

I can't think of a good workaround here at the moment ☹ One thing we discussed in the past is have a unique id for each log line so in case events are sent twice, they would not be duplicated but we are not there yet.

---

<div class="post-metadata">

### Author: ![YvorL](https://avatars.discourse-cdn.com/v4/letter/y/9fc348/32.png) [@YvorL](https://discuss.elastic.co/u/YvorL)
#### Post date: [October 2, 2018, 8:39am UTC](https://discuss.elastic.co/t/intentionally-delaying-log-harvesting-at-startup/149855/7 "2018-10-02T08:39:52Z")

</div>

I see. Thank you for the info!

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [October 30, 2018, 8:40am UTC](https://discuss.elastic.co/t/intentionally-delaying-log-harvesting-at-startup/149855/8 "2018-10-30T08:40:03Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
