# Logstash is processing logs which are already processed before on some later date

**URL:** <https://discuss.elastic.co/t/logstash-is-processing-logs-which-are-already-processed-before-on-some-later-date/124129>\
**Category:** Logstash\
**Created:** [March 15, 2018, 3:24pm UTC](https://discuss.elastic.co/t/logstash-is-processing-logs-which-are-already-processed-before-on-some-later-date/124129 "2018-03-15T15:24:39Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![aniketkk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aniketkk/32/96814_2.png) [@aniketkk](https://discuss.elastic.co/u/aniketkk)\
**Post date:** [March 15, 2018, 3:24pm UTC](https://discuss.elastic.co/t/logstash-is-processing-logs-which-are-already-processed-before-on-some-later-date/124129/1 "2018-03-15T15:24:39Z")

</div>

Hi, I am using Nginx as my app server. I have installed filebeat where Nginx is installed.  
Filebeat pushes these logs to Logstash. Then logstash pushes forward to elasticsearch.

This setup used to work perfectly fine. But on one particular date ( say Feb 13 for example purpose), I saw a sudden spike in the number of logs. Usually, my application receives around 100K logs day. But on that particular date, there were around 10M log records.

When checked, that day had logs of all the days from Dec 10 to Dec 28. So 13 Feb contained all the previous log records of December. But the logstash was running perfectly before Feb 13. Feb 12, 11, 10 had the expected number of records. But not sure what happened on Feb 13 that triggered a sudden reprocessing of older December logs.

Then I checked in Elasticsearch for log records of December. Those logs had been processed on that particular day. But it again got processed on Feb 13.

Since then it has happened at least 3 more times on random dates.

Not sure how to debug it. Is filebeat a culprit or it is problem with logstash.  
Where to debug it?

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![yaauie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yaauie/32/23363_2.png) [@yaauie](https://discuss.elastic.co/u/yaauie)\
**Post date:** [March 15, 2018, 5:45pm UTC](https://discuss.elastic.co/t/logstash-is-processing-logs-which-are-already-processed-before-on-some-later-date/124129/2 "2018-03-15T17:45:11Z")

</div>

I would start with Filebeat; there are a number of configurations that could cause it to re-emit events (after all, sometimes that is a desired behaviour).

- Do you have any evidence that Filebeat could have been restarted around those times?
- Are the files being prospected on a network-attached volume, and if so, do you have any evidence that the volume could have been remounted around those times?

You may also be interested in setting an [`ignore_older`](https://www.elastic.co/guide/en/beats/filebeat/current/configuration-filebeat-options.html#ignore-older) directive, which will ignore all files not modified in the given timespan.

---

<div class="post-metadata">

**Author:** ![aniketkk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aniketkk/32/96814_2.png) [@aniketkk](https://discuss.elastic.co/u/aniketkk)\
**Post date:** [March 16, 2018, 10:11am UTC](https://discuss.elastic.co/t/logstash-is-processing-logs-which-are-already-processed-before-on-some-later-date/124129/3 "2018-03-16T10:11:05Z")

</div>

Thank you @yaauie.

Yes, indeed the filebeat is restarting. We have dockerized the filebeat and whenever the new deployment goes, it fetches the latest image and is restarted. But the logs are stored on instance running the container (using docker volumes). So the old log files always remain. But the registry files are lost when the docker is restarted. So maybe that is the reason for it to reprocess older logs. Correct me if I am wrong.

Now regarding the solution.

So is using **ignore\_older** directive the solution? If yes, what should be the suitable value for it. If we set **ignore\_older** , is it mandatory to set **clean\_inactive**. Please suggest me the best-suited configuration for my case (an nginx server)

I was thinking about one more solution. Why how about exposing **/usr/share/filebeat/data** directory which contains the registery files to the host instance running the container. In this way, even if the filebeat container restarts during the deployment, the registery files would remain in the host instance like the logs. So it would pick from the correct offset and send only relevant logs.

I am thinking second solution perhaps would be better. Please suggest me the best solution.

Regards

---

<div class="post-metadata">

**Author:** ![yaauie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yaauie/32/23363_2.png) [@yaauie](https://discuss.elastic.co/u/yaauie)\
**Post date:** [March 16, 2018, 12:04pm UTC](https://discuss.elastic.co/t/logstash-is-processing-logs-which-are-already-processed-before-on-some-later-date/124129/4 "2018-03-16T12:04:50Z")

</div>

It may be best to ask a new question over in the [Beats](https://discuss.elastic.co/c/beats) forum; I am not familiar enough with Beats to give recommendations about its use under Docker.

That said, as a mitigating factor within `logstash-output-elasticsearch` it is possible to specify the document's id instead of letting it autogenerate; if we were to specify a checksum of relevant fields, we could ensure that any duplicate processing would _overwrite_ previously-written entries instead of creating duplicate entries.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 13, 2018, 12:05pm UTC](https://discuss.elastic.co/t/logstash-is-processing-logs-which-are-already-processed-before-on-some-later-date/124129/5 "2018-04-13T12:05:03Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
