# How to avoid data loss in logstash?

**URL:** <https://discuss.elastic.co/t/how-to-avoid-data-loss-in-logstash/149582>\
**Category:** Logstash\
**Created:** [September 23, 2018, 9:51am UTC](https://discuss.elastic.co/t/how-to-avoid-data-loss-in-logstash/149582 "2018-09-23T09:51:08Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nikhil\_Jaiswal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nikhil_jaiswal/32/39326_2.png) [@Nikhil\_Jaiswal](https://discuss.elastic.co/u/Nikhil_Jaiswal)\
**Post date:** [September 23, 2018, 9:51am UTC](https://discuss.elastic.co/t/how-to-avoid-data-loss-in-logstash/149582/1 "2018-09-23T09:51:08Z")

</div>

Hi All,

I observed there are data loss in my elasticsearch cluster, I am comparing the log count from two log monitoring tools,one is ELK and another a third party vendor.

I have verified syslog configuration in all network devices(ASA, Palo alto) and all the configuration are same both destination IPs, however the log count of all devices in third part vendor SIEM is more than the log count in elasticsearch. As per my calculation there is 90% of data loss.

Any help would be appreciated.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [September 25, 2018, 6:34pm UTC](https://discuss.elastic.co/t/how-to-avoid-data-loss-in-logstash/149582/2 "2018-09-25T18:34:19Z")

</div>

How are the logs sent to Logstash? TCP or UDP? Have you checked the Logstash and/or Elasticsearch logs to make sure no documents are rejected by ES?

---

<div class="post-metadata">

**Author:** ![Nikhil\_Jaiswal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nikhil_jaiswal/32/39326_2.png) [@Nikhil\_Jaiswal](https://discuss.elastic.co/u/Nikhil_Jaiswal)\
**Post date:** [September 26, 2018, 7:28am UTC](https://discuss.elastic.co/t/how-to-avoid-data-loss-in-logstash/149582/3 "2018-09-26T07:28:19Z")

</div>

Hi @magnusbaeck, currently i am using UDP and i tried with TCP as well but it was not working.

Logstash was sending RST to firewall so i again reverted to UDP.

I went through the logs of logstash i found below log.

[2018-09-26T05:29:53,563][WARN][logstash.outputs.elasticsearch] Could not index event to Elasticsearch. {:status=\>400, :action=\>["index", {:\_id=\>nil, :\_index=\>"logstash-syslog-2018.09.25", :\_type=\>"doc", :\_routing=\>nil}, #\<LogStash::Event:0xd47f1cd\>], :response=\>{"index"=\>{"\_index"=\>"logstash-syslog-2018.09.25", "\_type"=\>"doc", "\_id"=\>"AWYTLMWtieAf\_e\_rdBMv", "status"=\>400, "error"=\>{"type"=\>"mapper\_parsing\_exception", "reason"=\>"failed to parse [timestamp]", "caused\_by"=\>{"type"=\>"illegal\_argument\_exception", "reason"=\>"Invalid format: "Sep 26 2018 05:29:50""}}}}}

I am not sure why it's storing 26th Sep data to Sep 25th index also note i have multiple devices where i am getting logs in "Sep 26 2018 05:29:50"" format.

Is there any way to change the time format for all logs which is coming in logstash?

But again this is for only type of device for other devices like ASA , i do not have logs as per my comparison with other tool.

One more question:

Is there any chances that logstash is dropping the packets from since the total ram consumption of the server is 98%?

Thanks

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [September 26, 2018, 9:07am UTC](https://discuss.elastic.co/t/how-to-avoid-data-loss-in-logstash/149582/4 "2018-09-26T09:07:43Z")

</div>

> Logstash was sending RST to firewall so i again reverted to UDP.

Logstash shouldn't normally close connections like that.

> I am not sure why it's storing 26th Sep data to Sep 25th

Probably because UTC is used for index names.

> Is there any way to change the time format for all logs which is coming in logstash?

You should use a date filter to parse timestamps that ES recognizes.

> Is there any chances that logstash is dropping the packets from since the total ram consumption of the server is 98%?

That's a possibility, but I'd start by addressing the problems listed in the Logstash log. If things are working fine Logstash typically won't log much at all.

---

<div class="post-metadata">

**Author:** ![Nikhil\_Jaiswal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nikhil_jaiswal/32/39326_2.png) [@Nikhil\_Jaiswal](https://discuss.elastic.co/u/Nikhil_Jaiswal)\
**Post date:** [September 26, 2018, 10:25am UTC](https://discuss.elastic.co/t/how-to-avoid-data-loss-in-logstash/149582/5 "2018-09-26T10:25:08Z")

</div>

Hi @magnusbaeck,

Thanks for quick response.

Yes, i agreed on the problem which is listed in logstash.log, but i am still not clear why i am losing the data for other devices which is having proper @timestamp.

---

<div class="post-metadata">

**Author:** ![Nikhil\_Jaiswal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nikhil_jaiswal/32/39326_2.png) [@Nikhil\_Jaiswal](https://discuss.elastic.co/u/Nikhil_Jaiswal)\
**Post date:** [September 27, 2018, 2:35pm UTC](https://discuss.elastic.co/t/how-to-avoid-data-loss-in-logstash/149582/6 "2018-09-27T14:35:24Z")

</div>

Hi @magnusbaeck,

I want to highlight one more point which i forgot to mention in previous comment.

Total RAM of the machine is 16 Gb  
Speed: 667 Mhz  
Interface type: DDR2

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 25, 2018, 2:35pm UTC](https://discuss.elastic.co/t/how-to-avoid-data-loss-in-logstash/149582/7 "2018-10-25T14:35:37Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
