# We lost some metrics from time to time blocked by logstash

**URL:** <https://discuss.elastic.co/t/we-lost-some-metrics-from-time-to-time-blocked-by-logstash/148554>\
**Category:** Logstash\
**Created:** [September 14, 2018, 4:51am UTC](https://discuss.elastic.co/t/we-lost-some-metrics-from-time-to-time-blocked-by-logstash/148554 "2018-09-14T04:51:19Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Robin\_Guo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robin_guo/32/42297_2.png) [@Robin\_Guo](https://discuss.elastic.co/u/Robin_Guo)\
**Post date:** [September 14, 2018, 4:51am UTC](https://discuss.elastic.co/t/we-lost-some-metrics-from-time-to-time-blocked-by-logstash/148554/1 "2018-09-14T04:51:19Z")

</div>

Dear Elastic ，

We lost some metrics from time to time recently.  
Could someone give some suggestion to solve this problem?

**Error from Metricbeat**

```
2018-09-14T11:03:37+08:00 ERR Failed to publish events caused by: EOF
2018-09-14T11:17:37+08:00 ERR Failed to publish events caused by: EOF
2018-09-14T11:20:37+08:00 ERR Failed to publish events caused by: EOF

```

**Error from Cisco ISR**

> 10.0.0.1 is our HAproxy server, our logstash cluster is behind the HAproxy server.

`Sep 13 2018 12:26:35: %SYS-3-LOGGINGHOST_FAIL: Logging to host 10.0.0.1 port 5050 failed`

**Error from Cisco ASA**

`Sep 13 2018 12:27:29 testfw01 : %ASA-3-414003: TCP Syslog Server NETW:10.0.0.1/5049 not responding, New connections are permitted based on logging permit-hostdown policy`

**configure from Haproxy**

```
vim /etc/haproxy/haproxy.cfg
defaults
    log global
    mode tcp
    option dontlognull
    option redispatch
    retries 3
    #timeout http-request 60s
    timeout queue 300s
    timeout connect 300s
    timeout client 300s
    timeout server 300s
    #timeout http-keep-alive 60s
    timeout check 60s
    maxconn 50000

#---------------------------------------------------------------------
# main frontend which proxys to the backends
#---------------------------------------------------------------------
frontend metricbeat
     bind *:5044
     mode tcp
     timeout client 300s
     default_backend metricbeat2ES
     maxconn 2000
#---------------------------------------------------------------------
# round robin balancing between the various backends
#---------------------------------------------------------------------
backend metricbeat2ES
    mode tcp
    timeout server 300s
    balance roundrobin
    server robinlogstash01 10.0.0.1:5044 check maxconn 2000
    server robinlogstash02 10.0.0.2:5044 check maxconn 2000
    server robinlogstash03 10.0.0.3:5044 check maxconn 2000
    server robinlogstash04 10.0.0.4:5044 check maxconn 2000

```

**logstash pipeline for metricbeat**

```
vim /etc/logstash/conf.d/metricbeat.conf

#logstash for pipeline metricbeat

input {
  beats {
    port => 5044
    client_inactivity_timeout => 300
    #ssl => false
    #ssl_verify_mode => "none"
   # codec => json {
   # charset => "UTF-8"
   #}
  }
}

filter {

  if "_jsonparsefailure" in [tags] {
        drop { }
  }

}

output {
  file => "/tmp/metricbeat.log"

}

```

---

<div class="post-metadata">

**Author:** ![yaauie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yaauie/32/23363_2.png) [@yaauie](https://discuss.elastic.co/u/yaauie)\
**Post date:** [September 14, 2018, 6:21pm UTC](https://discuss.elastic.co/t/we-lost-some-metrics-from-time-to-time-blocked-by-logstash/148554/2 "2018-09-14T18:21:44Z")

</div>

The `EOF` means `End-Of-File`, meaning that whatever Metricbeat was connected to hung up on it, and it was unable to send the metric.

The Beats protocol uses long-lived TCP connections instead of establishing a new connection each time there is an event to be sent. Interrupting this connection could be the cause of the above issues.

I am not an HAProxy expert, but the configuration posted looks like it is attempting to limit the lifetime of the connections to 300s (with `timeout server` and `timeout client` directives), which would be a probable cause for early termination of connections. I would advise setting the `timeout connect` to something like `30s`, and setting the `timeout server` and `timeout client` values much, much higher.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 12, 2018, 6:21pm UTC](https://discuss.elastic.co/t/we-lost-some-metrics-from-time-to-time-blocked-by-logstash/148554/3 "2018-10-12T18:21:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
