# Filebeat: If tcp write or read error then filebeat stops harvesting files until restart

**URL:** <https://discuss.elastic.co/t/filebeat-if-tcp-write-or-read-error-then-filebeat-stops-harvesting-files-until-restart/94590>\
**Category:** Beats\
**Tags:** filebeat\
**Created:** [July 26, 2017, 7:57am UTC](https://discuss.elastic.co/t/filebeat-if-tcp-write-or-read-error-then-filebeat-stops-harvesting-files-until-restart/94590 "2017-07-26T07:57:24Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Alexkl](https://avatars.discourse-cdn.com/v4/letter/a/858c86/32.png) [@Alexkl](https://discuss.elastic.co/u/Alexkl)\
**Post date:** [July 26, 2017, 7:57am UTC](https://discuss.elastic.co/t/filebeat-if-tcp-write-or-read-error-then-filebeat-stops-harvesting-files-until-restart/94590/1 "2017-07-26T07:57:24Z")

</div>

Hello  
I am currently using the elk stack in v5.5 and i have a quite big issue:  
Everytime i have a tcp write or read error i have to restart filebeat because it stops sending messages.

> 2017-07-25T11:14:46+02:00 ERR Failed to publish events caused by: write tcp 163.172.15.176:53430-\>163.172.99.57:5000: write: connection reset by peer  
> 2017-07-25T11:14:46+02:00 ERR Failed to publish events caused by: write tcp 163.172.15.176:53428-\>163.172.99.57:5000: write: connection reset by peer  
> 2017-07-25T11:14:47+02:00 INFO Non-zero metrics in the last 30s: libbeat.logstash.call\_count.PublishEvents=9 libbeat.logstash.publish.read\_bytes=350 libbeat.logstash.publish.write\_bytes=5351704 libbeat.logstash.publish.write\_errors=4 libbeat.logstash.published\_and\_acked\_events=30181 libbeat.logstash.published\_but\_not\_acked\_events=12220 libbeat.publisher.published\_events=30901 publish.events=30181 registrar.states.update=30181 registrar.writes=5  
> 2017-07-25T11:15:17+02:00 INFO Non-zero metrics in the last 30s: libbeat.logstash.call\_count.PublishEvents=2 libbeat.logstash.publish.read\_bytes=4408 libbeat.logstash.publish.write\_bytes=1727247 libbeat.logstash.published\_and\_acked\_events=24497 libbeat.publisher.published\_events=13739 publish.events=12277 registrar.states.update=12277 registrar.writes=2  
> 2017-07-25T11:15:47+02:00 INFO No non-zero metrics in the last 30s  
> 2017-07-25T11:16:17+02:00 INFO No non-zero metrics in the last 30s  
> 2017-07-25T11:16:47+02:00 INFO No non-zero metrics in the last 30s  
> 2017-07-25T11:17:17+02:00 INFO No non-zero metrics in the last 30s  
> 2017-07-25T11:17:47+02:00 INFO No non-zero metrics in the last 30s  
> 2017-07-25T11:18:17+02:00 INFO No non-zero metrics in the last 30s  
> 2017-07-25T11:18:47+02:00 INFO No non-zero metrics in the last 30s  
> 2017-07-25T11:19:17+02:00 INFO No non-zero metrics in the last 30s  
> 2017-07-25T11:19:47+02:00 INFO No non-zero metrics in the last 30s  
> 2017-07-25T11:20:17+02:00 INFO No non-zero metrics in the last 30s  
> When i restart filebeat he push the missing messages and the new ones.

Here is my filebeat conf:

```
filebeat:
  name: "host7"
  spool_size: 16384
  prospectors:
  -
    paths:
      - /var/log/varnish/varnish.log
    input_type: log
    fields_under_root: true
    fields:
      tags: ['json', 'varnish']
      platform: boxes
    document_type: varnish-logs
    close_inactive: 5m

output.logstash:
  hosts: ["ls1:5000","ls2:5000"]
  loadbalance: true
  pipelining: 5
  worker: 2
  bulk_max_size: 8192
  ssl:
     certificate_authorities: ["/etc/filebeat/wildcard.ls.dev.logstash.crt"]

```

I have to send the logs to a distant datacenter, my logstash usualy gets 12k messages/s and i have the same pb on 5 differents plateforms (especialy the ones that don't send a lot of messages)

I started to have an extensive usage of filebeat (and started to loadbalance) since i migrated from 5.4 to 5.5 so i am not sure the problem happened since the 5.5 migration of if it would occur in 5.4.

Thanks !

---

<div class="post-metadata">

**Author:** ![steffens](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steffens/32/79630_2.png) [@steffens](https://discuss.elastic.co/u/steffens)\
**Post date:** [July 26, 2017, 11:04am UTC](https://discuss.elastic.co/t/filebeat-if-tcp-write-or-read-error-then-filebeat-stops-harvesting-files-until-restart/94590/2 "2017-07-26T11:04:58Z")

</div>

Can you run filebeat with debug logs enabled?

```auto
logging.level: debug
logging.selectors: ["output", "logstash"]

```

Logstash will print `close connection` and `connect` messages on reconnect. Plus debug messages on number of events send. The 'output' selector might add messages like: `add non-published events back into pipeline` and `async bulk publish success`.

Please note, upon failure the client uses exponential backoff (but only up to 1 minute).

When you can kill filebeat with `kill -ABRT <pid>`, it will print a stack trace. Alternatively you can start filebeat with `-httpprof :6060` and get a stack-trace of all go-routines via `curl http://localhost:6060/debug/pprof/goroutine`.

Having multiple stack-traces + debug logs can be helpful trying to identify if/where the outputs actually might hang.

The spool\_size is only twice bulk\_max\_size. Why have 2 workers, with pipelining set to 5? Does the problem still occur If you set pipelining to 0?

---

<div class="post-metadata">

**Author:** ![Alexkl](https://avatars.discourse-cdn.com/v4/letter/a/858c86/32.png) [@Alexkl](https://discuss.elastic.co/u/Alexkl)\
**Post date:** [July 26, 2017, 12:05pm UTC](https://discuss.elastic.co/t/filebeat-if-tcp-write-or-read-error-then-filebeat-stops-harvesting-files-until-restart/94590/3 "2017-07-26T12:05:59Z")

</div>

My logstash servers are not in the same datacenter so ... while reading the configuration about pipeling i understood that setting pipelining (i put a random number for test) would permit to push other batches without waiting for an ack.

For the bulk\_max\_size and spool\_size it's seems i misunderstood the lock step. Workers are not needed (events to N hosts in lock-step).

I add the debugging and httprof and will let you know when it break again

Thanks.

---

<div class="post-metadata">

**Author:** ![Alexkl](https://avatars.discourse-cdn.com/v4/letter/a/858c86/32.png) [@Alexkl](https://discuss.elastic.co/u/Alexkl)\
**Post date:** [July 27, 2017, 2:41pm UTC](https://discuss.elastic.co/t/filebeat-if-tcp-write-or-read-error-then-filebeat-stops-harvesting-files-until-restart/94590/4 "2017-07-27T14:41:30Z")

</div>

It looks like it s the pipelining option.  
I removed it, i wiil see in the next days if it breaks again.  
Thank you.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 24, 2017, 2:41pm UTC](https://discuss.elastic.co/t/filebeat-if-tcp-write-or-read-error-then-filebeat-stops-harvesting-files-until-restart/94590/5 "2017-08-24T14:41:51Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
