# Filebeat / Kafka bug?

**URL:** <https://discuss.elastic.co/t/filebeat-kafka-bug/185843>\
**Category:** Beats\
**Tags:** filebeat\
**Created:** [June 14, 2019, 11:33am UTC](https://discuss.elastic.co/t/filebeat-kafka-bug/185843 "2019-06-14T11:33:46Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Rob3](https://avatars.discourse-cdn.com/v4/letter/r/a4c791/32.png) [@Rob3](https://discuss.elastic.co/u/Rob3)\
**Post date:** [June 14, 2019, 11:33am UTC](https://discuss.elastic.co/t/filebeat-kafka-bug/185843/1 "2019-06-14T11:33:46Z")

</div>

Filebeat Version: 7.x (testing on 7.1.1 and 7.1.2)  
Kafka Version: Azure Event Hubs Kafka surface

Logstash and Fluentd both work with Event Hubs Kafka interface, Filebeat not so much.

For some reason it appears the Event Hub is not happy with how filebeat is authenticating, at a guess. As seen in the log snippet below, it appears the EH is closing the connection abruptly. This happens many times per minutes, and seems to block Filebeat for some time.

Every approx 1-3 minutes Filebeat successfully publishes 1 or 2 thousand events to the topic, and then enters an endless loop of the below auth/network drop cycle for another 1-3 minutes or more.

Since other producers do not experience this (logstash nor fluentd) it seems a bug in filebeat, or the Kafka library it uses, that's triggering Event Hubs to RST the connection.

Anyone have any ideas?

Filebeat Config

```auto
  output:
    kafka:
      hosts:
        - "azure_eh_kafka.servicebus.windows.net:9093"
      topic: test
      version: "1.0.0"
      ssl.enabled: true
      username: "$ConnectionString"
      password: "Endpoint=sb://azure_eh_kafka.servicebus.windows.net/;SharedAccessKeyName=<KEYNAME>;SharedAccessKey=<KEY>"
      compression: none

```

Filebeat logs

```auto
{"level":"info","timestamp":"2019-06-14T10:22:02.343Z","caller":"kafka/log.go:53","message":"kafka message: Successful SASL handshake"}
{"level":"info","timestamp":"2019-06-14T10:22:02.344Z","caller":"kafka/log.go:53","message":"SASL authentication successful with broker azure_eh_kafka.servicebus.windows.net:9093:4 - [0 0 0 0]\n"}
{"level":"info","timestamp":"2019-06-14T10:22:02.344Z","caller":"kafka/log.go:53","message":"Connected to broker at azure_eh_kafka.servicebus.windows.net:9093 (registered as #0)\n"}
{"level":"info","timestamp":"2019-06-14T10:22:02.375Z","caller":"kafka/log.go:53","message":"producer/broker/0 state change to [closing] because write tcp 10.32.nnn.nnn:41188->13.69.nnn.nnn:9093: write: connection reset by peer\n"}
{"level":"info","timestamp":"2019-06-14T10:22:02.375Z","caller":"kafka/log.go:53","message":"Error while closing connection to broker azure_eh_kafka.servicebus.windows.net:9093: write tcp 10.32.nnn.nnn:41188->13.69.nnn.nnn:9093: write: broken pipe\n"}

{"level":"info","timestamp":"2019-06-14T10:22:02.496Z","caller":"kafka/log.go:53","message":"SASL authentication successful with broker azure_eh_kafka.servicebus.windows.net:9093:4 - [0 0 0 0]\n"}
{"level":"info","timestamp":"2019-06-14T10:22:02.496Z","caller":"kafka/log.go:53","message":"Connected to broker at azure_eh_kafka.servicebus.windows.net:9093 (registered as #0)\n"}
{"level":"info","timestamp":"2019-06-14T10:22:02.763Z","caller":"kafka/log.go:53","message":"producer/broker/0 state change to [closing] because write tcp 10.32.nnn.nnn:41190->13.69.nnn.nnn:9093: write: connection reset by peer\n"}
{"level":"info","timestamp":"2019-06-14T10:22:02.763Z","caller":"kafka/log.go:53","message":"Error while closing connection to broker azure_eh_kafka.servicebus.windows.net:9093: write tcp 10.32.nnn.nnn:41190->13.69.nnn.nnn:9093: write: broken pipe\n"}
{"level":"debug","timestamp":"2019-06-14T10:22:02.767Z","logger":"kafka","caller":"kafka/client.go:251","message":"finished kafka batch"}
{"level":"debug","timestamp":"2019-06-14T10:22:02.767Z","logger":"kafka","caller":"kafka/client.go:265","message":"Kafka publish failed with: write tcp 10.32.nnn.nnn:41190->13.69.nnn.nnn:9093: write: connection reset by peer"}
{"level":"debug","timestamp":"2019-06-14T10:22:02.767Z","logger":"kafka","caller":"kafka/client.go:251","message":"finished kafka batch"}
{"level":"debug","timestamp":"2019-06-14T10:22:02.767Z","logger":"kafka","caller":"kafka/client.go:265","message":"Kafka publish failed with: write tcp 10.32.nnn.nnn:41190->13.69.nnn.nnn:9093: write: connection reset by peer"}

```

---

<div class="post-metadata">

**Author:** ![Rob3](https://avatars.discourse-cdn.com/v4/letter/r/a4c791/32.png) [@Rob3](https://discuss.elastic.co/u/Rob3)\
**Post date:** [June 17, 2019, 6:31am UTC](https://discuss.elastic.co/t/filebeat-kafka-bug/185843/2 "2019-06-17T06:31:46Z")

</div>

This appears to be caused by:

[https://github.com/elastic/beats/pull/12254/files#diff-ea065a113b39ad622b84e2dcedd8ffeeR215](https://github.com/elastic/beats/pull/12254/files#diff-ea065a113b39ad622b84e2dcedd8ffeeR215)

Compile from master and set `bulk_max_size` to something reasonable and it appears to be working perfectly now.

Had tried setting `bulk_max_size` in 7.1.x versions, but that didn't help.

---

<div class="post-metadata">

**Author:** ![pierhugues](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pierhugues/32/48383_2.png) [@pierhugues](https://discuss.elastic.co/u/pierhugues)\
**Post date:** [June 19, 2019, 6:57pm UTC](https://discuss.elastic.co/t/filebeat-kafka-bug/185843/3 "2019-06-19T18:57:26Z")

</div>

Looking at the code overriding the `bulk_max_size` value in the configuration it should have the same behavior as having the value hardcoded. I am surprised you don't get the same behavior because all settings are read and unpacked at the same time?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 17, 2019, 7:30pm UTC](https://discuss.elastic.co/t/filebeat-kafka-bug/185843/5 "2019-07-17T19:30:37Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
