# High rate Remote Syslog into filebeat

**URL:** <https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092>\
**Category:** Beats\
**Created:** [October 27, 2015, 5:25pm UTC](https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092 "2015-10-27T17:25:04Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![matt.koivisto](https://avatars.discourse-cdn.com/v4/letter/m/7ba0ec/32.png) [@matt.koivisto](https://discuss.elastic.co/u/matt.koivisto)\
**Post date:** [October 27, 2015, 5:25pm UTC](https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092/1 "2015-10-27T17:25:04Z")

</div>

I had a setup working, using logstash with udp input and rabbitmq output, to consume a high rate of remote syslog messages and publish it into elastic search (with another logstash instance using rabbitmq as input, and output to elastic search). I found that java was using 3 cores at 100% to handle the load (though it was handling it).

I am trying to convert to using filebeat as the transport now, instead of rabbitmq, with the hope of filebeat being able to work with much lower performance hit.

So before I had:

syslog src -\> udp 514 -\> logstash (local) -\> rabbitmq (AWS) -\> logstash (AWS) -\> ES (AWS)

And I am now trying to move to:

syslog src -\> udp 514 -\> rsyslog -\> /var/log/file.log -\> filebeat -\> logstash (AWS) -\> ES (AWS)

but I am finding that most messages are being dropped (likely due to the unnecessary file io). I also tried a pipe via mkfifo, but IIRC filebeat didn't load.

So the question is, is there any way for filebeat (or packetbeat if it can pull the message for that matter) to listen directly on port 514? This would then allow me to publish the syslog data w/o the file i/o overhead.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 28, 2015, 2:35am UTC](https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092/2 "2015-10-28T02:35:45Z")

</div>

Filebeat will only ever read files. Packetbeat probably won't work as it's application level.

Have you tried redis instead of MQ?

---

<div class="post-metadata">

**Author:** ![ruflin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ruflin/32/3116_2.png) [@ruflin](https://discuss.elastic.co/u/ruflin)\
**Post date:** [October 28, 2015, 7:19am UTC](https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092/3 "2015-10-28T07:19:21Z")

</div>

Filebeat currently does not support mkfifo. It supports stdin in case this helps.

Where in the above chain are the packages dropped? Why don't you install filebeat directly on the machine of "syslog src"?

---

<div class="post-metadata">

**Author:** ![matt.koivisto](https://avatars.discourse-cdn.com/v4/letter/m/7ba0ec/32.png) [@matt.koivisto](https://discuss.elastic.co/u/matt.koivisto)\
**Post date:** [October 28, 2015, 1:21pm UTC](https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092/4 "2015-10-28T13:21:04Z")

</div>

> [@warkolm](#):
>
> MQ

I have not tried redis, is there reason to believe that the rabbitmq output from logstash is what is consuming the most CPU in java?

---

<div class="post-metadata">

**Author:** ![matt.koivisto](https://avatars.discourse-cdn.com/v4/letter/m/7ba0ec/32.png) [@matt.koivisto](https://discuss.elastic.co/u/matt.koivisto)\
**Post date:** [October 28, 2015, 1:29pm UTC](https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092/5 "2015-10-28T13:29:30Z")

</div>

@ruflin

I can't install filebeat directly on the src as it's a 3rd party embedded system, I can only configure it to remote syslog to an arbitrary ip/port.

I believe the messages are being dropped between rsyslog and the kernel as rsyslog can't consume fast enough with trying to write to the filesystem. I will try running an application to print the messages to stdout and pipe that to filebeat consuming from stdin, that should eliminate any filesystem io bottlenecks.

---

<div class="post-metadata">

**Author:** ![matt.koivisto](https://avatars.discourse-cdn.com/v4/letter/m/7ba0ec/32.png) [@matt.koivisto](https://discuss.elastic.co/u/matt.koivisto)\
**Post date:** [October 28, 2015, 3:06pm UTC](https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092/6 "2015-10-28T15:06:38Z")

</div>

I wrote a small go program to listen to port 514/udp and just print the received bytes to stdout. I run that piping it to filebeat configured to listen to stdin and forward to logstash. The performance is fantastic, CPU sits around 50% of one core, which is great compared to my first solution.

The only hiccup was filebeat puts the contents from stdin into a field called "text" instead of "message", so I had to write a mutate filter in logstash before my other filters to get the expected behavior:

```
if [source] == "-" {
  mutate {
    remove_field => ["message"]
  }
  mutate {
    rename => { "text" => "message" }
  }
}

```

Thanks for your help/suggestions!

---

<div class="post-metadata">

**Author:** ![tudor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tudor/32/3753_2.png) [@tudor](https://discuss.elastic.co/u/tudor)\
**Post date:** [October 28, 2015, 3:39pm UTC](https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092/7 "2015-10-28T15:39:50Z")

</div>

Nice, it might make sense to transform this into a "Syslogbeat" so you only have to run one process. What do you think?

---

<div class="post-metadata">

**Author:** ![matt.koivisto](https://avatars.discourse-cdn.com/v4/letter/m/7ba0ec/32.png) [@matt.koivisto](https://discuss.elastic.co/u/matt.koivisto)\
**Post date:** [October 28, 2015, 4:02pm UTC](https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092/8 "2015-10-28T16:02:46Z")

</div>

I think any of the existing logstash input filters could have a case made for them to have a corresponding "beat" implementation, if running logstash on the remote machine was too resource intensive for an individual's use case.

Ultimately, lumberjack/logstash-forwarder was born out of the same need, so really the question becomes what is the desired vision for elastic's recommended deployment model? Logstash on all the "source" nodes, with "beats" to replace them iff performance is an issue? Or recommend always deploy an appropriate "beat" and only deploy logstash on a central well provisioned server?

If the later, then from an architectural perspective, we are basically recommending re-writing all the input filters of logstash in go instead of java, and breaking them out into separate executables (which someone is going to suggest become a single executable again 🙂 )

Anyways, back to your question, in the short term if anyone else has a similar need, it makes a lot of sense IMO to make a syslogbeat. I'll try to get some spare time to look at contributing back to help 😄

---

<div class="post-metadata">

**Author:** ![ruflin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ruflin/32/3116_2.png) [@ruflin](https://discuss.elastic.co/u/ruflin)\
**Post date:** [October 29, 2015, 7:28am UTC](https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092/9 "2015-10-29T07:28:38Z")

</div>

@matt.koivisto Which version of filebeat are you using? The reason I ask because about 7 days ago we changed from text to message and beat4 should have this change already inside: [https://github.com/elastic/filebeat/blob/89410957be0163fe2e999cab0efffdeb6ab926c5/input/file.go#L56](https://github.com/elastic/filebeat/blob/89410957be0163fe2e999cab0efffdeb6ab926c5/input/file.go#L56)

---

<div class="post-metadata">

**Author:** ![ruflin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ruflin/32/3116_2.png) [@ruflin](https://discuss.elastic.co/u/ruflin)\
**Post date:** [October 29, 2015, 7:35am UTC](https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092/10 "2015-10-29T07:35:40Z")

</div>

@matt.koivisto About your architecture question: Beats will not replace Logstash on all source nodes because Logstash has lots of additional capabilities. But in some cases it will definitively as described in your second option, especially as soon as we support filtering and multiline. Beat should be used if a lightweight solution is need to "just" forward the data without processing it. The goal is to keep beats as lightweight as possible so also the resource usage stays low. I'm quite sure you are right and lots if different "input" beats will be also created by the community. As the beats project is still a very young project it is still open to see where this will lead.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 9:58pm UTC](https://discuss.elastic.co/t/high-rate-remote-syslog-into-filebeat/33092/11 "2017-07-05T21:58:32Z")

</div>


