# Problems receiving data from large number of logstash-forwarders on logstash

**URL:** <https://discuss.elastic.co/t/problems-receiving-data-from-large-number-of-logstash-forwarders-on-logstash/33584>\
**Category:** Logstash\
**Created:** [November 3, 2015, 12:34am UTC](https://discuss.elastic.co/t/problems-receiving-data-from-large-number-of-logstash-forwarders-on-logstash/33584 "2015-11-03T00:34:29Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Medelwr](https://avatars.discourse-cdn.com/v4/letter/m/ccd318/32.png) [@Medelwr](https://discuss.elastic.co/u/Medelwr)\
**Post date:** [November 3, 2015, 12:34am UTC](https://discuss.elastic.co/t/problems-receiving-data-from-large-number-of-logstash-forwarders-on-logstash/33584/1 "2015-11-03T00:34:29Z")

</div>

Hi,

I have 60+ clients that I am trying to connect to my logstash instance. They are sending logs on a short interval using LSF (Lumberjack).

Versions:  
Logstash 2.0  
Elasticsearch 2.0  
Kibana 4.2  
Logstash-forwarder 0.4.0

Logstash setup:  
port 5000 used to receive from all LSF data.  
20 workers  
5g heap

input = lumberjack on port 5000  
filters = basic grok and timestamp filters  
output = file with filters + elasticsearch

The problem:

When I start up logstash all of the clients establish their connections but after a delay period (~15 seconds, the default LSF timeout period) the majority of them change state to CLOSE\_WAIT or SYN\_RECV. A handful remain ESTABLISHED and continue to send data without any problems. This is usually only 5-9 connections that remain ESTABLISHED at best and processing correctly (Not always the same connections out of the 60 that remain so I have crossed off specific LSF clients being the problem)

There is data in the RECV\_Q of the CLOSE\_WAIT connections. i.e. RECV\_Q is not zero.  
The CLOSE\_WAIT and SYN\_RECV connections have no PID attached (PID is '-' when running netstat -p as root) so I cannot close them manually as far as I am aware.  
The SYN\_RECV connections on the elk server are ESTABLISHED on the client side and have data in their SEND\_Q

logstash logs intermittently shows ":message=\>"Lumberjack input: the pipeline is blocked, temporary refusing new connection."

If this occurs I simply restart logstash. It does not happen consistently. I don't notice any discernible difference between runs when this line appears in the log and when it does not.

Doing a tcpdump on port 5000 shows constant traffic on the port when there shouldn't be between logging intervals. I am assuming it is clients attempting to connect.

Every connection that isn't ESTABLISHED on the main elk server usually has multiple duplicate connections showing on netstat..  
e.g. 1 client could have 3 CLOSE\_WAIT connections with data in their RECV\_Q and a SYN\_RECV

LSF is installed on the main elk server and is usually in the ESTABLISHED state but there is usually data in it's SEND\_Q and it is not being processed.

Before LSF was rolled out to all clients it was tested using 5 clients of different OS and worked fine consistently.

Originally I was running logstash using default config. After increasing increasing heap from 500m -\> 5g I noticed an increase in the ESTABLISHED connections initially and also when I increased the worker counter but it normally kills of connections eventually and stabilizes around 3 ESTABLISHED.

Can logstash handle this many connections?

Any help on this would be much appreciated.

Thanks.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [November 3, 2015, 6:06am UTC](https://discuss.elastic.co/t/problems-receiving-data-from-large-number-of-logstash-forwarders-on-logstash/33584/2 "2015-11-03T06:06:13Z")

</div>

Are you doing all your ingestion/receipt, filtering and output in a single instance?  
Cause it could be something further down the chain. Adding a broker may help alleviate this.

---

<div class="post-metadata">

**Author:** ![Thorsten\_Nickel](https://avatars.discourse-cdn.com/v4/letter/t/3be4f8/32.png) [@Thorsten\_Nickel](https://discuss.elastic.co/u/Thorsten_Nickel)\
**Post date:** [November 3, 2015, 8:24am UTC](https://discuss.elastic.co/t/problems-receiving-data-from-large-number-of-logstash-forwarders-on-logstash/33584/3 "2015-11-03T08:24:58Z")

</div>

You should be aware that Logstash only keeps 20 events maximum in it's processing pipeline, after which the pipeline gets into a blocking state. To mitigate this, either optimize your processing pipeline and/or use more worker processes with the new -w command line option.

Hope to help,  
Thorsten

---

<div class="post-metadata">

**Author:** ![Medelwr](https://avatars.discourse-cdn.com/v4/letter/m/ccd318/32.png) [@Medelwr](https://discuss.elastic.co/u/Medelwr)\
**Post date:** [November 3, 2015, 10:03am UTC](https://discuss.elastic.co/t/problems-receiving-data-from-large-number-of-logstash-forwarders-on-logstash/33584/4 "2015-11-03T10:03:13Z")

</div>

Hi, thanks for the replies,

@warkolm I am doing all my processing on a single instance yes. I was under the impression that I would not need a broker such as Redis if I am using LSF as it should stop sending when it is blocked but not break its connection. Do you believe that it would help?

@Thorsten_Nickel I am aware of the event cap. Right now I am using 20 workers and 5g of heap, do you believe I will need more? Or be able to support more?

Are there any reasons why there would be multiple duplicate connections for each client?

Also with half of my ESTABLISHED connections some data gets through but then data sits in the RECV\_Q...  
Once one client is processed it doesn't seem to release its hold on logstash

---

<div class="post-metadata">

**Author:** ![Thorsten\_Nickel](https://avatars.discourse-cdn.com/v4/letter/t/3be4f8/32.png) [@Thorsten\_Nickel](https://discuss.elastic.co/u/Thorsten_Nickel)\
**Post date:** [November 3, 2015, 10:36am UTC](https://discuss.elastic.co/t/problems-receiving-data-from-large-number-of-logstash-forwarders-on-logstash/33584/5 "2015-11-03T10:36:59Z")

</div>

Hi,

if you would test your setup with an increased amoutn of workers, say 40 or 50, and see a different behaviour, then you might be on the right track.  
I would also say, 5G of heap for the Logstash instance should be quite good.

Hope to help,  
Thorsten

---

<div class="post-metadata">

**Author:** ![Medelwr](https://avatars.discourse-cdn.com/v4/letter/m/ccd318/32.png) [@Medelwr](https://discuss.elastic.co/u/Medelwr)\
**Post date:** [November 3, 2015, 10:51am UTC](https://discuss.elastic.co/t/problems-receiving-data-from-large-number-of-logstash-forwarders-on-logstash/33584/6 "2015-11-03T10:51:47Z")

</div>

@Thorsten_Nickel Increased workers to 60. Noticed a decrease in stable ESTABLISHED connections i.e. connections that have 0 data in RECV\_Q and SEND\_Q

Increase in total ESTABLISHED connections with new workers but the majority have data in their RECV\_Q that logstash isn't picking up.

Like I said, I don't think logstash is working correctly. Right now I only have 4 stable connections that have finished sending their data but logstash is ignoring all the other connections. It doesn't seem to be rotating between clients as it should.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [November 3, 2015, 8:40pm UTC](https://discuss.elastic.co/t/problems-receiving-data-from-large-number-of-logstash-forwarders-on-logstash/33584/7 "2015-11-03T20:40:08Z")

</div>

It might be worth raising an issue on GH with as much info as possible.

---

<div class="post-metadata">

**Author:** ![Medelwr](https://avatars.discourse-cdn.com/v4/letter/m/ccd318/32.png) [@Medelwr](https://discuss.elastic.co/u/Medelwr)\
**Post date:** [November 4, 2015, 11:16am UTC](https://discuss.elastic.co/t/problems-receiving-data-from-large-number-of-logstash-forwarders-on-logstash/33584/8 "2015-11-04T11:16:53Z")

</div>

@warkolm I looked into using redis as a broker but it appears that logstash-forwarder does not support redis at this time. is this correct?

What alternative would you recommend? I believe I do need something to take the pressure off logstash

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [November 5, 2015, 2:50am UTC](https://discuss.elastic.co/t/problems-receiving-data-from-large-number-of-logstash-forwarders-on-logstash/33584/9 "2015-11-05T02:50:57Z")

</div>

You can go LSF \> LS \> redis \> LS \> ES, and run both LS instances on the same host.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 5:23am UTC](https://discuss.elastic.co/t/problems-receiving-data-from-large-number-of-logstash-forwarders-on-logstash/33584/10 "2017-07-06T05:23:58Z")

</div>


