# Losing messages during high traffic rate

**URL:** <https://discuss.elastic.co/t/losing-messages-during-high-traffic-rate/56449>\
**Category:** Logstash\
**Created:** [July 26, 2016, 9:00pm UTC](https://discuss.elastic.co/t/losing-messages-during-high-traffic-rate/56449 "2016-07-26T21:00:04Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![sharon.c](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sharon.c/32/16076_2.png) [@sharon.c](https://discuss.elastic.co/u/sharon.c)\
**Post date:** [July 26, 2016, 9:00pm UTC](https://discuss.elastic.co/t/losing-messages-during-high-traffic-rate/56449/1 "2016-07-26T21:00:04Z")

</div>

I am doing log analysis using Filebeat (1.2) -\> logstash(2.3) -\> Elasticsearch (2.3)

I have 4 filebeat instances, 4 logstash instances (6 cores each) , and Elasticsearch cluster (8 cores, 64G RAM) of 2 nodes

In Filebeat.yml logstash setting is as such:

> 

filebeat:  
prospectors:  
-  
paths:  
- /var/log/filebeat/_/_.json  
encoding: utf-8  
input\_type: log  
ignore\_older: 10m  
scan\_frequency: 1s  
exclude\_lines: ["^$"]  
spool\_size: 3072  
registry\_file: .filebeat

> output:  
> logstash:  
> enabled: true  
> hosts: ["logstash1:5044","logstash2:5044"]  
> worker: 8  
> loadbalance: true  
> index: elkstats\_record

Logstash’s Elasticsearch output setting is like this

```
    elasticsearch {
       hosts => ["node1", "node2"]
       index => "records_%{+YYYY.MM.dd}" # generate 1 index every month
       template_name => "template"
       document_id => "%{[@metadata][computed_id]}" # set documented
       workers => 2
       flush_size => 3500
    }

```

Logstash host environment variable for heap size is set to $LS\_HEAP\_SIZE = 2048M

Normally, if the traffic is 3000-4000 msg/sec, when the traffic exceeds 6000-7000 msg/sec, logstash side will have these repeated messages:

> 

CircuitBreaker::rescuing exceptions {:name=\>"Beats input", :exception=\>LogStash::Inputs::Beats::InsertingToQueueTakeTooLong, :level=\>:warn}  
Beats input: The circuit breaker has detected a slowdown or stall in the pipeline, the input is closing the current connection and rejecting new connection until the pipeline recover. {:exception=\>LogStash::Inputs::BeatsSupport::CircuitBreaker::HalfOpenBreaker, :level=\>:warn}

In system monitor tool, it shows logstash instance uses 50 - 80% of the cpu during the high traffic hours, but only 800MB RAM.

Elasticsearch side does not have obvious resource shortage. CPU usage is 15%-26%, memory usage is 26%

On elasticsearch error logs, it shows: org.apache.lucene.store.AlreadyClosedException, I have more details in this link

> [@org.apache.lucene.store.AlreadyClosedException](https://discuss.elastic.co/t/org-apache-lucene-store-alreadyclosedexception/56537):
>
> I am doing log analysis using Filebeat (1.2) -\> logstash(2.3) -\> Elasticsearch (2.3) I have 4 filebeat instances, 4 logstash instances (6 cores each, 8G RAM) , and Elasticsearch cluster (8 cores, 64G RAM) of 2 nodes When the traffic was pretty high \>6000 msg/sec, 10% to 20% of messges are lost. The errors in elasticsearch log shows: [2016-07-10 01:47:28,790][DEBUG][action.admin.cluster.node.stats] [Melasticsearch] failed to execute on node [e8rsvmGRReWzFOCem4Fxeg] RemoteTransportException[…

It looks like logstash is not processing messages fast enough, and filebeat takes long time to insert to logstash's queue. Then we end up losing the messages which are not inserted from filebeat to logstash queue.

Is there any way to tune logstash or filebeat (flush\_size, workers) so that Logstash will process messages faster?  
Also how to let logstash make use of all the LS\_HEAP\_SIZE of 2G to cache the unprocessed messages, instead of only 800MB?  
Is there any other way to prevent logstash from losing messages?

---

<div class="post-metadata">

**Author:** ![anhlqn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anhlqn/32/5454_2.png) [@anhlqn](https://discuss.elastic.co/u/anhlqn)\
**Post date:** [July 27, 2016, 6:33pm UTC](https://discuss.elastic.co/t/losing-messages-during-high-traffic-rate/56449/2 "2016-07-27T18:33:32Z")

</div>

How many Logstash filter workers for each LS instance?

Have you tried to increase the pipeline batch size for LS with the -b switch? The default one of `-b 125` is pretty low. Try to increase it gradually to see if it helps. Mine is set at `-b 1500` or `-b 3000` depending on the throughput.

> [@sharon.c](#):
>
> workers =\> 2

Try increasing this number to match the number of LS worker filters. It shoud be at least 6 since your LS instance has 6 CPU cores.

How many LS filters do you have? Too many filters hurt LS processing capacity.

---

<div class="post-metadata">

**Author:** ![sharon.c](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sharon.c/32/16076_2.png) [@sharon.c](https://discuss.elastic.co/u/sharon.c)\
**Post date:** [July 27, 2016, 6:35pm UTC](https://discuss.elastic.co/t/losing-messages-during-high-traffic-rate/56449/3 "2016-07-27T18:35:47Z")

</div>

Thank you for the prompt suggestions. Setting the batch size is very effective, there is no lost message after setting the appropriate batch size.

---

<div class="post-metadata">

**Author:** ![sharon.c](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sharon.c/32/16076_2.png) [@sharon.c](https://discuss.elastic.co/u/sharon.c)\
**Post date:** [July 27, 2016, 7:08pm UTC](https://discuss.elastic.co/t/losing-messages-during-high-traffic-rate/56449/4 "2016-07-27T19:08:09Z")

</div>

Can you also provide some link of article about how to optimise LS processes?

---

<div class="post-metadata">

**Author:** ![anhlqn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anhlqn/32/5454_2.png) [@anhlqn](https://discuss.elastic.co/u/anhlqn)\
**Post date:** [July 27, 2016, 7:12pm UTC](https://discuss.elastic.co/t/losing-messages-during-high-traffic-rate/56449/5 "2016-07-27T19:12:08Z")

</div>

This may be helpful [https://www.elastic.co/guide/en/logstash/2.3/pipeline.html](https://www.elastic.co/guide/en/logstash/2.3/pipeline.html). If possible, set Logstash output to /dev/null and test the filter workers and pipeline batch size first. Use LS metrics plugin to see how many msgs LS can handle.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:46am UTC](https://discuss.elastic.co/t/losing-messages-during-high-traffic-rate/56449/6 "2017-07-06T04:46:12Z")

</div>


