# Logstash high load and CPU usage

**URL:** <https://discuss.elastic.co/t/logstash-high-load-and-cpu-usage/357079>\
**Category:** Logstash\
**Created:** [April 9, 2024, 2:07pm UTC](https://discuss.elastic.co/t/logstash-high-load-and-cpu-usage/357079 "2024-04-09T14:07:51Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![pk92](https://avatars.discourse-cdn.com/v4/letter/p/f1d935/32.png) [@pk92](https://discuss.elastic.co/u/pk92)\
**Post date:** [April 9, 2024, 2:07pm UTC](https://discuss.elastic.co/t/logstash-high-load-and-cpu-usage/357079/1 "2024-04-09T14:07:51Z")

</div>

Hi,

I am running an Elastic Stack with multiple Logstash servers in different networks to aggregate, filter and forward the logs. For some time now I have the problem that some of these Logstash nodes regularly have a very high load and CPU usage. When I restart the Logstash service, it is all fine again for a while. You can see this behavior in this screenshot of the Elastic Monitoring

 ![Screenshot 2024-04-09 at 14.56.42](https://us1.discourse-cdn.com/elastic/original/3X/0/1/012f3c1a425e81a6207aa92336d935a191a42807.png)

I already searched for quite a bit on this problem, but still have no clue what exactly is causing these increased loads. I would be very happy if anyone has any idea what might be the cause and could point me in a direction!

Some more information on my configuration:

The specs of the nodes:

```auto
4 vCPUs
8GB RAM
6GB Heap

```

Logstash config:

```auto
# Ansible managed

pipeline.ordered: auto
path:
  data: /var/lib/logstash
  logs: /var/log/logstash

xpack.monitoring.enabled: false
monitoring.enabled: false
monitoring.cluster_uuid: "uuid"

xpack.management:
  enabled: true
  elasticsearch:
    hosts: ["https://elastic1:9200", "https://elastic2:9200", "https://elastic3:9200"]
    username: "logstash_internal"
    password: "password"
    ssl:
      verification_mode: certificate
      certificate_authority: /etc/logstash/certs/elastic-stack-ca.pem
  logstash.poll_interval: "5s"
  pipeline.id: ["lan"]

```

Pipeline:

```auto
input {
  elastic_agent {
    host => "${IP_ADDRESS}"
    port => 5044
    ssl_enabled => true
    ssl_certificate => "/etc/logstash/certs/logstash.crt.pem"
    ssl_key => "/etc/logstash/certs/logstash.key.pem"
    ssl_client_authentication => "none"
    type => "elastic_agent"
  }
  gelf {
    host => "${IP_ADDRESS}"
    use_udp => false
    use_tcp => true
    port => 12201
    type => "gelf"
  }
  syslog {
    host => "${IP_ADDRESS}"
    port => 10514
    type => "syslog"
    proxy_protocol => true
    ecs_compatibility => "v8"
  }
}

filter {
  if [host][hostname] in ["server1", "server2", "server3", "server4"] {
    mutate {
      add_field => {
        "[data_stream][type]" => "logs"
        "[data_stream][dataset]" => "webservices"
        "[data_stream][namespace]" => "dev"
      }
    }
  }
  if [host][hostname] in ["server5", "server6", "server7", "server8", "server9"] {
    mutate {
      add_field => {
        "[data_stream][type]" => "logs"
        "[data_stream][dataset]" => "webservices"
        "[data_stream][namespace]" => "prod"
      }
    }
  }
}

output {
  if ([data_stream][type] and [data_stream][type] != "" ) and ([data_stream][dataset] and [data_stream][dataset] != "" ) {
    elasticsearch {
      hosts => ["https://elastic1:9200", "https://elastic2:9200", "https://elastic3:9200"]
      data_stream => "true"
      user => "logstash_internal"
      password => "password"
      ssl_enabled => "true"
      ssl_verification_mode => "full"
      ssl_certificate_authorities => "/etc/logstash/certs/elastic-stack-ca.pem"
    }
  } else if [type] == "syslog" {
    elasticsearch {
      hosts => ["https://elastic1:9200", "https://elastic2:9200", "https://elastic3:9200"]
      ilm_enabled => true
      ilm_rollover_alias => "syslog"
      ilm_pattern => "{now/d}-000001"
      ilm_policy => "syslog"
      user => "logstash_internal"
      password => "password"
      ssl_enabled => "true"
      ssl_verification_mode => "full"
      ssl_certificate_authorities => "/etc/logstash/certs/elastic-stack-ca.pem"
    }
  }
  if [type] == "gelf" {
    elasticsearch {
      hosts => ["https://elastic1:9200", "https://elastic2:9200", "https://elastic3:9200"]
      ilm_enabled => true
      ilm_rollover_alias => "gelf"
      ilm_pattern => "{now/d}-000001"
      ilm_policy => "gelf"
      user => "logstash_internal"
      password => "password"
      ssl_enabled => "true"
      ssl_verification_mode => "full"
      ssl_certificate_authorities => "/etc/logstash/certs/elastic-stack-ca.pem"
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [April 9, 2024, 2:27pm UTC](https://discuss.elastic.co/t/logstash-high-load-and-cpu-usage/357079/2 "2024-04-09T14:27:08Z")

</div>

Are you using persistent queue or memory queues? It is not clear.

Is logstash the only service running on this server?

Do you have anything in the logs?

---

<div class="post-metadata">

**Author:** ![pk92](https://avatars.discourse-cdn.com/v4/letter/p/f1d935/32.png) [@pk92](https://discuss.elastic.co/u/pk92)\
**Post date:** [April 9, 2024, 3:15pm UTC](https://discuss.elastic.co/t/logstash-high-load-and-cpu-usage/357079/3 "2024-04-09T15:15:06Z")

</div>

Thanks for the quick reply. I gladyl provide more information:

My Logstash nodes are using the default memory queues

Apart from Logstash the only noteworthy services running on the servers would be the Elastic Agent and HAproxy/Keepalived for some load balancing of the syslog input between two Logstash nodes. But these processes hardly use any CPU or memory according to htop.

The only messages I get in the logstash-plain.log are like these:

```auto
[2024-04-09T17:04:27,340][INFO][logstash.inputs.syslog][lan][1fa9b49c9429d1d0a7fa6399b888deb8ae2ac1a205deaf8ccf38ff44b5e2ed5b] new connection {:client=>"10.20.30.40:52270"}

```

```auto
[2024-04-09T17:04:26,503][WARN][logstash.outputs.elasticsearch][lan][8da42123ea1df5ae3882151c817efd71ae382fac24ed2de9d61bc5c93419a5f5] Failed action {:status=>409, :action=>["create", {:_id=>nil, :_index=>"metrics-apache_tomcat.cache-default", :routing=>nil}, {"service"=>{"type"=>"prometheus", "address"=>"http://localhost:9090/metrics"}, "elastic_agent"=>{"id"=>"522874b7-bd30-487c-8c9f-a1fd3564e589", "version"=>"8.8.1", "snapshot"=>false}, "@version"=>"1", "type"=>"elastic_agent", "tags"=>["apache_tomcat-cache", "beats_input_raw_event"], "ecs"=>{"version"=>"8.0.0"}, "agent"=>{"name"=>"hostname", "type"=>"metricbeat", "id"=>"522874b7-bd30-487c-8c9f-a1fd3564e589", "version"=>"8.8.1", "ephemeral_id"=>"d55430eb-c910-417c-bf80-971cb2b62c25"}, "prometheus"=>{"labels"=>{"name"=>"Cache", "job"=>"prometheus", "context"=>"/manager", "host"=>"localhost", "instance"=>"localhost:9090"}, "metrics"=>{"Catalina_WebResourceRoot_maxSize"=>10240, "Catalina_WebResourceRoot_ttl"=>5000, "Catalina_WebResourceRoot_size"=>12, "Catalina_WebResourceRoot_objectMaxSize"=>512, "Catalina_WebResourceRoot_lookupCount"=>13, "Catalina_WebResourceRoot_hitCount"=>4}}, "metricset"=>{"name"=>"collector", "period"=>10000}, "data_stream"=>{"type"=>"metrics", "dataset"=>"apache_tomcat.cache", "namespace"=>"default"}, "event"=>{"dataset"=>"apache_tomcat.cache", "duration"=>151052696, "module"=>"prometheus"}, "host"=>{"name"=>"hostname", "id"=>"90ad598b369d41f68860a2898fb81488", "mac"=>["00-00-00-00-00-00"], "architecture"=>"x86_64", "hostname"=>"hostname", "os"=>{"platform"=>"ol", "name"=>"Oracle Linux Server", "kernel"=>"5.15.0-102.110.5.1.el9uek.x86_64", "type"=>"linux", "version"=>"9.2", "family"=>"redhat"}, "ip"=>["10.20.30.40"], "containerized"=>false}, "@timestamp"=>2024-04-09T15:04:25.221Z}], :response=>{"create"=>{"_index"=>".ds-metrics-apache_tomcat.cache-default-2024.03.23-000003", "_id"=>"8aSlLs8Em-fXACmJAAABjsNjjYU", "status"=>409, "error"=>{"type"=>"version_conflict_engine_exception", "reason"=>"[8aSlLs8Em-fXACmJAAABjsNjjYU][{agent.id=522874b7-bd30-487c-8c9f-a1fd3564e589, apache_tomcat.cache.application_name=/manager, host.name=hostname, service.address=http://localhost:9090/metrics}@2024-04-09T15:04:25.221Z]: version conflict, document already exists (current version [1])", "index_uuid"=>"6nhJvzKxT8axN1Z_4EzFew", "shard"=>"0", "index"=>".ds-metrics-apache_tomcat.cache-default-2024.03.23-000003"}}}}

```

But for me, these don't appear particularly related to my problem.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [April 9, 2024, 3:49pm UTC](https://discuss.elastic.co/t/logstash-high-load-and-cpu-usage/357079/4 "2024-04-09T15:49:03Z")

</div>

The screenshot suggests the JVM heap is continuously growing, and the CPU and system load are growing along with it. It looks like the heap only shrinks significantly when the JVM is restarted (leading to the brief gaps in the monitoring data).

That suggests a GC issue. I would enable GC logging (how to do that depends on the JVM, its version and the options you are using). That will show you the time spent on GC. Then get a heap dump and take a look at what is using up the heap. See [this](https://discuss.elastic.co/t/heapdumponoutofmemoryerror/298786/2) thread.

---

<div class="post-metadata">

**Author:** ![elasticforme](https://avatars.discourse-cdn.com/v4/letter/e/f05b48/32.png) [@elasticforme](https://discuss.elastic.co/u/elasticforme)\
**Post date:** [April 9, 2024, 4:13pm UTC](https://discuss.elastic.co/t/logstash-high-load-and-cpu-usage/357079/5 "2024-04-09T16:13:24Z")

</div>

what I have discover that if you have stdout in your output section and if you processing lot of documents then load will go super high.

---

<div class="post-metadata">

**Author:** ![pk92](https://avatars.discourse-cdn.com/v4/letter/p/f1d935/32.png) [@pk92](https://discuss.elastic.co/u/pk92)\
**Post date:** [April 10, 2024, 9:03am UTC](https://discuss.elastic.co/t/logstash-high-load-and-cpu-usage/357079/6 "2024-04-10T09:03:37Z")

</div>

> [@Badger](#):
>
> The screenshot suggests the JVM heap is continuously growing, and the CPU and system load are growing along with it. It looks like the heap only shrinks significantly when the JVM is restarted (leading to the brief gaps in the monitoring data).
> 
> That suggests a GC issue. I would enable GC logging (how to do that depends on the JVM, its version and the options you are using). That will show you the time spent on GC. Then get a heap dump and take a look at what is using up the heap. See [this](https://discuss.elastic.co/t/heapdumponoutofmemoryerror/298786/2) thread.

Thank you for the suggestion! I am not entirely convinced that it has something to do with the heap space though. That the heap and the load in the screenshot both drop at the time is because the Logstash service got restarted. At other times I also saw the garbage collector running and freeing up heap space and the load still remaining high.  
I will have a look into your suggestion anyway and see if I can find anything out!

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [April 10, 2024, 4:39pm UTC](https://discuss.elastic.co/t/logstash-high-load-and-cpu-usage/357079/7 "2024-04-10T16:39:27Z")

</div>

I typically run logstash with 200 MB of heap. If the heap is growing to 5 GB then I cannot think of any explanation other than a memory leak.

---

<div class="post-metadata">

**Author:** ![pk92](https://avatars.discourse-cdn.com/v4/letter/p/f1d935/32.png) [@pk92](https://discuss.elastic.co/u/pk92)\
**Post date:** [April 18, 2024, 9:02am UTC](https://discuss.elastic.co/t/logstash-high-load-and-cpu-usage/357079/8 "2024-04-18T09:02:05Z")

</div>

Like I already assumed, the problem wasn't the heap itself. It was the syslog input, like described here:

> <https://github.com/logstash-plugins/logstash-input-syslog/issues/74>
>
> \<!--
> GitHub is reserved for bug reports and feature requests; it is not the pla…ce
> for general questions. If you have a question or an unconfirmed bug , please
> visit the \[forums\](https://discuss.elastic.co/c/logstash). Please also
> check your OS is \[supported\](https://www.elastic.co/support/matrix#show\_os).
> If it is not, the issue is likely to be closed.
> 
> Logstash is located in a different organization: \[logstash\](https://github.com/elastic/logstash). For bugs specific to Logstash and not related to Logstash Plugins, please open it in the respective Logstash repository.
> 
> For security vulnerabilities please only send reports to security@elastic.co.
> See https://www.elastic.co/community/security for more information.
> 
> Please fill in the following details to help us reproduce the bug:
> \--\>
> 
> \*\*Logstash information\*\*:
> Logstash version: 8.4.3
> JVM version: bundled
> 
> \*\*Description of the problem including expected versus actual behavior\*\*:
> The plugin is not properly detecting client disconnections/resets. It seems it keeps trying to read from the closed socket, making the CPU usage grow every time a client disconnects. It may be a JRuby issue (https://github.com/jruby/jruby/issues/7961).
> 
> The problem is happening on the \[socket.each\](https://github.com/logstash-plugins/logstash-input-syslog/blob/main/lib/logstash/inputs/syslog.rb#L235) method. When the client connection is closed/reset, the code expects it to raise an \`ECONNRESET\`, stopping the loop, closing the connection and removing it from the connection counter. Instead, it doesn't raise any error, hangs and burns CPU.
> 
> An alternative solution is to change it to read using a non-blocking approach. Apparently, the \`read\_nonblock\` is not affected by this issue.
> 
> \`\`\`java
> Hot threads at 2023-09-25T10:41:31+02:00, busiestThreads=10000: 
> ================================================================================
> 61.74 % of cpu usage, state: runnable, thread name: 'input|syslog|tcp|10.1.1.12:5000}', thread id: 1726 
> app//org.jruby.util.io.PosixShim.read(PosixShim.java:158)
> app//org.jruby.util.io.OpenFile$2.run(OpenFile.java:1330)
> app//org.jruby.util.io.OpenFile$2.run(OpenFile.java:1316)
> ...
> \`\`\`
> 
> 
> \*\*Steps to reproduce\*\*:
> 1. Start Logstash with the following pipeline:
> \`\`\`
> input {
> syslog {
> port =\> 5555
> }
> }
>     
> output { stdout {} }
> \`\`\`
> 2. Run the following client code, checking the CPU usage: 
> \`\`\`ruby
> require 'socket'
>       
> HOST = 'localhost'
> PORT = 5555
>       
> def connect\_and\_close
> socket = TCPSocket.new(HOST, PORT)
> linger = \[1,0\].pack('ii')
> socket.setsockopt(Socket::SOL\_SOCKET, Socket::SO\_LINGER, linger)
> socket.close
> end
>       
> def tcp\_receiver(socket)
> socket.each { |line| puts line }
> rescue Errno::ECONNRESET
> puts "connection reset"
> end
>       
> server\_thread = Thread.new do
> server\_socket = TCPServer.new(HOST, PORT)
> loop do
> socket = server\_socket.accept
>       
> Thread.new(socket) do |socket|
> tcp\_receiver(socket)
> end
> end
> end
>       
> sleep 1
> 10.times { connect\_and\_close }
> \`\`\`

An update of the syslog input plugin to version 3.7.0 solved the problem.

---

<div class="post-metadata">

**Author:** ![diabedon](https://avatars.discourse-cdn.com/v4/letter/d/9dc877/32.png) [@diabedon](https://discuss.elastic.co/u/diabedon)\
**Post date:** [April 18, 2024, 1:42pm UTC](https://discuss.elastic.co/t/logstash-high-load-and-cpu-usage/357079/9 "2024-04-18T13:42:30Z")

</div>

I will have a look into your suggestion anyway and see if I can find anything out!
