# Logstash tcp input CPU perfomance

**URL:** <https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759>\
**Category:** Logstash\
**Created:** [June 21, 2022, 12:13pm UTC](https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759 "2022-06-21T12:13:28Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![perezdev](https://avatars.discourse-cdn.com/v4/letter/p/8e8cbc/32.png) [@perezdev](https://discuss.elastic.co/u/perezdev)\
**Post date:** [June 21, 2022, 12:13pm UTC](https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759/1 "2022-06-21T12:13:28Z")

</div>

Hello,

I have configured logstash using one tcp input as follows:

```auto
input {
  #logs01
  tcp {
    type => "logs"
    codec => "line"
    port => 9916
    add_field => { "event_dataset" => "logs_01" }
  }
}

```

With this configuration, we are expecting nearly 100 EPS from this source and everything is working as expected.

The thing is that we are expecting more logs from the same source on different port, so after adding another tcp input in the configuration file, logstash starts to increase the CPU utilization to more than 100% and we start to face a gap of time between the event indexed time and the event received timestamp in logstash.

```auto
input {
  #logs01
  tcp {
    type => "logs"
    codec => "line"
    port => 9916
    add_field => { "event_dataset" => "logs_01" }
  }
  #logs02
  tcp {
    type => "logs"
    codec => "line"
    port => 9917
    add_field => { "event_dataset" => "logs_02" }
  }
}

```

What could be the issue here? It seems something related to adding more than one tcp input in logstash. Could you help me with this issue?

Thank you.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 21, 2022, 1:11pm UTC](https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759/2 "2022-06-21T13:11:43Z")

</div>

You would need to share the full pipeline to give more insight into your issue.

What do your filters look like? What does your messages looks like? Do the messages from `logs_01` and `logs_02` have the same format or they are different?

---

<div class="post-metadata">

**Author:** ![perezdev](https://avatars.discourse-cdn.com/v4/letter/p/8e8cbc/32.png) [@perezdev](https://discuss.elastic.co/u/perezdev)\
**Post date:** [June 21, 2022, 1:27pm UTC](https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759/3 "2022-06-21T13:27:47Z")

</div>

I have the following configuration for logs\_01 (the same config will be for logs\_02, just changing the grok pattern and the input using port 9997)

```auto
input {
  #logs01
  tcp {
    type => "logs"
    codec => "line"
    port => 9916
    add_field => { "event_dataset" => "logs_01" }
  }
}
filter {
  #logs01 filtering
  grok {
	 patterns_dir => ["/etc/logstash/patterns"]
	 match => { "message" => "GROK PATTERN" }
  }
  if "_grokparsefailure" in [tags] {
	drop { }
  }
}
output {
  stdout { codec => rubydebug }
  http {
	  id => "logs_01"
	  headers => {
	   "x-api-key" => "APIKEY"
	  }
	  http_method => "post"
	  url => "API ENDPOINT"
	  proxy => "PROXY"
	  automatic_retries => 10
	  socket_timeout => 60
	  validate_after_inactivity => 3
	  request_timeout => 10
  }
}

```

I have one pipeline for each conf file, with the following parameters:

#logs\_01 have approx. 100 EPS

- pipeline.id: logs\_01  
path.config: /etc/logstash/conf.d/logs\_01.conf  
queue.type: memory  
pipeline.workers: 10  
pipeline.batch.size: 800  
pipeline.batch.delay: 1

#logs\_02 have approx. 8 EPS

- pipeline.id: logs\_02  
path.config: /etc/logstash/conf.d/logs\_02.conf  
queue.type: memory  
pipeline.workers: 2  
pipeline.batch.size: 1000  
pipeline.batch.delay: 1

Running in a logstash server with following requirements:

- CPU: 2
- RAM: 4GB

Hope you can help me with this issue.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 21, 2022, 1:50pm UTC](https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759/4 "2022-06-21T13:50:36Z")

</div>

Oh I see, the first configuration you shared is not the one you are using, if you are running `pipelines.yml` and have the two pipelines in different files, you do not have an `input` block with the two `tcp` inputs as you shared, so what I was thinking that could be the issue it is not true anymore, for example messages passing through grok filters that won't match.

With `pipelines.yml` your two pipelines are completely separated from each other.

But your pipeline configurations seems a little weird for the specs of your server, you have a server with 2 CPU and 4 GB, but you are using 10 workers for a pipeline with a batch size of 800, the recommendation for this setting is to change them from the default value if see that your server is not using all the resources, which is not your case now.

I would remove the `pipeline.workers` and `pipeline.batch.size` and also `pipeline.batch.delay` from your pipelines configurations and see how it performs, you have a low rate of events and the default values are more then enough for it, specially the `pipeline.batch.delay` which is in miliseconds and [rarely needs](https://www.elastic.co/guide/en/logstash/current/tuning-logstash.html#tuning-logstash) to be changed.

---

<div class="post-metadata">

**Author:** ![perezdev](https://avatars.discourse-cdn.com/v4/letter/p/8e8cbc/32.png) [@perezdev](https://discuss.elastic.co/u/perezdev)\
**Post date:** [June 21, 2022, 2:19pm UTC](https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759/5 "2022-06-21T14:19:40Z")

</div>

Thank you for your prompt reply.

I have just tried to remove `pipeline.workers`,`pipeline.batch.size` and `pipeline.batch.delay` as suggested, but the CPU utilization is still over 100%. Furthermore, with default parameters for workers we noticed that indexing rate decrease a lot for `logs_01`:

![image](https://us1.discourse-cdn.com/elastic/original/3X/d/7/d7bd2bb1df5e5cde392917c543e1207043483bde.png)

If I delete the pipeline for `logs_02`, the CPU utilization goes under 30%, but indexing rate doesn't go back to normal as shown in the screenshot.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 21, 2022, 2:57pm UTC](https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759/6 "2022-06-21T14:57:48Z")

</div>

If you add the `logs_02` pipeline and leave both pipelines running with the default configurations, the CPU will also increase to 100% as before?

Since without the `logs_02` pipeline your CPU utilization is pretty low, you could then start changing the parameters for the `logs_01` pipeline, I would start doubling the workers and batch size from its default value until you find the optimal value.

But you need first to find what is the issue with the `logs_02` pipeline, can you share some example messages for this pipeline and also the grok patterns that you are using?

---

<div class="post-metadata">

**Author:** ![perezdev](https://avatars.discourse-cdn.com/v4/letter/p/8e8cbc/32.png) [@perezdev](https://discuss.elastic.co/u/perezdev)\
**Post date:** [June 21, 2022, 3:28pm UTC](https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759/7 "2022-06-21T15:28:05Z")

</div>

That's right. I don't have CPU issues when running just `log_01` pipeline and after changing workers, indexing rate comes back to normal and CPU is still okay.

But after adding `logs_02` pipeline with default workers and parameters, CPU increase over 100% and indexing rate in other pipeline decrease as mentioned.

This is the log format for `logs_02`: (fields between double quotes, separated by comma)

`"Tue Jun 21 14:03:23 2022","user@mail.com","Data","Location","80","50","443","45","X.X.X.X","Y.Y.Y.Y","Z.Z.Z.Z","A.A.A.A","B.B.B.B","0","Data","Drop","No","Yes","No","HTTPS","app","TCP","SSL","Country","115","POLICY","197","683","0","115","1","None","None","None","user","hostuser"`

This is the grok pattern (similar to the one used for `log_01` as they are events from the same techonology but with some differences in the log format):

`match => { "message" => "\"%{TIMESTAMP_01:eventcreated}\",\"%{DATA:useremail}\",\"%{DATA:organizationdep}\",\"%{DATA:organizationloc}\",\"%{NUMBER:cdestinationport}\",\"%{NUMBER:courceport}\",\"%{NUMBER:sdestinationport}\",\"%{NUMBER:ssourceport}\",\"%{IP:csourceip}\",\"%{IP:cdestinationip}\",\"%{IP:ssourceip}\",\"%{IP:sdestinationip}\",\"%{IP:sourceip}\",\"%{NUMBER:source_port}\",\"%{DATA:eventtype}\",\"%{DATA:eventaction}\",\"%{GREEDYDATA}\",\"%{GREEDYDATA}\",\"%{GREEDYDATA}\",\"%{DATA:networktype}\",\"%{DATA:networkapplication}\",\"%{DATA:networkprotocol}\",\"%{DATA:eventcategory}\",\"%{DATA:countryname}\",\"%{NUMBER:eventduration}\",\"%{DATA:rulename}\",\"%{NUMBER:clientbytes}\",\"%{NUMBER:destinationbytes}\",\"%{GREEDYDATA}\",\"%{GREEDYDATA}\",\"%{GREEDYDATA}\",\"%{DATA:ruleset}\",\"%{DATA:indicatordescription}\",\"%{DATA:indicatorname}\",\"%{DATA:username}\",\"%{DATA:hostname}\"" }`

using pattern: `TIMESTAMP_01 %{DAY} %{MONTH} %{MONTHDAY} %{HOUR}:%{MINUTE}:%{SECOND} %{YEAR}`

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 21, 2022, 4:24pm UTC](https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759/8 "2022-06-21T16:24:07Z")

</div>

Your grok pattern is very expensive with a lot of `GREEDYDATA` and `DATA` patterns, it is also not anchored with a `^`, which helps improve the grok performance.

Normally you should anchor your groks and avoid using many `GREEDYDATA` , which is expensive.

But in your case you do not even need to use grok if your messages from `log_02` are like this, you have a `CSV` message, so it is better to use the [csv filter](https://www.elastic.co/guide/en/logstash/current/plugins-filters-csv.html) to parse it.

The main difference is that it won't validate the value of the field like grok, but if your messages have all the same format and the values do not change, for example, the column where you have an IP address will always have an IP address, then there is no need to use grok.

The following filter will parse your message, it will but to unnamed fields where you are using `GREEDYDATA` in a field named `not_used` and it will remove this field if the `csv` filter is successful.

```auto
filter {
    csv {
        source => "message"
        separator => ","
        skip_empty_columns => true
        columns => [
            "eventcreated", "useremail", "organizationdep", "organizationloc", "cdestinationport", "courceport", "sdestinationport", "ssourceport", "csourceip", "cdestinationip", "ssourceip", "sdestinationip", "sourceip", "source_port", "eventtype", "eventaction", "not_used", "not_used", "not_used", 
            "networktype", "networkapplication", "networkprotocol", "eventcategory", "countryname", "eventduration", "rulename", "clientbytes", 
            "destinationbytes", "not_used", "not_used", "not_used", "ruleset", "indicatordescription", "indicatorname", "username", "hostname"
        ]
        remove_field => ["not_used"]
    }
}

```

Besides that, if you still want to use `grok`, you should anchor your pattern using a `^` and also try to use other thing in the place of `GREEDYDATA` or `DATA`, maybe a `NOTSPACE` helps, and I think it is less expensive than `GREEDYDATA`.

---

<div class="post-metadata">

**Author:** ![perezdev](https://avatars.discourse-cdn.com/v4/letter/p/8e8cbc/32.png) [@perezdev](https://discuss.elastic.co/u/perezdev)\
**Post date:** [June 22, 2022, 11:02am UTC](https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759/9 "2022-06-22T11:02:52Z")

</div>

Definitely, that was the issue. After using csv filter as suggested for `logs_02` my CPU went back to normal and indexing rate is working fine as well.

Thank you so much for the help!!

I'd like to adapt as well `logs_01` to use csv filter, but some fields have double quotes in the values, for example:

"Tue Jun 21 14:03:23 [2022","user@mail.com](mailto:2022%22,%22user@mail.com)","Data","Location","80","50","443","45",**"domainexample&events=[["pageview"%2c{}]]"**,"Y.Y.Y.Y","Z.Z.Z.Z","A.A.A.A","B.B.B.B","0","Data","Drop","No","Yes","No","HTTPS","app","TCP","SSL","Country","115","POLICY","197","683","0","115","1","None","None","None","user","hostuser"

So, I'm facing this error:  
`:exception=>#<CSV::MalformedCSVError: Missing or stray quote in line 1`

Is there a way to use csv filters with this type of fields value?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 22, 2022, 1:49pm UTC](https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759/10 "2022-06-22T13:49:26Z")

</div>

It is always the same field, in the same position in the csv or other fields in other positions could also have double quotes?

If it always the same position in the csv message, then maybe you could do a workaround using dissect to split your message and multiple parts.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 20, 2022, 1:49pm UTC](https://discuss.elastic.co/t/logstash-tcp-input-cpu-perfomance/307759/11 "2022-07-20T13:49:55Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
