# Logstah performance issues

**URL:** <https://discuss.elastic.co/t/logstah-performance-issues/99128>\
**Category:** Logstash\
**Created:** [September 1, 2017, 12:52pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128 "2017-09-01T12:52:40Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Beuhlet\_Reseau](https://avatars.discourse-cdn.com/v4/letter/b/e95f7d/32.png) [@Beuhlet\_Reseau](https://discuss.elastic.co/u/Beuhlet_Reseau)\
**Post date:** [September 1, 2017, 12:52pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/1 "2017-09-01T12:52:40Z")

</div>

Hello,

I am so tired about Logstash problem.

I have total 2 850 000 messages every 15 minutes (3 document type differents), and On one of type, **I have 2 hours of delays.**

2 850 000 message every 15 minutes = **190 000 messages minutes**

=\> I have **dedicated logstash serveur (24 cpu 24 Go RAM)**, 3 elasticsearch on cluster (each have 12 cpu 16 Go RAM) and dedicated Kibana.

Logstash configuration =

```
pipeline.workers: 24
pipeline.output.workers: 24
pipeline.batch.size: 250
pipeline.batch.delay: 5

```

I will think to search a better etl it's not normal this perf (also 3 shards and 1 replica by type)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 2, 2017, 9:32pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/2 "2017-09-02T21:32:06Z")

</div>

You will probably get better performance by running multiple instances of Logstash.

This also depends on what your config is, what version of the stack you are on, your OS and even your JVM.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 3, 2017, 6:39am UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/3 "2017-09-03T06:39:32Z")

</div>

What does your Logstash configuration look like? What does your data look like? How much indexing throughput is your Elasticsearch cluster able to handle?

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [September 6, 2017, 10:39pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/4 "2017-09-06T22:39:59Z")

</div>

You are lamenting Logstash's performance without showing us how you have it configured. 190,000 messages per minute is only 3,167 messages per second. Unless they're extremely complex, and/or have heavy-duty enrichment going on, this should be achievable with that single server. There may be some constraints on I/O, but without knowing how you've configured anything, there's nothing we can do to assist you.

---

<div class="post-metadata">

**Author:** ![Beuhlet\_Reseau](https://avatars.discourse-cdn.com/v4/letter/b/e95f7d/32.png) [@Beuhlet\_Reseau](https://discuss.elastic.co/u/Beuhlet_Reseau)\
**Post date:** [September 7, 2017, 8:32am UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/5 "2017-09-07T08:32:23Z")

</div>

My logstash conf look like :

```
input {
  file {
    path => "/data/logstash/data/edr/P_ASK_*.txt"
    type => "edr"
    max_open_files => 30000
  }
}
filter {
  if [type] == "edr" {
    csv {
      columns => ["Date","CTE","CT","SubsId","MastN","MDN","MCC","ID","ModFlag","ParentalControlFlag","RuleBase","FairUseFlag","Flows","Label","Notification","TotalOctets","Grantectets","AP","QoS-Rule","Usamit","Januomer"]
    }
    if [message] =~ "\bCTE\b" {
      drop { }
    }
    mutate {
      remove_field => ["message", "host", "path"]
    }
    date {
      match => ["Date" , "UNIX"]
      remove_field => ["Date"]
    }
  }
}
output {
  if [type] == "edr" {
    elasticsearch {
      hosts => ["opm1zels01.com:9200","opm1zels02.com:9200","opm1zels03.com:9200"]
      index => "ed-%{+YYYY.MM.dd}"
    }
  }
}
```

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [September 7, 2017, 3:43pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/6 "2017-09-07T15:43:36Z")

</div>

> [@Beuhlet\_Reseau](#):
>
> ```auto
> input {
> file {
> # ...
> type => "edr"
> # ...
> filter {
> if [type] == "edr" {
> # ...
> output {
> if [type] == "edr" {
> 
> ```

If this is your _complete_ configuration, and there are no other configuration files or inputs anywhere else in your pipeline, then you don't need either of the `if [type] == "edr" {` lines. Those conditionals would be checking every line, but every line would already be of type `edr` because of the `type => "edr"` line in your `file` input block.

Those conditionals will each add a small amount of latency for your messages.

> [@Beuhlet\_Reseau](#):
>
> ```auto
> if [message] =~ "\bCTE\b" {
> drop { }
> }
> 
> ```

This regular expression is reading each _full_ line to find `"\bCTE\b"`, which much more expensive in terms of processing time than to look for the value `CTE` in an individual field. You're already breaking down the csv into individual fields in the `csv` filter. Why check the entire message if the value will only be in a given field? This could be slowing things down dramatically.

> [@Beuhlet\_Reseau](#):
>
> ```auto
> csv {
> columns => ["Date","CTE","CT","SubsId","MastN","MDN","MCC","ID","ModFlag","ParentalControlFlag","RuleBase","FairUseFlag","Flows","Label","Notification","TotalOctets","Grantectets","AP","QoS-Rule","Usamit","Januomer"]
> }
> 
> ```

If everything is truly CSV, then you could replace this with the [dissect filter](https://www.elastic.co/blog/logstash-dude-wheres-my-chainsaw-i-need-to-dissect-my-logs) and get a non-trivial performance and throughput boost.

These are a few quick observations, in no particular order of relative or expected performance gain.

---

<div class="post-metadata">

**Author:** ![Beuhlet\_Reseau](https://avatars.discourse-cdn.com/v4/letter/b/e95f7d/32.png) [@Beuhlet\_Reseau](https://discuss.elastic.co/u/Beuhlet_Reseau)\
**Post date:** [September 8, 2017, 9:38am UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/7 "2017-09-08T09:38:18Z")

</div>

> [@theuntergeek](#):
>
> If this is your complete configuration, and there are no other configuration files or inputs anywhere else in your pipeline, then you don’t need either of the if [type] == "edr" { lines. Those conditionals would be checking every line, but every line would already be of type edr because of the type =\> "edr" line in your file input block.

Top thank you.

> [@theuntergeek](#):
>
> if [message] =~ "\bCTE\b" {  
> drop { }  
> }

in fact, I have a header on each file. I do not see how to do what you're talking about  
So, maybe like that :

```
if [CTE] = "CTENAME" { drop {} }
csv {
      columns => ["Date","CTE","CT","SubsId","MastN","MDN","MCC","ID","ModFlag","ParentalControlFlag","RuleBase","FairUseFlag","Flows","Label","Notification","TotalOctets","Grantectets","AP","QoS-Rule","Usamit","Januomer"]
    }

```

?

> [@theuntergeek](#):
>
> If everything is truly CSV, then you could replace this with the dissect filter and get a non-trivial performance and throughput boost.

I don't know how to make that ☹ and if it's installed (ELK 5.5)

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [September 8, 2017, 1:15pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/8 "2017-09-08T13:15:12Z")

</div>

> [@Beuhlet\_Reseau](#):
>
> I have a header on each file. I do not see how to do what you're talking about  
> So, maybe like that :
> 
> ```auto
> if [CTE] = "CTENAME" { drop {} }
> 
> ```

Because it's a conditional, it should be:

```auto
if [CTE] == "CTENAME" { drop {} }

```

> [@Beuhlet\_Reseau](#):
>
> I don't know how to make that ☹ and if it's installed (ELK 5.5)

5.5 should have the dissect filter installed by default. The linked blog post in my earlier answer shows how to start configuring the dissect filter.

---

<div class="post-metadata">

**Author:** ![Beuhlet\_Reseau](https://avatars.discourse-cdn.com/v4/letter/b/e95f7d/32.png) [@Beuhlet\_Reseau](https://discuss.elastic.co/u/Beuhlet_Reseau)\
**Post date:** [September 8, 2017, 1:28pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/9 "2017-09-08T13:28:24Z")

</div>

Ok thank you @theuntergeek  
My config look like :

```
input {
  file {
    path => "/data/logstash/data/edr/P_ASK_*.txt"
    max_open_files => 30000
  }
}
filter {
    csv {
      columns => ["Date","CTE","CT","SubsId","MastN","MDN","MCC","ID","ModFlag","Parental","RuleBase","Faig","Flows","Label","Notification","TotalOctets","Grantectets","AP","QoS-Rule","Usamit","Januomer"]
    }

    if [CTE] == "CTENAME" { drop {} }

    mutate {
      remove_field => ["message", "host", "path"]
    }
    date {
      match => ["Date" , "UNIX"]
      remove_field => ["Date"]
    }
}
output {
    elasticsearch {
      hosts => ["opm1zels01.com:9200","opm1zels02.com:9200","opm1zels03.com:9200"]
      index => "ed-%{+YYYY.MM.dd}"
    }
}

```

Last thing, i must put "max\_open\_files =\> 30000 " in input because without, i have somme error message about max open file reached (even so ulimit = infiny ...)

---

<div class="post-metadata">

**Author:** ![Beuhlet\_Reseau](https://avatars.discourse-cdn.com/v4/letter/b/e95f7d/32.png) [@Beuhlet\_Reseau](https://discuss.elastic.co/u/Beuhlet_Reseau)\
**Post date:** [September 8, 2017, 2:10pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/10 "2017-09-08T14:10:53Z")

</div>

So, that good like that @theuntergeek :

```
dissect {
      mapping => {
        "message" => "%{Date},%{CTE},%{CT},%{SubsId},%{MastDN},%{MDN},%{MCC},%{ID},%{ModFlag},%{ParentalControl},%{RBe},%{FFlag},%{Status},%{RedLabel},%{NotificatD},%{UsedT},%{TotalOctets},%{AP},%{List},%{Usmit},%{Janomer}"
      }
    }

```

What is :

%{priority} or %{?priority}

or the difference between %{CTE} and %{+CTE} (i think it's to add multiple field into one no ?)

Is a problem if i define date field (for example : CTE1,dd/MM/yyyy HH:mm:ss,CT12 ) like that : %{CTE},%{Date},%{CT} ?

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [September 8, 2017, 3:00pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/11 "2017-09-08T15:00:57Z")

</div>

With regards to the dissect filter, since your delimiter is always a comma, you shouldn't have to worry about using the `?` or `+` in your field names.

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [September 8, 2017, 3:02pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/12 "2017-09-08T15:02:43Z")

</div>

> [@Beuhlet\_Reseau](#):
>
> Last thing, i must put "max\_open\_files =\> 30000 " in input because without, i have somme error message about max open file reached (even so ulimit = infiny ...)

You may want to ingest fewer files at once. I understand that you have many files you want to read in, but this message indicates you may be taxing Logstash by trying to open too many files at once. Try limiting the scope of your glob/wildcard and see if that helps.

---

<div class="post-metadata">

**Author:** ![Beuhlet\_Reseau](https://avatars.discourse-cdn.com/v4/letter/b/e95f7d/32.png) [@Beuhlet\_Reseau](https://discuss.elastic.co/u/Beuhlet_Reseau)\
**Post date:** [September 8, 2017, 3:14pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/13 "2017-09-08T15:14:51Z")

</div>

> [@theuntergeek](#):
>
> You may want to ingest fewer files at once. I understand that you have many files you want to read in, but this message indicates you may be taxing Logstash by trying to open too many files at once. Try limiting the scope of your glob/wildcard and see if that helps.

I see, Dissect filter incorporates a field suppression section, it has better than a traditional mutate

So i can remove

```
mutate {
      remove_field => ["message", "host", "path"]
    }

```

TO

```
dissect {
      mapping => {
        "message" => "%{xxxxx}...."
      }
    remove_field => ["message", "host", "path"]
}

```

Right ? 44% of efficacity it's amazing

```
Events per second
Dissect 16396	
Mutate (rename)	29248

```

EDIT : With DISSECT i have some error message type :

```
 Dissector mapping, key not found in event
 Dissector mapping, key not found in event
 Dissector mapping, key not found in event
 Dissector mapping, key not found in event
 Dissector mapping, key not found in event

```

I know what, i have one conf file by type of data (3 currently):

1. logstash-ed.conf
2. logstash-ca.conf
3. logstash-ms.conf

By removing the "if type == xx", Lostash tries to apply Dissect filter to each conf I believe. I have not yet changed the others

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [September 8, 2017, 3:41pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/14 "2017-09-08T15:41:00Z")

</div>

> [@Beuhlet\_Reseau](#):
>
> By removing the "if type == xx", Lostash tries to apply Dissect filter to each conf I believe. I have not yet changed the others

I did warn about that in the beginning. You have such a strong server you might want to consider a separate pipeline for each configuration file. In 5.x, that means a separate instance of Logstash for each (not necessarily multiple installs, just one with 3 different configurations). In 6.0, you'll be able to define multiple pipelines within one instance.

Or you could go back to the conditionals the way you had it before.

---

<div class="post-metadata">

**Author:** ![Beuhlet\_Reseau](https://avatars.discourse-cdn.com/v4/letter/b/e95f7d/32.png) [@Beuhlet\_Reseau](https://discuss.elastic.co/u/Beuhlet_Reseau)\
**Post date:** [September 8, 2017, 3:44pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/15 "2017-09-08T15:44:17Z")

</div>

> [@theuntergeek](#):
>
> I did warn about that in the beginning. You have such a strong server you might want to consider a separate pipeline for each configuration file. In 5.x, that means a separate instance of Logstash for each (not necessarily multiple installs, just one with 3 different configurations). In 6.0, you'll be able to define multiple pipelines within one instance.
> 
> Or you could go back to the conditionals the way you had it before.

Ok @theuntergeek ! but it's bad if i have 3 conf files instead of a single containing the 3 with conditionnal if?  
What do you think about that ?

EDIT : we agree that there is no need to define an conditionnal if type in the output section, I send everything in elasticsearch In any case

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [September 8, 2017, 4:03pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/16 "2017-09-08T16:03:32Z")

</div>

> [@Beuhlet\_Reseau](#):
>
> we agree that there is no need to define an conditionnal if type in the output section, I send everything in elasticsearch In any case

If they are not bounded by conditionals, then you will export each line to elasticsearch 3 times.

---

<div class="post-metadata">

**Author:** ![Beuhlet\_Reseau](https://avatars.discourse-cdn.com/v4/letter/b/e95f7d/32.png) [@Beuhlet\_Reseau](https://discuss.elastic.co/u/Beuhlet_Reseau)\
**Post date:** [September 8, 2017, 4:04pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/17 "2017-09-08T16:04:59Z")

</div>

Ok @theuntergeek Thank you vey much for your help ! You are a good man and you are my sensei now.

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [September 8, 2017, 4:05pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/18 "2017-09-08T16:05:13Z")

</div>

> [@Beuhlet\_Reseau](#):
>
> but it's bad if i have 3 conf files instead of a single containing the 3 with conditionnal if?
> 
> What do you think about that ?

Logstash merges the files into one when it reads them in. You still would need conditionals or separate instances of Logstash.

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [September 8, 2017, 5:08pm UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/19 "2017-09-08T17:08:47Z")

</div>

@Beuhlet_Reseau  
One thing to note about using Dissect instead of CSV is:  
Dissect does not check for a `,` comma in a quoted section like the CSV filter does.  
e.g. a message line like this

```auto
Adam Andrews, Beth Bell, "Cliff, Clive", Dave Dent 

```

and a dissect like this:

```auto
%{name_1}, %{name_2}, %{name_3}, %{name_4}, %{others}

```

will give (not what is expected)

```auto
name_1: Adam Andrews
name_2: Beth Bell
name_3: "Cliff
name_4: Clive"
others: Dave Dent

```

PROTIP 1: Always include a `others` or `rest` field at the end. Then check that this field is always empty - if its not, then your data has changed in some way. Output to a file or send an email or put it in redis.

PROTIP 2: use a named skip field if you know you don't need that data. e.g. `%{?host}`

---

<div class="post-metadata">

**Author:** ![Beuhlet\_Reseau](https://avatars.discourse-cdn.com/v4/letter/b/e95f7d/32.png) [@Beuhlet\_Reseau](https://discuss.elastic.co/u/Beuhlet_Reseau)\
**Post date:** [September 11, 2017, 9:22am UTC](https://discuss.elastic.co/t/logstah-performance-issues/99128/20 "2017-09-11T09:22:32Z")

</div>

Hello,

> [@guyboertje](#):
>
> PROTIP 1: Always include a others or rest field at the end. Then check that this field is always empty - if its not, then your data has changed in some way. Output to a file or send an email or put it in redis.

Excuse me but, i don't understand 😊 ., What's the aim ? Sometimes my fields are empty it's a problem ?

> [@guyboertje](#):
>
> PROTIP 2: use a named skip field if you know you don't need that data. e.g. %{?host}

Ok, it's better to use `%{?onething}` or use

```
dissect {
      mapping => {
        "message" => "...|%{onething}" ]
      }
     remove_field => ["%{onething}"]
}

```

[Next page](https://discuss.elastic.co/t/logstah-performance-issues/99128.md?page=2)
