# Logstash slow processing of events

**URL:** <https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392>\
**Category:** Logstash\
**Created:** [February 1, 2023, 7:36am UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392 "2023-02-01T07:36:54Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![PRASHANT\_MEHTA](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/prashant_mehta/32/101764_2.png) [@PRASHANT\_MEHTA](https://discuss.elastic.co/u/PRASHANT_MEHTA)\
**Post date:** [February 1, 2023, 7:36am UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/1 "2023-02-01T07:36:54Z")

</div>

Hello,

I'm trying to process events from logstash and I'm facing issue of slow processing of events.There are around 100k records.In logstash.yml I've enabled log.level debug.  
So far I can observe in 2 hours around 11000 records were processed.I want the event to be processed faster.  
I'm testing it in test instance with following config:  
JVM Heap :2 gb  
cpu:4 core

LOGSTASH 7.9.1  
logstash configfile:

```auto
- pipeline.id: tcgeometrytransfer
   **queue.type: persisted**
   path.config: "/l/custom/TCS/logstash/logstash-7.9.1/scripts/tcgeometry/tc_geometry.cfg"

```

configfile:

```auto
input {
   exec {
      command => '/l/custom/TCS/logstash/logstash-7.9.1/scripts/tcsgeometry/run_tcs_geometry.sh'
      schedule => "0 57 05 * * *"
   }
}
filter {
   if [message] =~ "^\{.*\}[\s\S]*$" {
      json {
         source => "message"
         target => "parsed_json"
		 remove_field => "message"
      }
	  
      split {
         field => "[parsed_json][geoMonitorResponse]"
         target => "geometry"
         remove_field => ["parsed_json"]
      }	
      
      if [geometry][graph][monitorDate] {
	         mutate {
				convert => { "[geometry][graph][monitorDate]" => "string" }
			}
			date {
              match => ["[geometry][graph][monitorDate]", "yyyy-MM-dd'T'HH:mm:ssZ"]
              timezone => "UTC"
              target => "@timestamp"
              
			}
		}
   
     if [geometry][position][monitorDate] {
	         mutate {
				convert => { "[geometry][position][monitorDate]" => "string" }
			}
			date {
              match => ["[geometry][position][monitorDate]", "yyyy-MM-dd'T'HH:mm:ssZ"]
              timezone => "UTC"
              target => "@timestamp"
              
			}
		}
   
    if [geometry][line][monitorDate] {
	         mutate {
				convert => { "[geometry][line][monitorDate]" => "string" }
			}
			date {
              match => ["[geometry][line][monitorDate]", "yyyy-MM-dd'T'HH:mm:ssZ"]
              timezone => "UTC"
              target => "@timestamp"
              
			}
		}
      
   }
   else {
     drop { }
   }
}
output {
   elasticsearch {
      hosts => "http://abc:9200"
	  ilm_pattern => "{now/d}-000001"
      ilm_rollover_alias => "cis-monitor-geometry"
	  ilm_policy => "tcs-monitor-geometry-policy"
	  doc_as_upsert => true
	  document_id => "%{[geometry][uniqueId]}" 
   } 
}

```

Also would be helpful if someone could suggest best config setup for below attributes:  
currently using defaults,Not sure what could be best batch size and delay to be setup so log processing is fast.

pipeline.batch.size: 125  
pipeline.batch.delay: 50

Some questions:  
1)If batch size is increased to 1000 events then what should be the batch delay? , not sure how this works.Intention to increase so that events are processed faster.  
What other requirements would be needed,like does it require more core or jvm heap size to be increased?

2)For other usecase where data is around 17000 its process all data in 30 min.What is the reason that all processed data of 17000 comes in index together i.e all data is inserted at a time in index and not incremently.  
I would like to understand if batch size and delay is not given then its process with default size and delay. what might be the reason that in "index" data is not coming according to batches and if comes then come whole.Really a confusion.

3)Does batch\_size and batch delay dosent work with persisted queue?

would be interested to know how logstash could be best configured to process events faster and with best optimization.

Thanks

---

<div class="post-metadata">

**Author:** ![PRASHANT\_MEHTA](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/prashant_mehta/32/101764_2.png) [@PRASHANT\_MEHTA](https://discuss.elastic.co/u/PRASHANT_MEHTA)\
**Post date:** [February 3, 2023, 6:39am UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/2 "2023-02-03T06:39:00Z")

</div>

Hello @yaauie ,

Can you plz guide or suggest something on this,how could this be resolved for fast processing of events?

Thanks

---

<div class="post-metadata">

**Author:** ![Sunile\_Manjee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sunile_manjee/32/111461_2.png) [@Sunile\_Manjee](https://discuss.elastic.co/u/Sunile_Manjee)\
**Post date:** [February 3, 2023, 2:18pm UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/3 "2023-02-03T14:18:54Z")

</div>

Does it require more heap by increasing batch size, yes. Here are the docs on both arguments which may help you determine the values based your use case: [Tuning and Profiling Logstash Performance | Logstash Reference [8.6] | Elastic](https://www.elastic.co/guide/en/logstash/current/tuning-logstash.html)

---

<div class="post-metadata">

**Author:** ![Rios](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rios/32/95745_2.png) [@Rios](https://discuss.elastic.co/u/Rios)\
**Post date:** [February 3, 2023, 9:20pm UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/4 "2023-02-03T21:20:46Z")

</div>

Please use `http://localhost:9600/_node/stats/pipelines?pretty` on LS to show you execution time per plugin. It's useful to add IDs per plugin.

```auto
input {
   exec {
      command => '/l/custom/TCS/logstash/logstash-7.9.1/scripts/tcsgeometry/run_tcs_geometry.sh'
      schedule => "0 57 05 * * *"
      id => "exec"
   }
}

```

It might be tc\_geometry.cfg is slow, 2 hours around 11000 records is too much for LS, that should be processed inside filter/pipeline in a minute or less. Do you have some logging on .sh side? If you don't have add. Just start, end time per lines, something basic.

At the glance, the json plugin should consume the most time. Others are basic activities.  
Increasing memory is useful if you have a large amount data like XML structure with 10-50 000 nodes.

---

<div class="post-metadata">

**Author:** ![PRASHANT\_MEHTA](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/prashant_mehta/32/101764_2.png) [@PRASHANT\_MEHTA](https://discuss.elastic.co/u/PRASHANT_MEHTA)\
**Post date:** [February 6, 2023, 5:20am UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/6 "2023-02-06T05:20:27Z")

</div>

Hello @Sunile_Manjee ,

Thanx for looking into this and pointing out the links.Will go through the documents and  
check what best can be done to rectify the slow processing.

Thanx

---

<div class="post-metadata">

**Author:** ![PRASHANT\_MEHTA](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/prashant_mehta/32/101764_2.png) [@PRASHANT\_MEHTA](https://discuss.elastic.co/u/PRASHANT_MEHTA)\
**Post date:** [February 6, 2023, 5:27am UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/7 "2023-02-06T05:27:28Z")

</div>

Hello @Rios ,

Thanx for yout time to look into this and pointing out the stats plugin.I will add the id and check with logger as u mentioned.The filter plugin seems working fine,and current challenge is in pipeline I've mentioned batch size:1000 and batch delay 200.It dosent seems processing 1000 events .I'm using queue.type:persisted and when changed queue.type:memory then also it dosent pick 1000 events.  
As recommended by you will add id and check with stats plugin to check execution time.Its still pending to test.Will keep posted latest outcome

Many Thanx

---

<div class="post-metadata">

**Author:** ![Sunile\_Manjee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sunile_manjee/32/111461_2.png) [@Sunile\_Manjee](https://discuss.elastic.co/u/Sunile_Manjee)\
**Post date:** [February 6, 2023, 5:43am UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/8 "2023-02-06T05:43:48Z")

</div>

You may have looked into this, curious. Have you verified that your batch size (1000) arrives within your batch delay (per worker thread)? These settings are per pipeline worker thread. What is your pipeline.workers set to?

---

<div class="post-metadata">

**Author:** ![PRASHANT\_MEHTA](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/prashant_mehta/32/101764_2.png) [@PRASHANT\_MEHTA](https://discuss.elastic.co/u/PRASHANT_MEHTA)\
**Post date:** [February 6, 2023, 5:58am UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/9 "2023-02-06T05:58:12Z")

</div>

Hello,

Below is the pipeline config:

```auto
- pipeline.id: tcsgeometrytransfer
  queue.type: persisted
  path.config: "/l/custom/TCS/logstash/logstash7.9.1/scripts/tcsgeometry/tcs_geometry.cfg"
  pipeline.batch.size: 1000
  pipeline.batch.delay: 200

```

---

<div class="post-metadata">

**Author:** ![Sunile\_Manjee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sunile_manjee/32/111461_2.png) [@Sunile\_Manjee](https://discuss.elastic.co/u/Sunile_Manjee)\
**Post date:** [February 6, 2023, 6:15am UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/10 "2023-02-06T06:15:27Z")

</div>

pipeline\_workers is not set in your config, therefore it defaults to number of cores on your host.

---

<div class="post-metadata">

**Author:** ![PRASHANT\_MEHTA](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/prashant_mehta/32/101764_2.png) [@PRASHANT\_MEHTA](https://discuss.elastic.co/u/PRASHANT_MEHTA)\
**Post date:** [February 6, 2023, 6:19am UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/11 "2023-02-06T06:19:32Z")

</div>

Yes,I'm having 4 core cpu for this test instance, I suppose then pipeline\_workers would be taken up as 4 by default then,but still its should process those 1000 events as it takes deafult workers?

---

<div class="post-metadata">

**Author:** ![Sunile\_Manjee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sunile_manjee/32/111461_2.png) [@Sunile\_Manjee](https://discuss.elastic.co/u/Sunile_Manjee)\
**Post date:** [February 6, 2023, 1:53pm UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/12 "2023-02-06T13:53:23Z")

</div>

My speculation is that events may not be generated(source side, input) fast enough (within batch delay) to fill up 1000 events (batch size) per thread. Therefore processing (filter) what ever is available within the batch queue

---

<div class="post-metadata">

**Author:** ![Rios](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rios/32/95745_2.png) [@Rios](https://discuss.elastic.co/u/Rios)\
**Post date:** [February 6, 2023, 5:02pm UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/13 "2023-02-06T17:02:58Z")

</div>

The same opinion share with Sunile. Without logger inside run\_tcs\_geometry.sh, it's hard to say.  
In general, increasing memory will help only in case of large data. Pipeline reconfiguration is useful when you processing enormous number of messages.

@PRASHANT_MEHTA have you checked the pipeline static? Can you post JSON with customized IDs?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 6, 2023, 5:03pm UTC](https://discuss.elastic.co/t/logstash-slow-processing-of-events/324392/14 "2023-03-06T17:03:02Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
