# Logstash - using jdbc input plugin as input. And pipiline delay as 30 seconds and batch size as 1000. Still input plugin reads all data from database doesnt get impacted by pipeline configuration of size and delay

**URL:** <https://discuss.elastic.co/t/logstash-using-jdbc-input-plugin-as-input-and-pipiline-delay-as-30-seconds-and-batch-size-as-1000-still-input-plugin-reads-all-data-from-database-doesnt-get-impacted-by-pipeline-configuration-of-size-and-delay/344480>\
**Category:** Logstash\
**Created:** [October 5, 2023, 12:46pm UTC](https://discuss.elastic.co/t/logstash-using-jdbc-input-plugin-as-input-and-pipiline-delay-as-30-seconds-and-batch-size-as-1000-still-input-plugin-reads-all-data-from-database-doesnt-get-impacted-by-pipeline-configuration-of-size-and-delay/344480 "2023-10-05T12:46:40Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Lovin\_Saini](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lovin_saini/32/124749_2.png) [@Lovin\_Saini](https://discuss.elastic.co/u/Lovin_Saini)\
**Post date:** [October 5, 2023, 12:46pm UTC](https://discuss.elastic.co/t/logstash-using-jdbc-input-plugin-as-input-and-pipiline-delay-as-30-seconds-and-batch-size-as-1000-still-input-plugin-reads-all-data-from-database-doesnt-get-impacted-by-pipeline-configuration-of-size-and-delay/344480/1 "2023-10-05T12:46:40Z")

</div>

JDBC input plugin

```auto
input {
    jdbc {
          jdbc_driver_library => "/usr/share/logstash/driver/mysql-connector-java-8.0.32.jar"
          jdbc_driver_class => "com.mysql.cj.jdbc.Driver"
          jdbc_connection_string => "${MYSQL_URL}?zeroDateTimeBehavior=convertToNull"
          last_run_metadata_path => "/tmp/sql_last_value_customer.yml"
          jdbc_user => "${MYSQL_USER}"
          jdbc_password => "${MYSQL_PASS}"
          statement => "SELECT cust.id AS 'customers.id', cust.name AS 'customers.name' FROM customers cust WHERE cust.created_at > :sql_last_value ORDER BY cust.created_at limit :size offset :offset"
          use_column_value => true
          jdbc_paging_mode => "explicit"
          tracking_column => "customers.id"
          tracking_column_type => "numeric"
          record_last_run => true
          jdbc_paging_enabled => true
          jdbc_page_size => 500
          enable_metric => true
    }
}

```

Pipeline config:

```auto
pipeline.batch.size: 1000
pipeline.batch.delay: 30000

```

I expect jdbc input plugin to fetch 1000 records and wait for 30 sec then again fetch next 1000 records and wait for 30 sec goes on....

But it doesnt happen it fetch contineously all data from database doesnt wait for 30 secs.

---

<div class="post-metadata">

**Author:** ![PodarcisMuralis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/podarcismuralis/32/122606_2.png) [@PodarcisMuralis](https://discuss.elastic.co/u/PodarcisMuralis)\
**Post date:** [October 5, 2023, 7:54pm UTC](https://discuss.elastic.co/t/logstash-using-jdbc-input-plugin-as-input-and-pipiline-delay-as-30-seconds-and-batch-size-as-1000-still-input-plugin-reads-all-data-from-database-doesnt-get-impacted-by-pipeline-configuration-of-size-and-delay/344480/2 "2023-10-05T19:54:41Z")

</div>

Hi.  
I did not test it but firstly, using jdbc\_page\_size may cause this issue because it means plugin will be reading large numbers of rows from database.

I used also once this jdbc input plugin and as far as i remember i used this to control the delay.  
`schedule => "*/30 * * * *"`

---

<div class="post-metadata">

**Author:** ![Lovin\_Saini](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lovin_saini/32/124749_2.png) [@Lovin\_Saini](https://discuss.elastic.co/u/Lovin_Saini)\
**Post date:** [October 6, 2023, 5:16am UTC](https://discuss.elastic.co/t/logstash-using-jdbc-input-plugin-as-input-and-pipiline-delay-as-30-seconds-and-batch-size-as-1000-still-input-plugin-reads-all-data-from-database-doesnt-get-impacted-by-pipeline-configuration-of-size-and-delay/344480/3 "2023-10-06T05:16:05Z")

</div>

Hi,  
Seems like jdbc input plugin doesn't cater pipeline config. So even if pipeline config is there to control pipeline it can happen that this config gets ignored, all it depends on plugin used. This is my understanding.

I tried Schedule but it is usefull in case there is no pagination and the sql query is fetching latest data only if any. Because it re runs the plugin as per schedule so it doesn't go with pagination where there is a dynamic query with limit/offset.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [October 6, 2023, 6:30pm UTC](https://discuss.elastic.co/t/logstash-using-jdbc-input-plugin-as-input-and-pipiline-delay-as-30-seconds-and-batch-size-as-1000-still-input-plugin-reads-all-data-from-database-doesnt-get-impacted-by-pipeline-configuration-of-size-and-delay/344480/4 "2023-10-06T18:30:56Z")

</div>

> [@Lovin\_Saini](#):
>
> `pipeline.batch.delay: 30000`

pipeline.batch.delay tells logstash how long to wait before flushing a batch to the pipeline if the batch has fewer than pipeline.batch.size events in it. If a batch reaches pipeline.batch.size then it will be flushed immediately.

So if your jdbc input is fetching 10500 rows I would expect 10 batches to be flushed immediately and the last 500 rows to follow 30 seconds later.

---

<div class="post-metadata">

**Author:** ![Lovin\_Saini](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lovin_saini/32/124749_2.png) [@Lovin\_Saini](https://discuss.elastic.co/u/Lovin_Saini)\
**Post date:** [October 18, 2023, 1:56pm UTC](https://discuss.elastic.co/t/logstash-using-jdbc-input-plugin-as-input-and-pipiline-delay-as-30-seconds-and-batch-size-as-1000-still-input-plugin-reads-all-data-from-database-doesnt-get-impacted-by-pipeline-configuration-of-size-and-delay/344480/5 "2023-10-18T13:56:51Z")

</div>

Thanks a lot for your reply. It is happening in similar way what you are explaining.

I wanted to control pipeline pace so that it won't hit database a lot. Currently database cpu is getting in higher limits whenever i run it and i am not seeing any ways to control logstash pipeline pace. Only thing is i can reduce worker count to 1 but then also if query is complex it will again causes database cpu to get higher.

As a solution i am planning to move to cron based prepared statement where i can specify frequency probably that can help.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 15, 2023, 1:57pm UTC](https://discuss.elastic.co/t/logstash-using-jdbc-input-plugin-as-input-and-pipiline-delay-as-30-seconds-and-batch-size-as-1000-still-input-plugin-reads-all-data-from-database-doesnt-get-impacted-by-pipeline-configuration-of-size-and-delay/344480/6 "2023-11-15T13:57:41Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
