# Scaling logstash nodes

**URL:** <https://discuss.elastic.co/t/scaling-logstash-nodes/259792>\
**Category:** Logstash\
**Created:** [December 29, 2020, 10:13am UTC](https://discuss.elastic.co/t/scaling-logstash-nodes/259792 "2020-12-29T10:13:58Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![David\_Beradze](https://avatars.discourse-cdn.com/v4/letter/d/e480ec/32.png) [@David\_Beradze](https://discuss.elastic.co/u/David_Beradze)\
**Post date:** [December 29, 2020, 10:13am UTC](https://discuss.elastic.co/t/scaling-logstash-nodes/259792/1 "2020-12-29T10:13:58Z")

</div>

Hello,

We have oracle (single table) \>\> logstash \>\> elasticsearch. What is the way to scale horizontally logstash nodes to prevent the same data selection from the same source? (oracle table)

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [December 29, 2020, 11:19am UTC](https://discuss.elastic.co/t/scaling-logstash-nodes/259792/2 "2020-12-29T11:19:09Z")

</div>

I don't think Logstash scales well on this case.

I'm assuming that you are using the `jdbc` input to query your oracle database, this plugin needs to store metadata about the last time it ran, this is stored on a file set by the configuration option `last_run_metadata_path`.

The `last_run_metadata_path` default value is a file inside the Logstash instance home folder.

What you could try is to set this value to a network shared folder and mount that folder in all the machines where you will run logstash, but this would not be enough, because you also have the `schedule` option, which can not be the same because you could have two or more instances trying to write into the same file at the same time.

You would also need to use different `schedule` options for each one of your instances and make sure that they do not overlap.

For example, you have one instance making a query every minutes and the other one making a query every two minutes, and your query would also need to take less the one minutes to run.

Maybe this could work, but you would need to test it with different `schedule` combinations.

---

<div class="post-metadata">

**Author:** ![David\_Beradze](https://avatars.discourse-cdn.com/v4/letter/d/e480ec/32.png) [@David\_Beradze](https://discuss.elastic.co/u/David_Beradze)\
**Post date:** [December 29, 2020, 1:12pm UTC](https://discuss.elastic.co/t/scaling-logstash-nodes/259792/3 "2020-12-29T13:12:23Z")

</div>

Thank you for your answer. I have to pull the data every 1 second so I can not set different schedules.  
What about linux Keepalived to keep both node (it work only for two node) sync, as soon as one node goes down I can update pipeline. What do you think will it work?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [December 29, 2020, 1:48pm UTC](https://discuss.elastic.co/t/scaling-logstash-nodes/259792/4 "2020-12-29T13:48:39Z")

</div>

Keepalived would make sense if you were sending data to logstash, but your case is different, it is the logstash process that is making the query to your database, and it needs to keep track of the last queried data, so you can't have two logstash nodes making the same query at the same time.

Also, for what I remember, the `schedule` option in `jdbc` input uses the `cron` format and the lower time you get is to run it every minute, I don't think you would be able to run it every second.

One solution would be putting a Kafka cluster between your database and your logstash, you would need to use a jdbc connector to put your database data into the kafka cluster and then you could use as many logstash nodes you want, all of them consuming from kafka.

When you use the same `group_id` in your logstash input, Kafka will track the already consumed messages between the consumer group.

But this would also add another layer in your infrastructure.

---

<div class="post-metadata">

**Author:** ![David\_Beradze](https://avatars.discourse-cdn.com/v4/letter/d/e480ec/32.png) [@David\_Beradze](https://discuss.elastic.co/u/David_Beradze)\
**Post date:** [December 30, 2020, 5:03pm UTC](https://discuss.elastic.co/t/scaling-logstash-nodes/259792/5 "2020-12-30T17:03:03Z")

</div>

Is there any option to sync oracle with ES with scaling feature? I need to sync for every one second and the data is more than 1000 insert per second.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [December 30, 2020, 5:29pm UTC](https://discuss.elastic.co/t/scaling-logstash-nodes/259792/6 "2020-12-30T17:29:07Z")

</div>

> [@David\_Beradze](#):
>
> Is there any option to sync oracle with ES with scaling feature?

Not in a native way, you will need to implement some connector between your database and logstash to fit your requirements.

The `jdbc` input plugin has a resolution down to the minute only, and to scale logstash you need to use other tools or services like HAProxy, Keepalived, Kafka, Redis etc.

You can for example write a python script to query your oracle database and send the data to Logstash using one of the available inputs, like `tcp`, `udp` or `http` or you can send it to a Kafka cluster and configure logstash to use the `kafka` input.

You can also write an API to query your oracle database and use the `http_poller` input to query this API, I think this way you can configure an schedule of `every 1s`.

If you want to send it directly to Elasticsearch you will can just skip the logstash part when you query your data, just need to sent it in a format that Elasticsearch will understand.

But either way, it is a logic that you will need to implement to fit your requirements.

---

<div class="post-metadata">

**Author:** ![David\_Beradze](https://avatars.discourse-cdn.com/v4/letter/d/e480ec/32.png) [@David\_Beradze](https://discuss.elastic.co/u/David_Beradze)\
**Post date:** [December 30, 2020, 7:43pm UTC](https://discuss.elastic.co/t/scaling-logstash-nodes/259792/7 "2020-12-30T19:43:11Z")

</div>

Jdbc input plugin supports pull the data every seconds, it works.

For single node, will multiple worker work to pull the data from jdbc? Will they use the same last\_run\_metadata\_path?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 27, 2021, 7:43pm UTC](https://discuss.elastic.co/t/scaling-logstash-nodes/259792/8 "2021-01-27T19:43:12Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
