# 1 huge pipelines vs 2 medium ones

**URL:** <https://discuss.elastic.co/t/1-huge-pipelines-vs-2-medium-ones/309925>\
**Category:** Logstash\
**Created:** [July 18, 2022, 3:28pm UTC](https://discuss.elastic.co/t/1-huge-pipelines-vs-2-medium-ones/309925 "2022-07-18T15:28:45Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![anon90868141](https://avatars.discourse-cdn.com/v4/letter/a/7ab992/32.png) [@anon90868141](https://discuss.elastic.co/u/anon90868141)\
**Post date:** [July 18, 2022, 3:28pm UTC](https://discuss.elastic.co/t/1-huge-pipelines-vs-2-medium-ones/309925/1 "2022-07-18T15:28:46Z")

</div>

Hello,

let's say I've got a cluster with 10 machines a 12 cores and I see data not being collected fast enough, so I need to raise my threads which is already at 12.  
Is better to go with the same pipeline and raise the threads to 16 or even 24 or just build another pipeline which collects the same data, so I got 2 with 12 threads each?

Thanks in advance  
Marcel

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 18, 2022, 4:17pm UTC](https://discuss.elastic.co/t/1-huge-pipelines-vs-2-medium-ones/309925/2 "2022-07-18T16:17:33Z")

</div>

This is not the kind of question that can easily be addressed in a forum like this. You need to identify what the bottleneck is and then address it. How to address it will depend on what the bottleneck is. Remember the bottleneck may not be in logstash, it could be in whatever logstash is reading from or writing to.

---

<div class="post-metadata">

**Author:** ![anon90868141](https://avatars.discourse-cdn.com/v4/letter/a/7ab992/32.png) [@anon90868141](https://discuss.elastic.co/u/anon90868141)\
**Post date:** [July 18, 2022, 7:56pm UTC](https://discuss.elastic.co/t/1-huge-pipelines-vs-2-medium-ones/309925/3 "2022-07-18T19:56:10Z")

</div>

Well, the bottleneck definitely is the Logstash as it's pulling data out of a redis db and then writing to another one. Both got enough ressources to handle the requests, so..

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [July 18, 2022, 8:22pm UTC](https://discuss.elastic.co/t/1-huge-pipelines-vs-2-medium-ones/309925/4 "2022-07-18T20:22:22Z")

</div>

What's your redis input and output looks like? What changes did you tried already? Please share your redis input and output.

---

<div class="post-metadata">

**Author:** ![anon90868141](https://avatars.discourse-cdn.com/v4/letter/a/7ab992/32.png) [@anon90868141](https://discuss.elastic.co/u/anon90868141)\
**Post date:** [July 19, 2022, 5:19am UTC](https://discuss.elastic.co/t/1-huge-pipelines-vs-2-medium-ones/309925/5 "2022-07-19T05:19:45Z")

</div>

I've tried many different scenarios, ranging from a single pipeline with 12 - 36 workers vs. 2 - 3 pipelines with 12 workers each. As I'm using 10 GB of heap, I multiplied the inflight count by 2,5 (cause 2,5x the standard heap) and then set the batchsize, so I'm barely below that threshold.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [July 19, 2022, 12:28pm UTC](https://discuss.elastic.co/t/1-huge-pipelines-vs-2-medium-ones/309925/6 "2022-07-19T12:28:41Z")

</div>

As I said in the previous post, what changes did you specifically did in the `redis` input and output?

What is your `redis` input? Did you change the number of `threads` and `batch_count` in the `redis` input? Changing those settings may help a lot with the ingestion.

Changing `pipeline.workers` and `pipeline.batch.size` will impact basically the `filter` and `output` blocks, but will make little to no difference in the `input`, which seems to be your issue.

Please share your logstash configuration with your input and your `logstash.yml` and `pipelines.yml`.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 16, 2022, 12:28pm UTC](https://discuss.elastic.co/t/1-huge-pipelines-vs-2-medium-ones/309925/7 "2022-08-16T12:28:54Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
