# Logstash Kafka consumer count

**URL:** <https://discuss.elastic.co/t/logstash-kafka-consumer-count/325671>\
**Category:** Logstash\
**Created:** [February 16, 2023, 2:00am UTC](https://discuss.elastic.co/t/logstash-kafka-consumer-count/325671 "2023-02-16T02:00:36Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![rsk0](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rsk0/32/124810_2.png) [@rsk0](https://discuss.elastic.co/u/rsk0)\
**Post date:** [February 16, 2023, 2:00am UTC](https://discuss.elastic.co/t/logstash-kafka-consumer-count/325671/1 "2023-02-16T02:00:36Z")

</div>

According to the [Logstash guide](https://www.elastic.co/guide/en/logstash/current/tips.html#tip-kafka):

> "How many partitions should I use per topic?"
> 
> At least the number of Logstash nodes multiplied by consumer threads per node.
> 
> **Better yet, use a multiple of the above number**. Increasing the number of partitions for an existing topic is extremely complicated. Partitions have a very low overhead. Using 5 to 10 times the number of partitions suggested by the first point is generally fine, so long as the overall partition count does not exceed 2000.

(Emphasis mine.)

Why is it better to have more partitions per consumer, rather than just 1-to-1?

Is there an efficiency benefit? A redundancy benefit?

And are there costs, too, like could there be time-spent-iterating-over-consumers growing with consumer count, or maybe contention at some point, or maybe it costs in Kafka memory overhead to manage lots of consumers ... or?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [February 16, 2023, 5:47am UTC](https://discuss.elastic.co/t/logstash-kafka-consumer-count/325671/2 "2023-02-16T05:47:58Z")

</div>

The guide makes no sense to me as-is. I suspect that original-brownbear meant "extremely cheap" rather than "extremely complicated". karenzone just copied the FAQ that original-brownbear wrote to summarize questions he was repeatedly being asked.

If your threads and partitions are numbered in 2 digits I would not worry about it. If you have thousands then benchmark it.

A larger number of partitions will result in a smoother distribution of partitions across threads.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 16, 2023, 5:48am UTC](https://discuss.elastic.co/t/logstash-kafka-consumer-count/325671/3 "2023-03-16T05:48:40Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
