Kafka Partitions and Consumer Groups
Problem Explain Kafka partitions and how consumers interact with them.
Be ready to discuss
- A topic is split into partitions, each an ordered, append-only, immutable log; partitioning is the unit of parallelism for both producing and consuming.
- Within a consumer group, each partition is assigned to exactly one consumer, so a group's useful parallelism is capped by partition count - extra consumers sit idle.
- Ordering is guaranteed only within a partition, never across a topic; events needing relative order (same user, same order ID) must share a partition key.
- How the partition is chosen: hash of the key modulo partition count, or round-robin when the key is null - and why adding partitions later breaks existing key-to-partition mapping.
- Offsets are tracked per partition per consumer group, so independent groups read the same topic at their own pace without interfering.
- Rebalancing: when consumers join, leave, or miss a heartbeat, partitions are reassigned - and the classic stop-the-world pause it causes.
- Rebalance storms: slow consumers exceeding
max.poll.interval.msget evicted, triggering a rebalance that slows things further; cooperative/incremental rebalancing mitigates this. - Replication and durability: leader/follower replicas per partition, in-sync replicas (ISR), and
ackssettings. - Delivery semantics: at-least-once vs exactly-once, and why offset commit timing determines which you get.
- Choosing partition count: too few caps throughput, too many inflate rebalance time, metadata, and open file handles.
asked …