Kafka Consumer Throughput and Lag Handling
Problem How can Kafka consumer throughput be increased, and what happens when consumption lags behind production?
Be ready to discuss
- The parallelism ceiling: within a consumer group a partition is consumed by at most one consumer, so partition count caps consumer parallelism — adding instances beyond the partition count leaves them idle.
- Fetch tuning:
max.poll.records,fetch.min.bytes,fetch.max.wait.msto batch more per poll, balanced againstmax.poll.interval.ms— process too much per poll and the consumer is evicted from the group and rebalances. - Processing model: keep per-record work fast, offload heavy or blocking work to a worker pool, and understand what that does to offset-commit correctness and ordering guarantees.
- Increasing partitions: raises the parallelism ceiling but is one-way, reshuffles key-to-partition mapping (breaking per-key ordering for existing keys), and forces a rebalance.
- Consumer lag: the offset gap between producer and consumer; growing lag means delayed processing, and if retention expires before the lagging consumer catches up, records are lost unread.
- Mitigations: alert on lag as a first-class metric, scale consumers out, apply backpressure upstream, and in the worst case seek to a later offset or to the end — explicitly trading data loss for recovery.
- Root-cause framing: distinguish a sustained throughput deficit from a transient spike, and check for a hot partition caused by a skewed key before blaming consumer count.
asked …