ZZomato·Tech KnowledgeL3System Design

Kafka Consumer Throughput and Lag Handling

Problem How can Kafka consumer throughput be increased, and what happens when consumption lags behind production?

Be ready to discuss

  • The parallelism ceiling: within a consumer group a partition is consumed by at most one consumer, so partition count caps consumer parallelism — adding instances beyond the partition count leaves them idle.
  • Fetch tuning: max.poll.records, fetch.min.bytes, fetch.max.wait.ms to batch more per poll, balanced against max.poll.interval.ms — process too much per poll and the consumer is evicted from the group and rebalances.
  • Processing model: keep per-record work fast, offload heavy or blocking work to a worker pool, and understand what that does to offset-commit correctness and ordering guarantees.
  • Increasing partitions: raises the parallelism ceiling but is one-way, reshuffles key-to-partition mapping (breaking per-key ordering for existing keys), and forces a rebalance.
  • Consumer lag: the offset gap between producer and consumer; growing lag means delayed processing, and if retention expires before the lagging consumer catches up, records are lost unread.
  • Mitigations: alert on lag as a first-class metric, scale consumers out, apply backpressure upstream, and in the worst case seek to a later offset or to the end — explicitly trading data loss for recovery.
  • Root-cause framing: distinguish a sustained throughput deficit from a transient spike, and check for a hot partition caused by a skewed key before blaming consumer count.
asked …
LeaderboardSalaryAccount