ZZomato·Tech KnowledgeL3System Design

How Kafka Works Internally

Problem Explain how Kafka works internally.

Be ready to discuss

  • Core abstraction: a distributed, append-only commit log. Producers write records to topics; topics split into partitions, which are the unit of parallelism and the only scope where ordering is guaranteed.
  • Partitioning: how a key hashes to a partition (and how a null key round-robins), why key choice determines both ordering guarantees and hot-partition risk.
  • Replication and durability: leader plus follower replicas, the in-sync replica (ISR) set, acks=0/1/all and min.insync.replicas, and leader election when a broker dies.
  • Why it is fast: sequential disk I/O plus the OS page cache rather than a userspace cache, zero-copy sends, and producer-side batching and compression to cut network overhead.
  • Consumers: consumer groups, one partition to at most one consumer per group, offsets tracked in the internal __consumer_offsets topic (historically ZooKeeper), and rebalancing when membership changes.
  • Delivery semantics: at-most-once, at-least-once, and exactly-once via idempotent producers plus transactions — and why at-least-once with idempotent consumers is the common production answer.
  • Retention and compaction: time/size-based retention versus log compaction for keyed state, and how retention interacts with a lagging consumer.
  • Metadata plane: ZooKeeper historically versus KRaft today for controller/quorum management.
asked …
LeaderboardSalaryAccount