Questions in System Design
Design a distributed data-processing engine (Spark/Delta-Lake style) — partitioning, shuffle, and job scheduling. The onsite also includes a dedicated concurrency/multithreading coding round (e.g., an efficient threaded logger). Cover functional/non-functional requirements, capacity estimation, API and data model, scaling strategy, caching, and fault tolerance.