Design a Distributed GPU Training Job Scheduler
Problem
Design a scheduler that places distributed deep-learning training jobs across a cluster of GPU nodes.
Requirements
Functional:
- Submit multi-GPU/multi-node jobs
- Gang-schedule all workers of a job together
- Topology-aware placement (NVLink/InfiniBand locality)
- Preempt and requeue lower-priority jobs
Non-functional:
- 1000s of GPUs, mixed job sizes
- High utilization, fair sharing
- Fault tolerance for node failures
Discussion points
- Gang scheduling and bin-packing
- Topology-aware placement for collective bandwidth
- Checkpoint/restart on failure
- Priority, preemption, and fairness
- Telemetry: GPU utilization, queue wait
added …