Design a Distributed GPU Training Job Scheduler

Problem

Design a scheduler that places distributed deep-learning training jobs across a cluster of GPU nodes.

Requirements

Functional:

  • Submit multi-GPU/multi-node jobs
  • Gang-schedule all workers of a job together
  • Topology-aware placement (NVLink/InfiniBand locality)
  • Preempt and requeue lower-priority jobs

Non-functional:

  • 1000s of GPUs, mixed job sizes
  • High utilization, fair sharing
  • Fault tolerance for node failures

Discussion points

  1. Gang scheduling and bin-packing
  2. Topology-aware placement for collective bandwidth
  3. Checkpoint/restart on failure
  4. Priority, preemption, and fairness
  5. Telemetry: GPU utilization, queue wait
added …
LeaderboardSalaryAccount