Design a Job Scheduler

Problem Design a distributed job scheduler: run one-off jobs at a given time and recurring (cron-like) jobs, reliably, across a worker fleet.

Functional requirements

  • submit(job, runAt | cronExpr), cancel, job status/history.
  • Per-job retry policy with backoff; priorities.
  • Misfire handling (what happens to runs missed during downtime).

Non-functional requirements

  • Scale to discuss: millions of scheduled jobs, thousands of executions/sec at peak.
  • A job must run even if the node that scheduled it dies.
  • No duplicate concurrent runs of the same job (or make duplicates safe).

Areas to go deep

  • The job store and the next_run_at query surface.
  • The trigger/claim mechanism that prevents double-firing.
  • Dispatch, worker leases for dead-worker recovery, partitioning, and exactly-once vs idempotency.
asked …
LeaderboardSalaryAccount