Design a Job Scheduler
Problem Design a distributed job scheduler: run one-off jobs at a given time and recurring (cron-like) jobs, reliably, across a worker fleet.
Functional requirements
submit(job, runAt | cronExpr), cancel, job status/history.- Per-job retry policy with backoff; priorities.
- Misfire handling (what happens to runs missed during downtime).
Non-functional requirements
- Scale to discuss: millions of scheduled jobs, thousands of executions/sec at peak.
- A job must run even if the node that scheduled it dies.
- No duplicate concurrent runs of the same job (or make duplicates safe).
Areas to go deep
- The job store and the
next_run_atquery surface. - The trigger/claim mechanism that prevents double-firing.
- Dispatch, worker leases for dead-worker recovery, partitioning, and exactly-once vs idempotency.
asked …