Bulk employee upload over a single-create API
Problem
Customers send us employee data as spreadsheets, and the ops team currently copies the rows into the product one at a time. Replace that with a bulk upload: the customer hands over a file and the system creates every employee in it. There is an existing employee service with a POST /employees endpoint that creates one employee per call. It is a separate service reached over the network — you do not write into its database, and you are expected to mock it. There is no product spec: agree one with the interviewer first, then build a working slice.
Requirements
POST /bulk-uploads— accept a spreadsheet (CSV/XLSX), return ajob_idimmediately. Processing is asynchronous.GET /bulk-uploads/{job_id}— poll job status: counts per state plus the per-row errors, so the customer can fix and re-submit.- Every row becomes one
POST /employeescall against the external service, which may fail transiently (timeout, 5xx, 429) or permanently (validation, duplicate). - Rows that fail transiently are retried; rows still failing after the attempt cap are dead-lettered and surfaced in the status response rather than silently dropped.
Areas to design
- What an employee record actually is — deciding the required columns, header mapping, and per-row validation is part of the exercise, and validation should happen before any network call.
- Row-level state machine — e.g.
pending → in_progress → succeeded | failed | dead_lettered— and where that state is persisted so a worker crash does not lose it. - Retry policy — which errors are retryable, backoff strategy, attempt cap, and the dead-letter store that catches the remainder.
- Idempotency — the same file, row, or redelivered message must not create duplicate employees.
- Concurrency and backpressure — how many rows are in flight, and how you avoid overwhelming a service built for single creates.
- Job rollup and partial success — the job finishes even when some rows fail; what "done" means and what the customer gets back.
What's evaluated How you cut an ambiguous problem down to a thin, working slice — what you build first versus explicitly defer — and how fluently you direct an AI coding agent to get there while still owning the design decisions.