Design Cross-Cloud Data Sync for Backup (Azure to AWS)
Problem All reads and writes are served from Azure; AWS must be maintained as an asynchronous backup copy of the same data. Design the cross-cloud syncing mechanism.
Functional requirements
- Every committed write on Azure eventually appears in AWS.
- Support insert, update and delete propagation.
- Detect and repair drift between the two clouds.
- Provide a measurable, alertable replication lag.
- Support a bulk backfill / re-seed of AWS from Azure.
Non-functional requirements
- ~50k writes/sec sustained on Azure, ~150k/sec at peak; average record ~2 KB → ~100 MB/sec steady, ~300 MB/sec peak of cross-cloud egress.
- Total dataset ~50 TB; a full re-seed must complete inside 24 h.
- Target replication lag p99 < 30 s; alert at > 5 min; RPO ≤ 1 min.
- Azure write path must not regress: added p99 latency < 5 ms, and AWS being fully down must never fail an Azure write.
- Cross-region egress cost is a first-class constraint at ~250 TB/month.
Key components
- Change capture: either dual-write from the application into a durable queue, or (preferred) CDC off the Azure database's transaction log so capture cannot diverge from what actually committed.
- Durable queue/log (Kafka / Event Hubs) partitioned by record key, sized to buffer several hours of backlog.
- Replicator worker pool: consumes the log, batches and compresses, writes to AWS with retries and a dead-letter queue.
- Batch reconciliation job: a scheduled bulk sync (e.g. hourly) that reconciles AWS against Azure when the real-time path falls behind or gaps are detected.
- Checksum/anti-entropy job: Merkle-tree or per-range checksums over key ranges to detect silent drift, feeding a repair queue.
- Monitoring: consumer lag, DLQ depth, drift-repair counts, egress bytes.
Deep dives / trade-offs
- Dual-write vs CDC: dual-write is trivial to build but is not atomic with the DB commit — a crash between commit and enqueue silently loses a record forever. CDC from the transaction log is exactly the durable, ordered source of truth you want, at the cost of schema-evolution handling and operational complexity.
- Ordering and idempotency: partition the queue by record id so per-record order is preserved; global order is neither achievable nor needed. Every apply must be idempotent — use a version/sequence number and last-write-wins on the target, so redelivery cannot resurrect a deleted row or roll back a newer value.
- Overload/backpressure: when the write rate exceeds drain capacity the backlog grows unbounded. Fall back to periodic batch reconciliation (a scheduled job that bulk-copies deltas) instead of trying to replay a hopeless real-time queue, and shed the log rather than block Azure.
- Deletes and tombstones: a delete that races ahead of a pending update, or a compacted log that drops the tombstone, resurrects data. Discuss tombstone retention windows.
- Failover semantics: this is asynchronous, so promoting AWS means accepting the in-flight RPO. Be explicit about what "backup" guarantees and what data is lost on an unplanned cutover.
asked …