Design Cross-Cloud Data Sync for Backup (Azure to AWS)

Problem All reads and writes are served from Azure; AWS must be maintained as an asynchronous backup copy of the same data. Design the cross-cloud syncing mechanism.

Functional requirements

  • Every committed write on Azure eventually appears in AWS.
  • Support insert, update and delete propagation.
  • Detect and repair drift between the two clouds.
  • Provide a measurable, alertable replication lag.
  • Support a bulk backfill / re-seed of AWS from Azure.

Non-functional requirements

  • ~50k writes/sec sustained on Azure, ~150k/sec at peak; average record ~2 KB → ~100 MB/sec steady, ~300 MB/sec peak of cross-cloud egress.
  • Total dataset ~50 TB; a full re-seed must complete inside 24 h.
  • Target replication lag p99 < 30 s; alert at > 5 min; RPO ≤ 1 min.
  • Azure write path must not regress: added p99 latency < 5 ms, and AWS being fully down must never fail an Azure write.
  • Cross-region egress cost is a first-class constraint at ~250 TB/month.

Key components

  • Change capture: either dual-write from the application into a durable queue, or (preferred) CDC off the Azure database's transaction log so capture cannot diverge from what actually committed.
  • Durable queue/log (Kafka / Event Hubs) partitioned by record key, sized to buffer several hours of backlog.
  • Replicator worker pool: consumes the log, batches and compresses, writes to AWS with retries and a dead-letter queue.
  • Batch reconciliation job: a scheduled bulk sync (e.g. hourly) that reconciles AWS against Azure when the real-time path falls behind or gaps are detected.
  • Checksum/anti-entropy job: Merkle-tree or per-range checksums over key ranges to detect silent drift, feeding a repair queue.
  • Monitoring: consumer lag, DLQ depth, drift-repair counts, egress bytes.

Deep dives / trade-offs

  • Dual-write vs CDC: dual-write is trivial to build but is not atomic with the DB commit — a crash between commit and enqueue silently loses a record forever. CDC from the transaction log is exactly the durable, ordered source of truth you want, at the cost of schema-evolution handling and operational complexity.
  • Ordering and idempotency: partition the queue by record id so per-record order is preserved; global order is neither achievable nor needed. Every apply must be idempotent — use a version/sequence number and last-write-wins on the target, so redelivery cannot resurrect a deleted row or roll back a newer value.
  • Overload/backpressure: when the write rate exceeds drain capacity the backlog grows unbounded. Fall back to periodic batch reconciliation (a scheduled job that bulk-copies deltas) instead of trying to replay a hopeless real-time queue, and shed the log rather than block Azure.
  • Deletes and tombstones: a delete that races ahead of a pending update, or a compacted log that drops the tombstone, resurrects data. Discuss tombstone retention windows.
  • Failover semantics: this is asynchronous, so promoting AWS means accepting the in-flight RPO. Be explicit about what "backup" guarantees and what data is lost on an unplanned cutover.
asked …
LeaderboardSalaryAccount