GitHub Actions · Data platform resilience
Led the zero-downtime re-architecture of the persistence layer behind GitHub Actions for higher availability and 100× scale.
I led the migration of critical workflow data from a 24 TiB MySQL cluster to Azure Cosmos DB while Actions continued handling more than 200 million jobs a day. We introduced store-agnostic routing, dual writes, shadow reads, parity gates, one-flag rollback, session consistency, and multi-region failover so every cutover remained observable and reversible.
The result: an architecture targeting at least five-nines availability with 100× scale runway, drastically reducing the platform-wide availability and recovery risk of a saturated, single-write-path database cluster.
- Daily scale
- 200M+ jobs
- Scale runway
- 100×
- Delivery
- 3 months