Site reliability engineer with 8 years keeping large edge and API platforms fast and available for millions of users. Automates toil, designs systems that fail gracefully, and leads calm, blameless incident response across time zones. Holds the CKA and has run on-call for teams on three continents.
- Owned reliability for an edge network serving 30B requests a day across 90 cities at 99.99% availability.
- Automated certificate rotation and node draining, removing 20 hours of manual toil per week.
- Introduced SLOs and error budgets for 14 services; paging volume dropped 45% in 6 months.
- Led the response to 3 major incidents and published blameless reviews that drove 25 follow-up fixes.
- Built chaos experiments for regional failover, proving recovery in under 2 minutes.
- Cut monthly cloud spend 18% by right-sizing clusters and scheduling batch jobs off-peak.
- Migrated 40 services to Kubernetes with zero customer-facing downtime.
- Built Terraform modules that cut new-environment setup from 2 days to 30 minutes.
- Replaced ad-hoc scripts with a CI/CD pipeline, taking releases from monthly to daily.
- Centralised logs and metrics for 200 hosts, cutting mean time to detect from 30 to 5 minutes.
- Designed disaster-recovery drills that cut restore time for the core database from 6 hours to 40 minutes.
- Mentored 3 engineers who went on to lead on-call for their teams.
- Administered 300 Linux servers for 12 enterprise clients with 99.9% uptime.
- Automated patching with Ansible, reducing monthly maintenance windows from 8 hours to 2.
- Documented runbooks for 40 common alerts, used by the whole on-call team.
- Monitored backup jobs for all clients and fixed 18 silent failures.