Five years into infrastructure, currently leading a team of five engineers across a genuinely multi-cloud estate — AWS, GCP, Azure, OVHcloud, plus a self-managed Kubernetes footprint on bare metal. Everything ships through Terraform and GitHub Actions: 100% IaC, zero manual changes, GitOps via ArgoCD, and an observability stack (Grafana, Loki, Prometheus) people actually check before things break. Clouds talk to each other over private P2P/VPC peering with VPN-only access — no public exposure — and branch protection, CODEOWNERS, and secret scanning are the default, not an afterthought.
Earlier in those five years: migrated 250+ CI/CD pipelines to GitHub Actions and built auto-scaling self-hosted runners, cutting CI spend 60%; trimmed cloud costs by $20K+/month through governance automation across 50+ AWS accounts; stood up Velero + ArgoCD disaster recovery for Kubernetes; got VM provisioning with Ansible down from hours to under 10 minutes; and supported a 3000+ server data-center migration while patching and securing 1800+ Linux VMs along the way. When production breaks, I'm usually the one running the incident and writing the RCA after.



