Summary
Bilal Ahmad is a Lead Site Reliability Engineer based in London with a decade of experience building resilient, automated cloud platforms and developer tooling. He leads SRE teams to full infrastructure automation and has architected observability and incident management practices using Grafana OSS, Prometheus, Loki, Mimir and Tempo. Previously he designed serverless SaaS on AWS, modernized monolithic systems into microservices, and cut cloud costs through optimization and Terraform/Ansible-driven IaC. Bilal blends hands-on coding in Go and Python with platform-level strategy, and has introduced governance like AWS Control Tower and Okta to scale multi-account environments securely. He’s equally focused on developer experience—building in-house CLI/tooling and sandbox environments to speed onboarding and on-call ownership. A practical systems thinker, he often surfaces organizational improvements (SLOs, RCA) as part of technical delivery rather than as separate initiatives.
10 years of coding experience