Summary
Andrei Blinov is a Staff Site Reliability Engineer with 14+ years building and operating large-scale distributed systems, currently architecting LLM inference and multi-agent deployments that serve 50M+ tokens/day across cloud and bare-metal GPU clusters. He combines deep SRE practice—SLO/SLI frameworks, unified observability, and incident response for orgs of 200+ engineers—with hands-on experience running 1,000+ node fleets and defending against massive DDoS and malicious traffic. Andrei has repeatedly improved availability and recovery (e.g., <1h disaster recovery for multi-terabyte validators) and scaled real-time media and streaming platforms handling hundreds of GB/s and millions of concurrent connections. He brings a product-minded, entrepreneurial streak—mentoring startups and shaping design-thinking strategy—while translating research-grade LLM tooling (vLLM, Vertex AI) into production-grade, per-agent SLO-driven services. Based in Amsterdam, he blends embedded-systems roots with modern cloud and bare-metal GPU deployment expertise, making him adept at bridging low-level reliability and cutting-edge AI infra.
11 years of coding experience
13 years of employment as a software developer
Ignite - Innovation & People Management, Ignite - Innovation & People Management at Stanford University Graduate School of Business
Microelectronics and Embedded Systems, B.S., Microelectronics and Embedded Systems, B.S. at Nizhny Novgorod State Technical University n.a. R.E. Alekseev (NNSTU)
English, Russian, Dutch, Hebrew