Alex Matveev is a machine learning systems engineer and researcher with six years of industry experience and a PhD in computer science, currently working as a Member of Technical Staff at Red Hat. He co-founded Neural Magic and served as Chief Scientist, driving R&D on high-performance parallel execution engines for ML and AI across CPUs and GPUs. His hands-on work includes core contributions to the vllm project—optimizing the Marlin kernel for quantized LLMs, adding GPTQ 8-bit support, and resolving low-level GPU (H100) stability and prefill performance issues. Combining deep academic roots at MIT and Tel Aviv University with startup productization, he uniquely blends kernel-level performance tuning with deployable inference systems. Based in Cambridge, MA, he is known for squeezing production-grade throughput and memory efficiency out of modern ML hardware.
A high-throughput and memory-efficient inference and serving engine for LLMs
Role in this project:
ML Engineer
Contributions:314 reviews, 83 PRs, 160 pushes in 2 years 3 months
Contributions summary:Alexander's primary contributions focused on enhancing the Marlin kernel within the vllm-project/vllm repository, specifically targeting improvements for quantized large language models. They introduced support for GPTQ 8-bit models, fine-tuned configurations for GPTQ marlin, and addressed kernel-level crashes, especially concerning H100 hardware. They also benchmarked and improved prefill performance and added tests to ensure functionality.
A high-throughput and memory-efficient inference and serving engine for LLMs
Contributions:4 pushes, 2 branches in 1 year 6 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.