Roman Ageev is a Staff Deep Learning Engineer specializing in vision-language and multimodal LLM inference, currently optimizing production stacks at NVIDIA to cut latency, boost throughput, and reduce memory/cost on GPUs. With ~3 years of focused industry experience across startups and tech leaders (Stochastic, Yandex, Huawei, Intel, JetBrains), he brings a rare blend of research-driven modeling and hands-on systems work—quantization, FP8/INT8 tuning, TensorRT-LLM/Triton integration, and upstream fixes in vLLM and HF Transformers. He has shipped high-throughput production systems (200k rps at Yandex), built fine-tuning tooling (xTuring), and repeatedly turned SOTA ideas into robust, deployable pipelines. Based in Amsterdam, he favors simple processes for complex problems and often pairs architectural tweaks with practical profiling and automated performance gates to ensure reliability at scale.
4 years of coding experience
4 years of employment as a software developer
Master's degree, Computer Science, Master's degree, Computer Science at Higher School of Economics
Data Science, Data Science at Yandex School of Data Analysis
Saint Petersburg Lyceum 239
Quantitative Finance, Quantitative Finance at Center of Mathematical Finance
Bachelor's degree, Applied Mathematics and Computer Science, Bachelor's degree, Applied Mathematics and Computer Science at ITMO University
Build, personalize and control your own LLMs. From data pre-processing to fine-tuning, xTuring provides an easy way to personalize open-source LLMs. Join our discord community: https://discord.gg/TgHXuSJEk6
Contributions:6 releases, 34 reviews, 63 PRs in 8 months
Contributions:3 reviews, 2 PRs, 2 pushes in 1 year 1 month
discordinferencestable-diffusiontensorrtpytorch
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.