Roman Ageev

Staff Deep Learning Engineer (IC4), VLM at NVIDIA

Amsterdam, North Holland, Netherlands
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Roman Ageev is a Staff Deep Learning Engineer specializing in vision-language and multimodal LLM inference, currently optimizing production stacks at NVIDIA to cut latency, boost throughput, and reduce memory/cost on GPUs. With ~3 years of focused industry experience across startups and tech leaders (Stochastic, Yandex, Huawei, Intel, JetBrains), he brings a rare blend of research-driven modeling and hands-on systems work—quantization, FP8/INT8 tuning, TensorRT-LLM/Triton integration, and upstream fixes in vLLM and HF Transformers. He has shipped high-throughput production systems (200k rps at Yandex), built fine-tuning tooling (xTuring), and repeatedly turned SOTA ideas into robust, deployable pipelines. Based in Amsterdam, he favors simple processes for complex problems and often pairs architectural tweaks with practical profiling and automated performance gates to ensure reliability at scale.
code4 years of coding experience
job4 years of employment as a software developer
bookMaster's degree, Computer Science, Master's degree, Computer Science at Higher School of Economics
bookData Science, Data Science at Yandex School of Data Analysis
bookSaint Petersburg Lyceum 239
bookQuantitative Finance, Quantitative Finance at Center of Mathematical Finance
bookBachelor's degree, Applied Mathematics and Computer Science, Bachelor's degree, Applied Mathematics and Computer Science at ITMO University
github-logo-circle

Github Skills (34)

peft10
language-model10
discord10
llama10
llm10
data-preprocessing10
deep-learning10
generative-ai10
fine-tuning10
adapter10
mistral10
lora10
quantization10
notebook9
onnxruntime9

Programming languages (2)

Jupyter NotebookPython

Github contributions (5)

github-logo-circle
stochasticai/xTuring

Mar 2023 - Dec 2023

Build, personalize and control your own LLMs. From data pre-processing to fine-tuning, xTuring provides an easy way to personalize open-source LLMs. Join our discord community: https://discord.gg/TgHXuSJEk6
Contributions:6 releases, 34 reviews, 63 PRs in 8 months
data-preprocessingdiscordfine-tuningdeep-learninggpt-2
Real-time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. Join our Discord communty: https://discord.com/invite/TgHXuSJEk6
Contributions:3 reviews, 2 PRs, 2 pushes in 1 year 1 month
discordinferencestable-diffusiontensorrtpytorch
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial