Alex Gladkov is a Deep Learning Engineer in the San Francisco Bay Area with 6 years focused on model optimization, CUDA/Triton acceleration, and productionizing computer vision models for embedded and RTOS environments. He has driven deep learning infrastructure and custom CUDA kernels at Rivian and Meta, and earlier built image processing, camera drivers, and DSP systems at Amazon Lab126 and Broadcom. A frequent contributor to open-source ML tooling, he improved operator support and CUDA conv kernels in Apache TVM and enhanced model serving internals in AWS's Multi Model Server—work that bridges research-grade compiler improvements with production inference reliability. Alex combines low-level Linux kernel and embedded expertise with high-level model graph optimizations (quantization, pruning, NAS), enabling efficient deployment from silicon to cloud. An M.Sc. in Computer Science from the National Aerospace University underpins a career that uniquely spans FPGA/video codec roots to cutting-edge DL compilers and runtime engineering.
7 years of coding experience
18 years of employment as a software developer
M.Sc. Computer Science and Engineering, M.Sc. Computer Science and Engineering at National Aerospace University, USSR / Ukraine
Multi Model Server is a tool for serving neural net models for inference
Role in this project:
Back-end Developer
Contributions:7 commits, 29 PRs, 7 pushes in 2 months
Contributions summary:Alex primarily contributed to the back-end server logic, focusing on model loading and inference. Their work involved implementing model pre-load functionality, addressing memory management issues, and enhancing the worker thread management within the multi-model server. Key changes included modifying the worker thread behavior, adjusting buffer allocation, and incorporating improvements to model unloading. They also worked on test enhancements related to logging within the model server.
Contributions:4 reviews, 13 commits, 15 PRs in 1 year 3 months
Contributions summary:Alex primarily contributed to enhancing the TVM compiler stack by adding support for various Tensorflow operators, including log1p, cos, sin, and arctan, which involved changes across multiple files like tests, frontend implementations, and mathematical function definitions. They extended the compiler's functionality by integrating MXNet operators such as pad, slice, cos, sin, arctan, and 1D convolution/deconvolution. Moreover, the user contributed to CUDA improvements by optimizing conv2d and conv2d_transpose for different layouts and refining the conv3d implementations. They also worked on fast exponent implementation.
compilermachine-learningtensordeep-learninggpu
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.