Sr. Software Engineer In Deep Learning Frameworks at NVIDIA
Gig Harbor, Washington, United States
Join Prog.AI to see contacts
Join Prog.AI to see contacts
Summary
🤩
Rockstar
🎓
Top School
Boris Fomitchev is a senior software engineer with 15+ years building high-performance, heterogeneous systems and deep learning frameworks, currently leading CUDA/CuDNN integration and deployment work at NVIDIA. He combines advanced C++ and low-level systems expertise with hands-on GPU acceleration for image processing and ML, contributing to flagship projects like Torch, Theano, PyTorch/TensorRT and NVIDIA NeMo. His background spans production-grade REST APIs and large-scale backend services for eBay and PayPal, plus long-term work on Adobe’s image-processing core and real-time demos such as GTC’s HD Video Style Transfer. A frequent open-source contributor, he has practical experience making CUDA/CMake tooling robust across architectures (including Volta and fp16) and enabling ONNX/TorchScript exports for model deployment. Notably, he authored compatibility patches and half-precision optimizations in widely used ML repos and submitted a proposal to add a short-float type into the C/C++ standards. Based in Gig Harbor, WA, he brings a rare blend of compiler/simulator roots, production engineering, and ML deployment know-how.
15 years of coding experience
16 years of employment as a software developer
Master of Science (MS), Computer Science, A-, Master of Science (MS), Computer Science, A- at Moscow Institute of Physics and Technology (State University) (MIPT)
Contributions:112 commits, 52 PRs, 15 pushes in 1 year 6 months
Contributions summary:Boris contributed significantly to the Torch-7 FFI bindings for NVIDIA CuDNN, focusing on adding and updating features. Their primary contribution was implementing compatibility patches for `cudnn_v4`. These included adding support for `cudnn4 Batch Normaliztion` and converting the existing `addTensor_v2` function to the latest API. The user also focused on updating various CuDNN API functions like `cudnnConvolution` and `cudnnActivation`, which demonstrates deep expertise in low-level numerical libraries.
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
Role in this project:
Back-end Developer & ML Engineer
Contributions:246 reviews, 81 commits, 161 PRs in 3 years 2 months
Contributions summary:Boris's commits focus on exporting and integrating the NeMo framework with ONNX and TorchScript for deployment. They modified existing code to enable ONNX export for several models, particularly those related to ASR, TTS and NLP, and also adjusted code to address inconsistencies with tools like TensorRT. The user also made adjustments related to mixed-precision training, indicating an awareness of performance optimization and deployment considerations in model serving.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.