Ritesh Patel

Senior Deep Learning Architect at NVIDIA

Sacramento, California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

👤
Senior
🎓
Top School
Ritesh Patel is a Senior Deep Learning Architect with 15 years of experience optimizing ML performance across GPUs and large-scale clusters, currently focusing on end-to-end training performance for LLMs at NVIDIA. He has led performance model development and software-hardware co-design at Google and Intel, building predictive cost models for core ML ops (GEMM, convs, fusions) and influencing TPU/GPU architecture decisions. Ritesh combines low-level kernel implementation expertise with system-level profiling to drive tangible speedups—authoring cycle-accurate simulators and highly optimized compute shaders earlier in his career. His background includes building control systems for precision manufacturing and translating hardware behavior into actionable software optimizations, a blend that helps him spot non-obvious bottlenecks across the stack. Based in Sacramento, he brings deep practical experience in performance pathfinding, sharding/scheduling strategies, and operation fusion techniques that accelerate production-scale ML training.
code15 years of coding experience
job14 years of employment as a software developer
bookMaster of Science (M.S.) Electrical and Computer Engineering, Master of Science (M.S.) Electrical and Computer Engineering at University of California, Davis
languagesEnglish, Gujarati
github-logo-circle

Github Skills (15)

transformers10
transformer-models10
large-language-models10
cuda9
data-parallel9
huggingface8
gpu7
python7
nvidia7
floating-point6
pytorch6
deep-learning6
machine-learning6
inference4
jax4

Programming languages (2)

PythonCuda

Github contributions (5)

github-logo-circle
rapatel/TransformerEngine

Sep 2025 - Jun 2026

A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
Contributions:7 pushes, 2 branches in 9 months
rapatel/Megatron-LM

Dec 2025 - Jun 2026

Ongoing research training transformer models at scale
Contributions:23 pushes, 8 branches in 6 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial