Machine Learning Engineer - ML Runtime And Optimization
Sunnyvale, California, United States
Join Prog.AI to see contacts
Join Prog.AI to see contacts
Summary
👤
Senior
🎓
Top School
Shivin Devgon is a machine learning engineer with nine years of experience specializing in ML runtime, optimization, and high-performance model deployment. At Waymo he authored Triton/CUDA/PTX kernels (matrix multiply, flash attention, GLU, quantization) that accelerated vision transformer training and inference up to 3x vs JAX XLA and enabled model-sharding speedups of 7x. Previously at Singular Genomics he led stateful model architecture and end-to-end pipeline optimizations that cut calibration error 60%, reduced reads, sped training 4x, and produced multi-fold CPU/GPU and IO savings including int8 encodings and Triton-based deployment. He combines deep research training from UC Berkeley (BS/MS EECS, highest honors) with hands-on systems engineering across Triton, TensorRT, PyTorch, and distributed parallelism. Notably, he reduced routine engineering toil by automating pipelines and saved substantial cloud costs through compression and inference optimizations—demonstrating an uncommon blend of algorithmic insight and low-level performance hacking. Located in Sunnyvale, he focuses on pushing model efficiency without sacrificing quality.
8 years of coding experience
5 years of employment as a software developer
Master's of Science, Electrical Engineering and Computer Sciences, 3.92, Master's of Science, Electrical Engineering and Computer Sciences, 3.92 at UC Berkeley College of Engineering
Contributions:7 pushes, 1 branch in 2 years 4 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial
Shivin Devgon - Machine Learning Engineer - ML Runtime And Optimization