Michael Goldfarb

Senior Deep Learning Performance Engineer at NVIDIA

Greater Chicago Area United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Michael Goldfarb is a Senior Deep Learning Performance Engineer at NVIDIA with a strong cross-disciplinary background spanning compilers, ASIC design, parallel and accelerated computing, and ML systems. He has repeatedly led compiler and performance efforts—bootstrapping EnCharge AI’s compiler team for in-memory compute, driving AI100 accelerator software at Qualcomm, and optimizing training stacks and CUDA kernels at NVIDIA. Michael specializes in ML compiler toolchains (LLVM, XLA, MLIR, TVM, Glow), performance modeling, and low-level C++/Python/assembly optimization for transformers and other modern architectures. His experience includes practical hardware/software codesign and the unusual ability to move between RTL and high-performance kernel code, enabling end-to-end performance wins from silicon to training loops. Based in the Chicago area with a Purdue ECE background, he brings a blend of research-grade systems thinking and hands-on engineering that accelerates LLM training at scale.
code2 years of coding experience
job13 years of employment as a software developer
bookMasters Electrical and Computer Engineering, Masters Electrical and Computer Engineering at Purdue University
github-logo-circle

Github Skills (20)

transformers10
pytorch10
python10
machine-learning10
inference10
transformer-models10
nvidia10
deep-learning10
gpu10
cuda10
floating-point10
jax10
large-language-models9
fine-tuning9
mistral9

Programming languages (2)

C++Python

Github contributions (5)

github-logo-circle
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper and Ada GPUs, to provide better performance with lower memory utilization in both training and inference.
Contributions:172 pushes, 14 branches in 1 year 3 months
floating-pointinferencenvidiatransformer-modelstransformers
mgoldfarb-nvidia/jax

Aug 2024 - Jul 2026

Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
Contributions:23 pushes, 5 branches in 1 year 11 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial