Junjie Bai is a senior technology leader with 12 years of experience building production-grade AI systems, currently leading Physical AI and World Foundation Models at NVIDIA. He combines research-grade expertise in reinforcement learning and simulation with product and startup experience as a former co-founder of Lepton AI and director-level roles at Alibaba and Facebook AI. His background in computational science and mathematics from TUM informs a rigorous approach to model engineering, system optimization, and scalable ML platform design. Junjie has hands-on contributions to high-impact open-source ML tooling, including work on PyTorch/TensorRT integration and performance-focused fixes in the Caffe2 codebase. He has moved between quant trading, academic-style engineering, and large-scale AI platform work, giving him a rare mix of low-latency systems thinking and foundation-model scale experience. Based in Sunnyvale, he blends entrepreneurial pragmatism with deep technical craftsmanship to turn complex physical-world RL problems into deployable systems.
12 years of coding experience
8 years of employment as a software developer
Master of Science (MSc) Computational Science and Engineering, Master of Science (MSc) Computational Science and Engineering at Technical University of Munich
Caffe2 is a lightweight, modular, and scalable deep learning framework.
Role in this project:
Back-end Developer & ML Engineer
Contributions:74 commits, 77 PRs, 45 pushes in 4 months
Contributions summary:Junjie's commits primarily involve deprecating and refactoring parts of the `CNNModelHelper` within the Caffe2 deep learning framework. They made code changes to the `crf.py` and `memonger_test.py` files, which suggest work related to Conditional Random Fields and memory optimization testing. In addition, there were changes related to enabling top-k in the GPU accuracy operator and fixing a broken sequence to sequence example. These changes reflect a focus on improving the framework's functionality and efficiency for deep learning tasks.
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT
Role in this project:
ML Engineer
Contributions:8 commits, 5 PRs, 3 comments in 5 days
Contributions summary:Junjie focused on implementing and extending the functionality of the PyTorch/TorchScript/FX compiler for NVIDIA GPUs (TensorRT). Their contributions included fixing parameter order issues and build-related problems, adding support for activation functions such as sigmoid and tanh, and expanding the coverage of unary operations. Furthermore, they ensured that the inputs for CUDA engines were contiguous before passing the data pointer to TensorRT and made adjustments to input dimension padding.
compilernvidiapytorchtensorrttorchscript
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.