Ning Wang is a software engineer with a decade of experience building large-scale machine learning and computer vision systems, currently working at Meta Superintelligence Labs after a stint focused on efficient GenAI model training at Databricks. He has deep expertise in distributed PyTorch model parallelism and recommendation-system infrastructure, contributing notable fixes and performance work to high-profile open-source projects like torchrec and mosaicml/composer. Ning’s background spans production computer vision at Baidu to optimizing inference and sharding strategies, and he routinely improves robustness in training (e.g., resumable profilers, FSDP resharding and checkpoint retry mechanisms). Based in Menlo Park, he blends research-caliber ML knowledge with pragmatic backend engineering and a track record of shipping maintainable, high-performance code across industry-scale systems.
11 years of coding experience
10 years of employment as a software developer
Bachelor of Science (BS) Mathematics and Computer Science, Bachelor of Science (BS) Mathematics and Computer Science at Northeastern University, China
Master’s Degree Computer Science, Master’s Degree Computer Science at New York University - Polytechnic School of Engineering
Contributions:3 releases, 119 reviews, 88 PRs in 1 year
Contributions summary:Ning primarily contributed to the composer/profiler and torch_profiler modules, addressing issues related to profiling and resuming training runs. They implemented fixes to accurately handle the skip_first parameter within the cyclic schedule during resumption. They also moved an after_load callback to the profiler, along with fixing unit tests. Further contributions included FSDP resharding after an OOM condition and the addition of retry mechanisms for checkpoint downloads.
Contributions:3 reviews, 14 commits, 14 PRs in 9 months
Contributions summary:Ning contributed to the PyTorch domain library for recommendation systems by implementing and improving core functionalities related to model performance and distributed training. They focused on optimizing inference performance by migrating code to C++ and improving the KeyedJaggedTensor (KJT) operations. Furthermore, the user made improvements to the embedding tower sharding and partitioner, supporting distributed model training and deployment. Additional contributions included fixing unit tests and adding documentation, showcasing a focus on code quality and maintainability.
pytorchrecommendation-systemgpudeep-learningcuda
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.