Boxiang Wang is a Senior Research Engineer based in California with six years of experience building scalable AI systems and machine learning infrastructure. He currently contributes to NVIDIA Research on physical AI and world models after roles spanning deep learning algorithm engineering and large-scale ML systems at organizations including Alibaba, NUS, and HPC-AI Tech. His academic background includes advanced degrees from MIT and Harvard in computer science and computational science, and he has taught scalable AI topics as a guest lecturer. An active open-source contributor, he has improved 2D parallelism and CUDA kernel integrations in the widely used ColossalAI project to make large models cheaper and faster. Pragmatic about trade-offs, he blends systems-level engineering with algorithmic insight to push distributed training performance. Colleagues describe him as someone who pursues practical, high-impact research that bridges prototypes to production.
6 years of coding experience
3 years of employment as a software developer
Master of Science - MS, Computer Science, Master of Science - MS, Computer Science at Massachusetts Institute of Technology
Master of Science - S.M., Computational Science and Engineering, Master of Science - S.M., Computational Science and Engineering at Harvard University
High School Diploma, General Science, High School Diploma, General Science at Nankai High School
Bachelor of Engineering - BE, Electrical and Electronics Engineering, Bachelor of Engineering - BE, Electrical and Electronics Engineering at Nanyang Technological University Singapore
Making large AI models cheaper, faster and more accessible
Role in this project:
ML Engineer
Contributions:7 reviews, 9 commits, 37 PRs in 10 months
Contributions summary:Boxiang primarily updated and documented layer integration within the ColossalAI framework, specifically focusing on 2D parallelism operations. The changes include modifications to matrix multiplication and classifier functions, along with the addition of documentation, suggesting an effort to improve the usability and understanding of the library's distributed training capabilities. These updates involve changes to CUDA kernels and the implementation of layer normalization functions within the framework. This user's contribution focuses on supporting and enhancing the performance of large AI models by improving the functionality of 2D parallelism operations.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.