Wanchao L is a Staff Software Engineer at Meta with 11 years of experience building distributed systems and high-performance ML infrastructure, currently contributing to PyTorch’s DTensor and core distributed operators. Based in Menlo Park, he has deep expertise in sharding, optimizer integrations, and performance-critical operators (including flash-attention), and has improved the sharding cost model and op-dispatch logic that power large-scale model training. His background spans research and production roles—from cloud storage and data-center indexing to parameter-server ML work and recommendation-system libraries like torchrec—bridging academic rigor with production-grade engineering. Known for tackling hard backend problems, he combines systems-level thinking with hands-on optimization to accelerate distributed deep learning workflows.
11 years of coding experience
4 years of employment as a software developer
Master's degree Computer Science, Master's degree Computer Science at Carnegie Mellon University
Bachelor of Science (BS) Computer Science, Bachelor of Science (BS) Computer Science at Shanghai Jiao Tong University
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Role in this project:
Back-end Developer
Contributions:2271 reviews, 987 commits, 779 PRs in 4 years 7 months
Contributions summary:Wanchao made significant contributions to the PyTorch framework's distributed tensor (DTensor) module. They focused on refactoring the op dispatch logic, improving the sharding cost model, implementing the scaled dot product attention (flash-attention) operator, and enabling various optimizer functionalities (e.g. foreach and FSDP2). Their work centered on enhancing the performance and functionality of DTensor, with a strong emphasis on core distributed operations.
Contributions summary:Wanchao's contributions primarily focused on improving the `torchrec` library, fixing bugs, and implementing new features related to recommendation systems. They addressed issues in the sharding logic, specifically correcting calculations within `cw_sharding` and making adjustments to fused optimizer behavior. Furthermore, the user updated APIs, including introducing an init API. These modifications involved interacting with core components related to sharding, embeddings, and optimizers within the `torchrec` framework.
pytorchrecommendation-systemgpudeep-learningcuda
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.