Shangdi Yu is a research scientist at Meta with a strong foundation in machine learning, systems, and graph mining developed during a PhD at MIT and earlier dual undergraduate majors in Computer Science and Operations Research from Cornell. He has nine years of hands-on experience across research and industry, including internships at Google and The New York Times and contributions to PyTorch and Functorch—notably improving aten.norm, batch-norm backward, and implementing CSE passes to make AOTAutograd compilation more memory- and compute-efficient. He combines production-facing engineering (rematerialization/checkpointing, profiling utilities, GPU utilization metrics) with rigorous academic research in network modeling and data-driven systems. A long-time teaching assistant and course grader, he communicates complex ideas clearly across undergrad and graduate audiences. His background in building tooling, visualization pipelines, and web/mobile prototypes shows fluency across the stack from algorithms to deployment. Less obvious: he pairs deep low-level ML runtime work with applied data analysis experience from projects that modeled real-world systems like bike-share rebalancing and music-streaming dynamics.
9 years of coding experience
4 years of employment as a software developer
Doctor of Philosophy - PhD Computer Science, Doctor of Philosophy - PhD Computer Science at Massachusetts Institute of Technology
Bachelor of Science (B.S.) Computer Science & Operations Research, Bachelor of Science (B.S.) Computer Science & Operations Research at Cornell University
Bachelor of Science - BS Computer Science & Operations Research, Bachelor of Science - BS Computer Science & Operations Research at Cornell Engineering
functorch is JAX-like composable function transforms for PyTorch.
Role in this project:
ML Engineer
Contributions:9 reviews, 200 commits, 21 PRs in 2 months
Contributions summary:Shangdi contributed to the `functorch` repository, which focuses on composable function transforms for PyTorch. Their work primarily involved optimizing and enhancing the compilation process within the framework, including implementing common subexpression elimination (CSE) in AOTAutograd to improve memory efficiency. Additionally, the user worked on utilities for profiling and performance analysis, such as dumping chrome traces and calculating GPU utilization metrics. Furthermore, the user modified code to prepare the compilation for jit.script.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Role in this project:
ML Engineer
Contributions:250 reviews, 85 commits, 186 PRs in 2 months
Contributions summary:Shangdi contributed to the PyTorch library by implementing and refining decompositions for various operations. They focused on improving the `aten.norm` and `aten.native_batch_norm_backward` operations. The user also worked on a Common Subexpression Elimination (CSE) pass within the AOTAutograd and Functorch components, aiming to optimize the compilation process, and added graph dumping utilities for debugging.
pythongpu-accelerationdeep-learninggpunumpy
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.