Shuang Ma

Research Scientist at Meta

Mountain View, California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Shuang Ma is a research scientist with a decade of experience building foundational language and multimodal models, currently working on GenAI and Llama training at Meta after leading Apple’s foundational language model efforts. At Apple she designed RLHF algorithms, built post-training pipelines, and drove data selection and synthetic data generation to scale Apple Intelligence Foundation Models, and earlier at Microsoft Research she advanced multimodal pretraining and representation learning for embodied agents. Her open-source contributions include impactful engineering to Microsoft’s torchscale—integrating BERT, Mixture-of-Experts components, and XPOS positional encodings—reflecting a strong bridge between research and production ML systems. Based in Mountain View with a PhD-focused background from University at Buffalo, she combines deep research rigor with hands-on pipeline and model engineering across both single- and multi-modal LLMs.
code10 years of coding experience
job5 years of employment as a software developer
bookDoctor of Philosophy (PhD) Computer Science, Doctor of Philosophy (PhD) Computer Science at University at Buffalo
languagesEnglish, Chinese
github-logo-circle

Github Skills (8)

pytorch10
transformer10
nlp10
bert10
fairseq9
machine-learning9
multimodal4
computer-vision3

Programming languages (3)

C++HTMLPython

Github contributions (5)

github-logo-circle
microsoft/torchscale

Nov 2022 - Mar 2023

Foundation Architecture for (M)LLMs
Role in this project:
userML Engineer
Contributions:40 commits, 23 PRs, 50 pushes in 3 months
Contributions summary:Shuang primarily focused on integrating and modifying machine learning models within the `torchscale` framework. Their contributions include incorporating BERT models, updating MoE (Mixture of Experts) components, and refactoring code for better compatibility with the latest versions of Fairseq. Furthermore, the user made changes to model configurations and introduced XPOS (XPositional Encoding) for enhanced performance. The user's work significantly impacts the core functionality of the project, aimed at building large language models.
pytorchnlptransformerspythonlanguage-modeling
lancopku/WEAN

Feb 2018 - Jan 2019

Contributions:12 commits, 11 pushes, 1 branch in 10 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial