Yineng Zhang

Senior Director at Together AI

San Francisco, California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Yineng Zhang is a Senior Director leading the inference team at Together AI, bringing six years of focused experience in production ML systems and model-serving optimization. He has progressed from hands-on core developer work on SGLang—where he implemented InternLM2 support, integrated FlashInfer, and improved test and build stability—to leading inference strategy and execution at multiple startups. Yineng combines deep backend engineering (GPU and compilation-aware optimizations) with cross-cutting research contributions showcased in MLSys and ACM venues, and has spoken at industry events including AMD AI Dev Day and PyTorch Conference. Based in San Francisco, he has driven model performance teams at Baseten and large-scale ML infrastructure at Meituan and Baidu, consistently turning research into deployable inference solutions. Colleagues describe him as pragmatic and curious—his GitHub bio “Just for fun 🌁” hints at a playful approach to complex systems engineering.
code6 years of coding experience
job7 years of employment as a software developer
bookBachelor of Engineering, ENGINEERING, Bachelor of Engineering, ENGINEERING at Jiangnan University
languagesChinese, English
github-logo-circle

Github Skills (11)

code-optimization10
pytorch10
integrate10
python10
integrations10
data-integration10
deep-q-learning9
deep-learning9
cuda9
cprogramming-language5
c-language5

Programming languages (11)

TypeScriptC++ShellJavaScriptVueHTMLSwiftJupyter Notebook

Github contributions (5)

github-logo-circle
sgl-project/sglang

Jul 2024 - Apr 2025

SGLang is a fast serving framework for large language models and vision language models.
Role in this project:
userBack-end Developer
Contributions:6 releases, 651 reviews, 1285 PRs in 8 months
Contributions summary:Yineng implemented support for the InternLM2 model within the SGLang framework, indicating a focus on expanding the framework's capabilities to support various language models. They added a pre-commit configuration to improve code quality. The user also added unit tests and corrected the code for the MTP in a follow-up commit, improving stability. They also integrated with FlashInfer and addressed potential issues during compilation, showcasing a focus on optimization and compatibility.
cudadeepseekdeepseek-llmdeepseek-r1deepseek-r1-zero
zhyncs/sglang

Jul 2024 - Oct 2024

SGLang is yet another fast serving framework for large language models and vision language models.
Contributions:627 pushes, 151 branches, 1 tag in 2 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial