Qun Song

GPU Architect at NVIDIA

Shanghai, Shanghai, China
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Qun Song is a GPU architect at NVIDIA Shanghai with eight years of experience accelerating machine learning workloads across CPUs and GPUs. He combines deep systems expertise from internships at Intel, YITU, and Huawei Noah’s Ark Lab with a master’s in computer science from USC to drive fast kernel and inference optimizations. An active open-source contributor, he has commits to PyTorch improving tensor semantics and compilation passes and maintains an AArch64 inference kernel project that surfaces low‑level performance insights. Qun’s background spans Linux kernel testing, ML inference frameworks, and cross‑architecture optimization, making him adept at turning research-grade techniques into production GPU kernels.
code8 years of coding experience
job1 year of employment as a software developer
bookMaster of Science - MS, Computer Science, Master of Science - MS, Computer Science at University of Southern California
bookBachelor of Science - BS, Electronic and Computer Engineering, Bachelor of Science - BS, Electronic and Computer Engineering at 上海交通大学
github-logo-circle

Github Skills (10)

pytorch10
machine-learning10
deep-learning10
python10
autograd10
aws-dynamodb9
dynamodb9
amazon-dynamodb9
tensor9
neural-network8

Programming languages (2)

C++Python

Github contributions (5)

github-logo-circle
pytorch/pytorch

Oct 2018 - Dec 2024

Tensors and Dynamic neural networks in Python with strong GPU acceleration
Role in this project:
userBack-end Developer
Contributions:7 reviews, 12 PRs, 33 comments in 6 years 3 months
Contributions summary:Qun contributed to the PyTorch codebase by implementing and modifying core functionalities. Their work involved adding support for the `is_complex` method for tensors within the `dynamo` module. They also addressed issues in code generation, specifically focusing on deleting unused values, and fixed reordering issues in the compiled autograd. Furthermore, the user worked on allowing symbolic integer input in the reinplace pass.
pythongpu-accelerationdeep-learninggpunumpy
YangQun1/Paddle

Mar 2023 - Jun 2023

PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署)
Contributions:1 PR, 44 pushes, 7 branches in 3 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial