Kaiyu Xie

Sr. Engineer at NVIDIA

Beijing, China
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Kaiyu Xie is a senior engineer with nine years of experience specializing in compute architecture and deep learning systems, currently working at NVIDIA in Beijing. He progressed from research and internship roles at SenseTime and Microsoft to engineering and senior engineering positions at NVIDIA, bringing practical research experience into production compute design. With a Master’s in Computer Science from Harbin Institute of Technology, he bridges algorithmic insight and systems-level implementation for high-performance AI workloads. Kaiyu’s background in both academic research and industry R&D suggests a knack for turning prototype models into scalable compute solutions that meet production constraints. Notably, his career path reflects a consistent focus on optimizing deep learning pipelines and compute stacks rather than only model research.
code9 years of coding experience
job3 years of employment as a software developer
bookBachelor's degree, Computer Science, Bachelor's degree, Computer Science at Harbin Engineering University
bookMaster's degree, Computer Science, Master's degree, Computer Science at Harbin Institute of Technology
github-logo-circle

Github Skills (49)

autoencoder10
toolbox9
benchmark9
pytorch9
openmmlab9
pan9
vision-transformer9
object-detection9
faster-rcnn9
transformer8
swin-transformer8
retinanet8
deep-learning8
mask-rcnn8
reinforcement-learning8

Programming languages (2)

C++Python

Github contributions (5)

github-logo-circle
NVIDIA/TensorRT-LLM

Oct 2023 - Apr 2025

TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
Contributions:7 releases, 83 reviews, 226 PRs in 1 year 5 months
The Triton TensorRT-LLM Backend
Contributions:9 reviews, 129 PRs, 149 pushes in 1 year 5 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial