Jiaxing Qi is a solution architect with 8 years of experience bridging high-performance computing, deep learning, and production AI infrastructure. Currently at NVIDIA, he applies expertise in CUDA and LLMs to optimize large-scale GPU workloads, building on prior work at Didi where he restructured multi-GPU C++ training pipelines and optimized neural inference for mobile and server. His background is rooted in computational science—PhD work and research roles focused on massively parallel Lattice Boltzmann solvers and CFD—giving him a rare combination of low-level performance tuning and applied ML engineering. Fluent in building end-to-end systems from algorithm to deployment, he’s equally comfortable writing high-performance kernels and iterating acoustic and speaker-recognition models for real-world products. An interesting thread through his career is moving complex simulation techniques into production-grade, scalable implementations that now inform his work on GPU-accelerated AI.
9 years of coding experience
5 years of employment as a software developer
Exchange student, Computer Science, Exchange student, Computer Science at University of Houston
Bachelor, Biomedical Engineering, Bachelor, Biomedical Engineering at 华中科技大学
M.Sc, Biomedical Engineering, M.Sc, Biomedical Engineering at Huazhong University of Science and Technology
Doctor of Philosophy (Ph.D.), Mechanical Engineering, Doctor of Philosophy (Ph.D.), Mechanical Engineering at RWTH Aachen University
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper and Ada GPUs, to provide better performance with lower memory utilization in both training and inference.
Contributions:1 push, 1 branch, 1 comment in 1 year 4 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.