Kaixi Hou is a Senior Deep Learning Software Engineer based in California with 11 years of experience building high-performance GPU and HPC software. At NVIDIA since 2018, Kaixi focuses on deep learning kernels, CUDA/cuDNN integration, and performance optimizations that power large-scale training and inference. He is an active open-source contributor to flagship projects like JAX, TensorFlow, XLA and Flax, where his work spans attention kernels, FP8 support, mixed-precision conv3d, and fused GPU ops—contributions that directly impact ML library performance on modern accelerators. His background includes research and engineering roles at Virginia Tech, Oak Ridge, and AMD, blending academic rigor with production-grade systems engineering. Beyond typical kernel work, Kaixi has repeatedly enabled subtle compatibility and precision improvements (CUDA version checks, padding/sliding-window attention, GQA/MQA support) that unlock real-world model performance. He brings a rare combination of low-level GPU expertise and practical ML engineering, translating compiler/runtime changes into measurable speedups for developers and researchers.
11 years of coding experience
Virginia Tech
Bachelor of Engineering (B.E.), Computer Science, Bachelor of Engineering (B.E.), Computer Science at Beijing University of Chemical Technology
DeepRec is a high-performance recommendation deep learning framework based on TensorFlow. It is hosted in incubation in LF AI & Data Foundation.
Role in this project:
ML Engineer
Contributions:52 commits in 1 year 9 months
Contributions summary:Kaixi's contributions primarily focused on enabling and testing Automatic Mixed Precision (AMP) support for 3D convolution operations within the TensorFlow-based deep learning framework. This involved modifying existing test files to include new tests for 3D convolution layers and their gradients, as well as updating configuration lists to include the relevant operators for mixed-precision optimization. Furthermore, the user worked on integrating and testing cuDNN CTC loss for the framework.
A machine learning compiler for GPUs, CPUs, and ML accelerators
Role in this project:
Back-end Developer
Contributions:10 reviews, 21 commits, 16 PRs in 3 years 6 months
Contributions summary:Kaixi primarily contributed to the XLA compiler, focusing on GPU-related functionalities. Their work involved implementing and modifying CUDA and cuDNN related components for batch normalization and fused operations, including Elu activation, by integrating cuDNN runtime fusion. Additionally, the user updated the code for MatMul+bias+tanh/sigmoid fusion and for BF16 support in MatMul operations. They also made changes to optimize the convolution algorithms and improve the device indexing in the XLA code.
compilercommunity-drivenmachine-learningmodular
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.