Haicheng Wu

Computer Architect at NVIDIA

Raleigh-Durham-Chapel Hill Area United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Haicheng Wu is a computer architect with 15 years of experience building high-performance compilers and GPU software, currently contributing to NVIDIA’s CUTLASS project to optimize CUDA linear algebra kernels. His background blends deep research—earning a PhD focused on heterogeneous CPU/GPU compilation and creating Red Fox, the first platform to run full TPC-H on GPUs—with industry impact at Qualcomm improving LLVM for ARM server CPUs. He has a track record of practical optimizations (kernel fusion/fission, resource allocation, race-condition fixes) that have produced multi‑order speedups in data‑intensive workloads. Based in the Raleigh-Durham-Chapel Hill area, he pairs rigorous academic training with hands-on open-source contributions to widely used GPU libraries, and is noted for surfacing subtle correctness and performance fixes in production-grade code.
code15 years of coding experience
job3 years of employment as a software developer
bookDoctor of Philosophy (Ph.D.), Electrical and Computer Engineering, Doctor of Philosophy (Ph.D.), Electrical and Computer Engineering at Georgia Institute of Technology
bookBachelor's degree, Electrical and Electronics Engineering, 86, Bachelor's degree, Electrical and Electronics Engineering, 86 at Shanghai Jiao Tong University
languagesChinese, English
stackoverflow-logo

Stackoverflow

Stats
1reputation
0reached
0answers
0questions
github-logo-circle

Github Skills (7)

cuda10
c-language10
cprogramming-language10
gpu10
linear-algebra9
gemfire9
cpp9

Programming languages (3)

C++PythonCuda

Github contributions (5)

github-logo-circle
NVIDIA/cutlass

Jul 2020 - Jan 2023

CUDA Templates for Linear Algebra Subroutines
Role in this project:
userBack-end Developer
Contributions:16 releases, 446 reviews, 65 commits in 2 years 6 months
Contributions summary:Haicheng contributed to the CUDA library by fixing typos, updating and improving existing code examples. The changes involve correcting matrix malloc sizes, adding bias vector support, and adding a leaky ReLU activation function. The user also added verification of the reduction tensor and addressed a race condition. Furthermore, the user refactored and optimized the GEMM reductionK fusion and enhanced the streamk load balance.
cudacpplinear-algebranvidiamatrix-multiplication
hwu36/cutlass

Jul 2020 - Oct 2024

CUDA Templates for Linear Algebra Subroutines
Contributions:17 pushes, 20 branches in 4 years 3 months
cudalinear-algebrandarraygpulinear
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial