Henry Ho

Principle Member Of Technical Staff Software Engineer at AMD

New Taipei, Taiwan
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

👤
Senior
🎓
Top School
Henry Ho is a Principal Member of Technical Staff at AMD with eight years of experience optimizing GPU performance for AI/ML workloads, focused on GEMM and tensor contractions across MI100/MI200/MI300 architectures. He drives low-level assembly kernel optimizations, kernel selection algorithms, and scheduling strategies that have measurably boosted the Tensile/hipBLASLt stacks used in high-performance deep learning. Prior to AMD he designed OpenCL FPGA accelerators and production ML monitors, giving him a rare blend of silicon-aware systems expertise and applied ML engineering. Henry’s contributions include backend performance work in the well-known ROCm/Tensile project—improving register usage, interweaving packing with computation, and hardening edge cases. He combines top academic credentials from National Taiwan University with hands-on kernel-writing and compiler-aware scheduling skills that translate architecture insights into tangible throughput gains.
code8 years of coding experience
job9 years of employment as a software developer
bookBachelor's degree, Electronic Engineering, GPA 4.0/4.0, Ranking 7/141, Bachelor's degree, Electronic Engineering, GPA 4.0/4.0, Ranking 7/141 at 國立臺灣科技大學
bookMaster's degree, Electronic Engineering, GPA 4.3/4.3, Ranking 9/131, Master's degree, Electronic Engineering, GPA 4.3/4.3, Ranking 9/131 at 國立臺灣大學
languagesChinese, English
github-logo-circle

Github Skills (18)

auto-tuning10
assembly10
performance-monitor10
tensorrt10
matrix-multiplication10
contract10
performance-measurement10
performance-testing10
gemfire10
performance-analysis10
gpu10
performance-tuning10
contraction10
tensor10
performance-monitoring10

Programming languages (4)

C++CAssemblyPython

Github contributions (5)

github-logo-circle
ROCm/Tensile

Feb 2020 - Nov 2022

Stretching GPU performance for GEMMs and tensor contractions.
Role in this project:
userBack-end Developer & Performance Engineer
Contributions:50 reviews, 94 commits, 70 PRs in 2 years 9 months
Contributions summary:Henry primarily contributed to optimizing the performance of the Tensile library, specifically focusing on matrix multiplication and tensor contraction operations. Their work involved code changes to improve register usage and modify the scheduling algorithm to raise the priority of matrix operations, leading to performance gains. They also added features like interweaving packing operations with matrix computations and addressing edge cases. The user made modifications to kernel writers and structures to meet these goals.
amdpythontensorhipauto-tuning
aazz44ss/hipBLASLt

Jan 2023 - Apr 2025

Contributions:367 pushes, 136 branches in 2 years 2 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial