Chao Liu is a Senior Staff Engineer in Menlo Park with nine years of experience specializing in high-performance parallel computing and GPU-accelerated ML infrastructure. He led development of AMD’s Composable Kernel library and drove algorithm work in MIOpen, delivering low-level tensor and kernel optimizations across FP16/BF16 and 3D convolution workloads. At Meta now, he brings a proven track record of squeezing performance from hardware—benchmarked across platforms via contributions to DeepBench—and a background in aerospace and computational engineering that informs his systems-level thinking. Colleagues rely on him to bridge research and production: he translates numerical methods and parallel-mesh experience into production-grade, performance-portable ML primitives.
9 years of coding experience
2 years of employment as a software developer
Ph.D. Computational Engineering, Ph.D. Computational Engineering at The University of Tennessee at Chattanooga
M.S. Aerospace Engineering, M.S. Aerospace Engineering at Beihang University
Composable Kernel: Performance Portable Programming Model for Machine Learning Tensor Operators
Role in this project:
Back-end Developer & Performance Engineer
Contributions:953 reviews, 891 commits, 439 PRs in 1 year 3 months
Contributions summary:Chao primarily focused on modifying and optimizing the code for the `Composable Kernel` project, which targets machine learning tensor operators. Their contributions involved merging branches and making code changes within the `composable_kernel` directory, suggesting work related to improving the performance of machine learning kernels. The user's edits included changes to header files and code related to reduction operations and XDL operations, potentially for performance enhancement.
Contributions:235 reviews, 406 commits, 44 PRs in 3 years 6 months
Contributions summary:Chao's contributions primarily involve implementing and optimizing low-level tensor operations, and kernel implementations within the AMD's Machine Intelligence Library (MIOpen). The commits showcase the modification of kernel code for tensor copying, scaling, and setting, incorporating performance improvements and supporting various data types, including FP16/BF16. Furthermore, their work includes refactoring and enhancing the implementation for efficient handling of 3D convolution operations.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.