Jinchen Ye is a software engineer based in the San Francisco Bay Area with six years of experience building high-performance ML systems and developer tools. A Rice University statistics master's graduate, he blends statistical rigor with hands-on systems work in C++, Python, R, SQL and GPU-focused toolchains like ROCm and CUDA. At nod.ai (acquired by AMD) and now AMD, he contributed to SHARK Studio’s dataset annotation UI and to the IREE compiler’s CUDA backend, implementing stream command buffers and other performance optimizations. He is comfortable across full-stack and backend domains—shipping Gradio-based UIs integrated with Google Cloud Storage as well as low-level compiler/runtime improvements. His profile combines practical data analysis roots with production ML infrastructure experience, including uncommon depth in both frontend annotation tooling and GPU codegen/streaming.
7 years of coding experience
4 years of employment as a software developer
Bachelor of Science - BS, Applied Statistics, Bachelor of Science - BS, Applied Statistics at South China Normal University
Master's degree, Statistics, Master's degree, Statistics at Rice University
AMD-SHARK Studio -- Web UI for SHARK+IREE High Performance Machine Learning Distribution
Role in this project:
Full-stack Developer
Contributions:25 reviews, 6 commits, 74 PRs in 28 days
Contributions summary:Jinchen contributed significantly to the SHARK Studio project, primarily focusing on developing the dataset annotation tool. Their work involved building the UI with Gradio, integrating with Google Cloud Storage, and implementing features for managing and annotating datasets. Further contributions included adding multiple prompt support, addressing TODOs related to the annotation process, and improving the overall user experience.
A retargetable MLIR-based machine learning compiler and runtime toolkit.
Role in this project:
Backend Engineer
Contributions:31 reviews, 1 commit, 31 PRs in 1 day
Contributions summary:Jinchen primarily contributed to the CUDA backend of the IREE compiler, focusing on implementing and integrating a stream command buffer for improved performance. Their work involved adding a block pool to the CUDA device, implementing the stream command buffer functionality, and integrating it into the CUDA backend. They also added a new operation for creating a parameter handle for tile sizes to be used with a MFMA based matmul codegen strategy. Furthermore, the user introduced a basic lowering pipeline without tiling and distribution for ROCm and LLVMGPU backends.
compilermachine-learningmlirvulkantensorflow
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.