Chen Fu is a Principal Software Engineer based in San Jose with seven years of recent industry experience and a long track record of building high-performance, reliable distributed systems across Microsoft, Alibaba, and Apple. He specializes in performance optimization and resource-efficient storage and database designs, having shipped innovations like the Exabyte Scavenger to tame write amplification and stabilize multi-tenant workloads. At Microsoft he has bridged research and production—applying model checking, Bayesian-driven incident mitigation, and automated diagnosis to shorten incident response from hours to minutes. Chen is also an active contributor to high-profile open-source ML infrastructure, optimizing ONNX Runtime for memory efficiency and 4-bit quantized matrix multiplications to accelerate inference. His PhD-level background in computer science underpins a pragmatic approach that blends deep systems research with measurable production impact. Quietly, he often focuses on reducing collateral resource spikes that disrupt collocated services, a recurring but underappreciated lever for improving cloud reliability.
7 years of coding experience
9 years of employment as a software developer
Master of Engineering Computer Engineering, Master of Engineering Computer Engineering at Institute of Computing, Chinese Academy of Science
Bachelor Computer Science, Bachelor Computer Science at Peking University
Ph.D Computer Science, Ph.D Computer Science at Rutgers University
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Role in this project:
ML Engineer
Contributions:330 reviews, 48 commits, 165 PRs in 1 year 10 months
Contributions summary:Chen primarily contributed to the performance optimization of the ONNX Runtime, focusing on memory usage and code efficiency within the context of machine learning inferencing. They identified and addressed issues related to buffer management in prepacked tensors, and worked on integrating and optimizing quantized matrix multiplication operations. This work included implementing optimizations for 4-bit quantization and implementing performance tests.
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Contributions:1 review, 1 PR, 1106 pushes in 3 years 5 months
pytorchdeep-learningruntimemachine-learningonnx
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.