Zhengyu C is a research scientist based in San Jose with a decade of experience building production-grade ML systems, currently focused on large language and multimodal foundation models at ByteDance/TikTok. Trained at Tsinghua and Carnegie Mellon with cross-continental research stints, he combines deep academic rigor with hands-on MLOps and systems engineering. His background spans perception and autonomy at TuSimple, scalable data tooling at A9, and open-source contributions to Kubernetes-centric projects like Koordinator and Kubeflow that improve runtime scheduling and distributed training reliability. Equally comfortable in C++ for embedded GPU deployments and Python/PyTorch for model work, he has a proven track record of turning terabytes of data into automated training pipelines and more reliable production workloads. Beyond research, he brings a pragmatic interest in quant and investment strategies, reflecting a broader appetite for systems that optimize both performance and value.
10 years of coding experience
2 years of employment as a software developer
Visiting Scholar, Computer Science, Visiting Scholar, Computer Science at University of California, Berkeley
Bachelor's degree, Software Engineering, 89/100, Bachelor's degree, Software Engineering, 89/100 at Tsinghua University
High School Diploma, High School Diploma at Wuhan Foreign Languages School
Master's degree, Computer Science, 3.83/4.0, Master's degree, Computer Science, 3.83/4.0 at Carnegie Mellon University
Exchange Student, Computer Science, 4.0/4.0, Exchange Student, Computer Science, 4.0/4.0 at The University of Texas at Austin
A QoS-based scheduling system brings optimal layout and status to workloads such as microservices, web services, big data jobs, AI jobs, etc.
Role in this project:
Back-end & DevOps Engineer
Contributions:32 reviews, 17 commits, 17 PRs in 2 months
Contributions summary:Zhengyu primarily contributed to the runtime proxy component of the Koordinator project, implementing and modifying APIs and interfaces related to container resource management. They added features like hooks for pre- and post-container operations, and container ID integration. Moreover, the user worked on ensuring proper cgroup parent configurations for Docker and systemd integrations, along with fixing PLEG-related cgroup path issues. These changes suggest a strong focus on runtime environment integration and management.
Distributed ML Training and Fine-Tuning on Kubernetes
Role in this project:
MLOps Engineer
Contributions:6 reviews, 6 commits, 9 PRs in 4 months
Contributions summary:Zhengyu primarily contributed to the improvement and maintenance of Kubeflow's training infrastructure. Their commits focused on modifying Kubernetes resource configurations like MPIJobs and TFJobs, including bug fixes related to gang scheduling and restart policies. They also updated job status handling for PyTorch and XGBoost jobs, suggesting a focus on monitoring, logging, and overall operational efficiency within the training environment.
xgboostkubernetesmachine-learningtrainingai
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.