Summary
Wang Zhang is a software engineer with 11 years of experience specializing in cloud-native AI/ML infrastructure and distributed training systems, currently contributing to ByteDance from San Jose. He has driven elastic and fault-tolerant training features across Kubeflow/training-operator and co-authored projects like Kube-Queue and FTLib that optimize GPU sharing, job scheduling, and resilience for large-scale ML workloads. At Tencent and Caicloud he built multi-cluster RL schedulers and GPU-sharing solutions that reclaim fragmented and spot resources for actors and learners, demonstrating practical expertise in resource elasticity and scheduling plugins. An approver on Kubeflow’s training-operator and author of FTLib, he combines academic rigor from Carnegie Mellon with hands-on systems engineering to bridge low-level GPU techniques and higher-level orchestration. Notably, he helped design veRL’s single-controller architecture and has a track record of enabling elastic training across TensorFlow, PyTorch, Horovod and more.
11 years of coding experience
7 years of employment as a software developer
Master's degree, Mechanical Engineering, Master's degree, Mechanical Engineering at 美国卡内基梅隆大学
Bachelor's degree, Mechanical Engineering, Bachelor's degree, Mechanical Engineering at 浙江大学