Ye Zhou is an engineering manager in Mountain View with 11 years of experience building and scaling data-intensive distributed systems, currently leading efforts to improve observability and reliability of LinkedIn’s data pipelines. Previously a Staff Software Engineer and tech lead for Spark at LinkedIn, Ye drove the migration of Spark workloads from YARN to Kubernetes and was a main contributor to the push-based shuffle feature in Apache Spark. Their work on push-based shuffle and Spark History Server stability includes RPC coordination and synchronization fixes that improved performance and resilience at large scale, and contributed to the widely used Apache Spark project. Ye combines deep hands-on systems engineering—co-authoring a VLDB paper on Magnet—with team leadership that unites infrastructure, compute unification, and operational tooling. Comfortable operating at the intersection of research and production, Ye has a strong academic grounding from Carnegie Mellon and Huazhong University and a knack for turning complex distributed-compute problems into robust, deployable solutions.
11 years of coding experience
12 years of employment as a software developer
Master of Computer Architecture Virtualization, Master of Computer Architecture Virtualization at Huazhong University of Science and Technology
Master’s Degree Computational Data Science, Master’s Degree Computational Data Science at Carnegie Mellon University
Apache Spark - A unified analytics engine for large-scale data processing
Role in this project:
Back-end Developer
Contributions:224 reviews, 3 commits, 14 PRs in 2 months
Contributions summary:Ye primarily contributed to the Apache Spark codebase, focusing on improving the stability and functionality of the shuffle service and history server. Their work involved adding synchronization mechanisms to prevent concurrent modification issues in the history server and implementing RPC-based coordination for push-based shuffle, improving performance and reliability. Additionally, the user addressed issues related to multi-application attempts and data persistence during NodeManager restarts, enhancing the push-based shuffle implementation.
Apache Spark - A unified analytics engine for large-scale data processing
Contributions:116 commits, 102 pushes, 8 branches in 2 years 2 months
analyticsdata-processingapachebig-dataspark
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.