Ye Zhou

Engineering Manager at LinkedIn

Mountain View, California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Ye Zhou is an engineering manager in Mountain View with 11 years of experience building and scaling data-intensive distributed systems, currently leading efforts to improve observability and reliability of LinkedIn’s data pipelines. Previously a Staff Software Engineer and tech lead for Spark at LinkedIn, Ye drove the migration of Spark workloads from YARN to Kubernetes and was a main contributor to the push-based shuffle feature in Apache Spark. Their work on push-based shuffle and Spark History Server stability includes RPC coordination and synchronization fixes that improved performance and resilience at large scale, and contributed to the widely used Apache Spark project. Ye combines deep hands-on systems engineering—co-authoring a VLDB paper on Magnet—with team leadership that unites infrastructure, compute unification, and operational tooling. Comfortable operating at the intersection of research and production, Ye has a strong academic grounding from Carnegie Mellon and Huazhong University and a knack for turning complex distributed-compute problems into robust, deployable solutions.
code11 years of coding experience
job12 years of employment as a software developer
bookMaster of Computer Architecture Virtualization, Master of Computer Architecture Virtualization at Huazhong University of Science and Technology
bookMaster’s Degree Computational Data Science, Master’s Degree Computational Data Science at Carnegie Mellon University
github-logo-circle

Github Skills (8)

javas10
big-data10
spark10
shuffle10
java10
scala10
sql4
jdbc4

Programming languages (2)

JavaScala

Github contributions (5)

github-logo-circle
apache/spark

Jul 2021 - Oct 2021

Apache Spark - A unified analytics engine for large-scale data processing
Role in this project:
userBack-end Developer
Contributions:224 reviews, 3 commits, 14 PRs in 2 months
Contributions summary:Ye primarily contributed to the Apache Spark codebase, focusing on improving the stability and functionality of the shuffle service and history server. Their work involved adding synchronization mechanisms to prevent concurrent modification issues in the history server and implementing RPC-based coordination for push-based shuffle, improving performance and reliability. Additionally, the user addressed issues related to multi-application attempts and data persistence during NodeManager restarts, enhancing the push-based shuffle implementation.
analyticspythondata-processingsqlapache
linkedin/spark

May 2020 - Jul 2022

Apache Spark - A unified analytics engine for large-scale data processing
Contributions:116 commits, 102 pushes, 8 branches in 2 years 2 months
analyticsdata-processingapachebig-dataspark
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial