Cheng Su is a seasoned distributed systems engineer with 10 years of experience building and optimizing large-scale data infrastructure, currently a Member of Technical Staff on OpenAI's data infrastructure team in San Francisco. Previously an engineering manager at Anyscale, he led Ray Data efforts—improving data loading and batch inference workflows used by companies like Pinterest and Spotify—while still contributing significant IC work to the Ray project. At Facebook he worked on Spark, Hive and Hadoop at production scale, specializing in Spark SQL query optimization and join execution, and is an active contributor to the Apache Spark codebase. His open-source contributions include performance-focused enhancements to Spark’s SQL engine and practical improvements to Ray Data’s dataset APIs and Parquet handling. Cheng combines hands-on backend engineering with people leadership and a track record of shipping measurable performance gains for distributed ML and analytics workloads. He holds an MS in Computer Science from University of Wisconsin–Madison and brings a researcher’s rigor to production systems.
10 years of coding experience
7 years of employment as a software developer
Master of Science Computer Science, Master of Science Computer Science at University of Wisconsin-Madison
Bachelor of Science Computer Science, Bachelor of Science Computer Science at Nanjing University
Hong Kong University of Science and Technology (HKUST)
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Role in this project:
Back-end Developer & Data Scientist
Contributions:1 release, 1558 reviews, 90 commits in 6 months
Contributions summary:Cheng primarily contributed to improving and implementing features within the Ray Data ecosystem. The code changes involved fixing actor pool strategies within datasets, adding support for the `drop_columns` API, and enhancing the documentation for this functionality. Additionally, the user addressed issues with Parquet file size estimations and improved the experience of reading from HTTP files. The user also made adjustments to test suite for different configurations.
Apache Spark - A unified analytics engine for large-scale data processing
Role in this project:
Back-end Developer
Contributions:631 reviews, 22 commits, 124 PRs in 1 year 5 months
Contributions summary:Cheng's commits focus on enhancing the Apache Spark SQL engine, specifically targeting performance improvements. They primarily worked on optimizing the execution of join operations, including shuffled hash joins and broadcast nested loop joins, adding code generation capabilities for better performance. Furthermore, they improved code organization by refactoring common logic, introducing more SQL metrics. Also, the user has made changes to support and enable new features to make more data sources compatible with spark.
analyticspythondata-processingsqlapache
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.