Cheng Su

Member Of Technical Staff at OpenAI

San Francisco, California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Cheng Su is a seasoned distributed systems engineer with 10 years of experience building and optimizing large-scale data infrastructure, currently a Member of Technical Staff on OpenAI's data infrastructure team in San Francisco. Previously an engineering manager at Anyscale, he led Ray Data efforts—improving data loading and batch inference workflows used by companies like Pinterest and Spotify—while still contributing significant IC work to the Ray project. At Facebook he worked on Spark, Hive and Hadoop at production scale, specializing in Spark SQL query optimization and join execution, and is an active contributor to the Apache Spark codebase. His open-source contributions include performance-focused enhancements to Spark’s SQL engine and practical improvements to Ray Data’s dataset APIs and Parquet handling. Cheng combines hands-on backend engineering with people leadership and a track record of shipping measurable performance gains for distributed ML and analytics workloads. He holds an MS in Computer Science from University of Wisconsin–Madison and brings a researcher’s rigor to production systems.
code10 years of coding experience
job7 years of employment as a software developer
bookMaster of Science Computer Science, Master of Science Computer Science at University of Wisconsin-Madison
bookBachelor of Science Computer Science, Bachelor of Science Computer Science at Nanjing University
bookHong Kong University of Science and Technology (HKUST)
github-logo-circle

Github Skills (29)

algorithms10
apache-spark10
python10
back-end-development10
data-science10
data-engineering10
machine-learning10
data-structure10
java10
datasets10
scala10
javas10
parquet10
data-structures10
performance-analytics9

Programming languages (8)

JavaC++RustScalaGoHTMLJupyter NotebookPython

Github contributions (5)

github-logo-circle
ray-project/ray

Jul 2022 - Jan 2023

Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Role in this project:
userBack-end Developer & Data Scientist
Contributions:1 release, 1558 reviews, 90 commits in 6 months
Contributions summary:Cheng primarily contributed to improving and implementing features within the Ray Data ecosystem. The code changes involved fixing actor pool strategies within datasets, adding support for the `drop_columns` API, and enhancing the documentation for this functionality. Additionally, the user addressed issues with Parquet file size estimations and improved the experience of reading from HTTP files. The user also made adjustments to test suite for different configurations.
pythonconsistsruntimetensorflowserving
apache/spark

Feb 2021 - Jul 2022

Apache Spark - A unified analytics engine for large-scale data processing
Role in this project:
userBack-end Developer
Contributions:631 reviews, 22 commits, 124 PRs in 1 year 5 months
Contributions summary:Cheng's commits focus on enhancing the Apache Spark SQL engine, specifically targeting performance improvements. They primarily worked on optimizing the execution of join operations, including shuffled hash joins and broadcast nested loop joins, adding code generation capabilities for better performance. Furthermore, they improved code organization by refactoring common logic, introducing more SQL metrics. Also, the user has made changes to support and enable new features to make more data sources compatible with spark.
analyticspythondata-processingsqlapache
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial