Simon Cheung

Senior Data Engineer

Milpitas, California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Simon Cheung is a Senior Data Engineer based in Milpitas with eight years of hands-on experience building and optimizing data pipelines, ETL processes, and analytics platforms across finance and tech. He combines strong SQL and data-modeling chops with DevOps and SDLC discipline, having modernized analytics at CLSA and designed central client databases and KYC/AML workflows in financial services. An active Apache committer, Simon contributes to high-profile projects like HugeGraph, SeaTunnel, Kyuubi and DolphinScheduler—work that highlights deep backend expertise in graph databases, distributed data integration, and data lake integrations (including Hudi and Spark). Comfortable translating business requirements into specs, diagrams, and working code, he pairs technical delivery with documentation and stakeholder-facing experience from integration engineering to product scoping. Colleagues value him as a creative, reliable problem-solver who moves quickly from prototype to production while resolving complex dependency and performance issues.
code8 years of coding experience
job8 years of employment as a software developer
bookSan José State University
bookApplied Mathematics, Applied Mathematics at Santa Clara University
github-logo-circle

Github Skills (35)

spark-sql10
spark10
apache-hudi10
data-pipelines10
back-end-development10
apidoc10
data-engineering10
configuration-management10
databases10
graph-database10
java10
data-ingestion10
graphdb10
javas10
api10

Programming languages (8)

MDXJavaShellScalaJavaScriptGoJupyter NotebookPython

Github contributions (5)

github-logo-circle
apache/dolphinscheduler

Dec 2019 - Apr 2021

Apache DolphinScheduler is the modern data orchestration platform. Agile to create high performance workflow with low-code
Role in this project:
userBack-end Developer
Contributions:26 reviews, 109 commits, 42 PRs in 1 year 4 months
Contributions summary:Simon primarily contributed to the `apache/dolphinscheduler` project by fixing issues and adding features related to the DataX task. Their work included addressing and fixing bugs, as well as supporting custom DataX configurations by introducing new parameters and modifying relevant classes. The changes involved modifications to both the common and server modules, suggesting a focus on enhancing data orchestration capabilities. The user also provided unit tests to validate these changes.
dependenciesuser-interfacepipelineorchestrationdata-workflow
apache/seatunnel

Jul 2021 - Mar 2022

SeaTunnel is a next-generation super high-performance, distributed, massive data integration tool.
Role in this project:
userBack-end Developer
Contributions:51 reviews, 45 commits, 50 PRs in 8 months
Contributions summary:Simon primarily contributes to the `apache/seatunnel` repository by implementing new plugins and features related to data ingestion and integration. Their work includes adding a new plugin for Spark to read files, as well as a sink plugin for Kafka to write data. The commits also include refactoring code and fixing conflicts. Furthermore, the user demonstrated expertise in data processing by merging branches and resolving code conflicts within the configuration files.
servicenowseatunneldata-streametl-frameworkintegration-platform
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial