Henry Wu is a Staff Software Engineer based in Bellevue with 17+ years of experience building large-scale data and real-time systems across Microsoft, LinkedIn, Uber, and Airbnb. He specializes in Big Data, Distributed Systems, Machine Learning, microservices, and streaming, and has a track record of turning weekly batch pipelines into near-real-time platforms that power analytics, alerts, and incentive systems. At Airbnb he built a Text-to-SQL layer and curated semantic access to warehouses; at Uber he cut payout latency by four orders of magnitude with a real-time incentive system. A hands-on leader, he architects data processing frameworks using Spark, Flink, StarRocks, Hive, and Airflow while also shipping production ML services and microservices in Python/Java/Go. He contributes to notable open-source projects like mage-ai (adding Pub/Sub streaming and Snowflake batch uploads) and improves onboarding and docs for community tools like knowledge-repo. Henry combines deep systems-level engineering with practical product impact and a history of scaling teams and pipelines for mission-critical analytics.
7 years of coding experience
12 years of employment as a software developer
Master of Science (M.S.) Computer Science, Master of Science (M.S.) Computer Science at University of South Carolina
A next-generation curated knowledge sharing platform for data scientists and other technical professions.
Role in this project:
Full-stack Developer
Contributions:107 reviews, 255 commits, 122 PRs in 11 months
Contributions summary:Henry primarily focused on updating the documentation and installation instructions for the `knowledge-repo` project. They added detailed setup commands for virtual environments and clarified deployment procedures, enhancing the onboarding experience for new users. The user also made several code improvements, including efficient code changes, string formatting fixes, and raw string fixes within the `org.py` file. Furthermore, the user worked on adjusting the import statements to be alphabetically sorted.
🧙 Build, run, and manage data pipelines for integrating and transforming data.
Role in this project:
Data Engineer
Contributions:59 reviews, 3 commits, 126 PRs in 10 days
Contributions summary:Henry primarily contributed to data engineering tasks within the project. Their work included simplifying and refining existing code related to data catalog management and source integrations, specifically with BigQuery. They also upgraded the project's dependencies, such as pip, and addressed code quality issues by removing redundant imports and adding return type hints. Furthermore, the user added support for a new streaming source, Google Cloud PubSub, and enabled batch uploads for the Snowflake destination.
pythondatadbttransformationdata-quality
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.