Hao Jiang is a Senior Software Engineer at Databricks with 13 years of industry experience and a PhD in Computer Science from the University of Chicago. He specializes in storage systems, columnar data stores, compression, key-value stores and FPGA-driven optimization, and currently contributes to Delta Lake core features that improve schema handling and Hive Metastore interoperability. Before industry research at Databricks he was a postdoc at Harvard building self-designing adaptive blockchain systems, blending academic rigor with production engineering. Hao began his career as a J2EE architect and database administrator, bringing deep expertise in scalable enterprise systems, Oracle/MySQL performance tuning, and ORM frameworks like Spring and Hibernate. He has led distributed development teams and has hands-on experience modifying MySQL replication internals, a signal of his willingness to work at both high and low levels of the stack. Based in Sunnyvale, he combines research-level problem solving with practical contributions to a widely used open-source Lakehouse project.
13 years of coding experience
18 years of employment as a software developer
Bachelor of Science (B.S.), Computer Science, Bachelor of Science (B.S.), Computer Science at Fudan University
Doctor of Philosophy (PhD), Computer Science, Doctor of Philosophy (PhD), Computer Science at The University of Chicago
Master of Science (M.S.), Computer Science, Master of Science (M.S.), Computer Science at Clarkson University
An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Trino, and Hive and APIs
Role in this project:
Back-end Developer
Contributions:49 reviews, 46 PRs, 21 comments in 1 year 6 months
Contributions summary:Hao primarily contributed to the core functionalities of the Delta project, specifically focusing on improving the handling of `varchar` data types and related schema comparisons within the Delta Lake framework. They addressed an issue where `varchar` columns were not correctly compared due to Spark's internal representation, implementing a fix by converting the schema. Additionally, the user refactored and tested various components, including those related to error messages. Furthermore, they were involved in adding support for the UniForm feature, enabling interoperability with Hive Metastore (HMS).
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.