Yingyi Bu is a Senior Staff Software Engineer in San Francisco with 13 years of deep expertise in distributed query engines, query optimization, and low-latency analytics. She has driven major performance wins at Databricks and Google—leading initiatives that halved median query latency and delivered multi‑fold speedups across production fleets—and was the first SWE in Databricks’ Query Optimization team. Her open-source contributions to projects like Apache Spark and Delta Lake show hands‑on improvements to query compilation, tree traversal pruning, and data skipping that directly boost large-scale analytics performance. A founder‑engineer for Couchbase Analytics and long‑time Apache AsterixDB contributor, she blends research rigor (PhD-level training) with product delivery to ship robust, scalable systems. Notably, she couples compiler‑level changes with system‑level tuning, a combination that repeatedly yields both faster queries and reduced compilation cost.
13 years of coding experience
11 years of employment as a software developer
The Chinese University of Hong Kong (CUHK)
Bachelor of Science Computer Science, Bachelor of Science Computer Science at Nanjing University
Contributions summary:Yingyi made several commits focused on addressing issues related to static type casting and record constructors in the AsterixDB runtime environment. The contributions involved modifying Java code to ensure correct type checking and handling of function expressions, specifically within the `NonTaggedDataFormat.java` and `FlowRecordDescriptor.java` files. The changes appear to involve modifying class descriptors for new features and to address existing bugs.
Apache Spark - A unified analytics engine for large-scale data processing
Role in this project:
Back-end Developer
Contributions:214 reviews, 3 commits, 24 PRs in 10 months
Contributions summary:Yingyi contributed to the Apache Spark codebase, focusing on query compilation and optimization. Their work involved modifying core components like `TreeNode` and `AnalysisHelper` to enable early stopping in the transform and resolve function families, thus improving query compilation time. They implemented changes related to rule ids and condition lambdas for tree traversal pruning, which directly impacts the performance of query analysis and optimization. The commits also involved changes to existing rules like ReorderJoin and OptimizeIn.
analyticspythondata-processingsqlapache
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.