Frank Hu

Engineering Manager Data Platform

San Francisco, California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Frank Hu is an engineering manager specializing in data platforms with 11 years of experience building high-performance, byte-efficient OLAP systems from startups to large-scale product teams. Based in San Francisco, he currently leads Data Platform engineering at TikTok after architecting real-time streaming pipelines and a distributed query engine at Optimizely that processed billions of events and petabytes of data. He combines hands-on systems work—contributing performance and Parquet support improvements to the widely used Presto project—with product-driven leadership, having co-founded MailTime and shipped an App Store–featured iOS product. Frank has deep expertise across Kafka, Samza, HBase, Presto, Parquet and Airflow, and has operationalized encryption-at-rest for petabyte-scale data. He is as comfortable optimizing compilers and memory management for query engines as he is scaling microservices and real-time analytics for customer-facing use cases. His background in research and UAV vision systems hints at a pragmatic curiosity that drives novel, cross-disciplinary solutions.
code11 years of coding experience
job4 years of employment as a software developer
bookExchange Student, Computer Science, Exchange Student, Computer Science at Dartmouth College
bookThe Chinese University of Hong Kong (CUHK)
languagesEnglish, Chinese, Chinese
github-logo-circle

Github Skills (10)

javas10
big-data10
presto10
sql10
java10
testing9
parquet9
query-optimization9
data-engineering8
hadoop8

Programming languages (7)

JavaC++ScalaJavaScriptGoObjective-CPython

Github contributions (5)

github-logo-circle
prestodb/presto

Jun 2020 - Apr 2021

The official home of the Presto distributed SQL query engine for big data
Role in this project:
userBack-end Developer
Contributions:6 reviews, 2 commits, 14 PRs in 10 months
Contributions summary:Frank contributed to the Presto distributed SQL query engine by optimizing common sub-expression handling within the `CursorProcessorCompiler`, improving query performance. Further contributions involved fixing memory management and data size calculations, ensuring data consistency. The user also enabled Parquet format support for TPC-H and TPC-DS tests, and addressed test flakiness. Finally, the user added TPC-DS test queries for Q2 and Q78 and disabled Q64.
distributed-sqlquerybigdataquery-enginesql
frankobe/presto

Apr 2021 - Jul 2023

The official home of the Presto distributed SQL query engine for big data
Contributions:58 pushes, 14 branches in 2 years 2 months
queryhivebigdataquery-enginesql
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial