Siying Dong

Senior Staff Software Engineer at Databricks

Menlo Park, California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Siying Dong is a Senior Staff Software Engineer with 12 years of production experience in distributed systems, storage and streaming, currently advancing Spark Structured Streaming at Databricks from Menlo Park. Previously a long-time tech lead and key developer of RocksDB at Facebook, she has deep expertise in embeddable key-value stores, range deletions, concurrency and memory optimizations that underpin high-throughput services. Her background includes work on HDFS, Hive and SQL Azure, giving her a rare end-to-end view of data platforms from ingestion to stateful stream processing. An active open-source contributor, she has improved Kafka integration and RocksDB-backed state stores in Apache Spark and contributed core features and fixes to facebook/rocksdb. Trained at Tsinghua and Brandeis, Siying combines rigorous academic foundations with pragmatic engineering leadership that focuses on performance, reliability and operational clarity.
code12 years of coding experience
job16 years of employment as a software developer
bookBachelor's Degree Computer Science, Bachelor's Degree Computer Science at Tsinghua University
bookMaster's Degree Computer Science, Master's Degree Computer Science at Brandeis University
languagesEnglish, Chinese
stackoverflow-logo

Stackoverflow

Stats
146reputation
7kreached
9answers
0questions
github-logo-circle

Github Skills (28)

filesystem10
c-language10
spark10
rocksdb10
databases10
big-data10
kafka10
data-structure10
java10
scala10
javas10
sql10
data-structures10
cprogramming-language10
relational-databases10

Programming languages (6)

LiquidC++CScalaJavaScriptNunjucks

Github contributions (5)

github-logo-circle
facebook/rocksdb

Oct 2013 - Jan 2023

A library that provides an embeddable, persistent key-value store for fast storage.
Role in this project:
userBack-end Developer & Database Engineer
Contributions:21 releases, 470 reviews, 1715 commits in 9 years 5 months
Contributions summary:Siying primarily contributed to the RocksDB database project, focusing on the implementation of core features and enhancements. Their contributions included adding functionality to the LDB command-line tool, refining the performance of range deletions, and optimizing the handling of internal keys within the database's architecture. They also addressed bugs related to memory management, file handling, and concurrency.
persistent-storagebigtablefast-storagelsm-treedatabase
apache/spark

Apr 2023 - Mar 2025

Apache Spark - A unified analytics engine for large-scale data processing
Role in this project:
userBack-end Developer
Contributions:102 reviews, 27 PRs, 171 comments in 1 year 11 months
Contributions summary:Siying primarily contributed to Apache Spark's streaming functionality, focusing on improvements to the Kafka integration and state store features. Their work includes enhancing logging for Kafka batch reading, addressing issues in FileStreamSource, and adding support for custom decimal types in Avro. They also implemented improvements to the RocksDB state store, including logging enhancements, compression configurations, and checkpoint ID handling, indicating involvement in optimizing and refining Spark's stateful streaming capabilities.
analyticspythondata-processingsqlapache
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial