Siying Dong is a Senior Staff Software Engineer with 12 years of production experience in distributed systems, storage and streaming, currently advancing Spark Structured Streaming at Databricks from Menlo Park. Previously a long-time tech lead and key developer of RocksDB at Facebook, she has deep expertise in embeddable key-value stores, range deletions, concurrency and memory optimizations that underpin high-throughput services. Her background includes work on HDFS, Hive and SQL Azure, giving her a rare end-to-end view of data platforms from ingestion to stateful stream processing. An active open-source contributor, she has improved Kafka integration and RocksDB-backed state stores in Apache Spark and contributed core features and fixes to facebook/rocksdb. Trained at Tsinghua and Brandeis, Siying combines rigorous academic foundations with pragmatic engineering leadership that focuses on performance, reliability and operational clarity.
12 years of coding experience
16 years of employment as a software developer
Bachelor's Degree Computer Science, Bachelor's Degree Computer Science at Tsinghua University
Master's Degree Computer Science, Master's Degree Computer Science at Brandeis University
A library that provides an embeddable, persistent key-value store for fast storage.
Role in this project:
Back-end Developer & Database Engineer
Contributions:21 releases, 470 reviews, 1715 commits in 9 years 5 months
Contributions summary:Siying primarily contributed to the RocksDB database project, focusing on the implementation of core features and enhancements. Their contributions included adding functionality to the LDB command-line tool, refining the performance of range deletions, and optimizing the handling of internal keys within the database's architecture. They also addressed bugs related to memory management, file handling, and concurrency.
Apache Spark - A unified analytics engine for large-scale data processing
Role in this project:
Back-end Developer
Contributions:102 reviews, 27 PRs, 171 comments in 1 year 11 months
Contributions summary:Siying primarily contributed to Apache Spark's streaming functionality, focusing on improvements to the Kafka integration and state store features. Their work includes enhancing logging for Kafka batch reading, addressing issues in FileStreamSource, and adding support for custom decimal types in Avro. They also implemented improvements to the RocksDB state store, including logging enhancements, compression configurations, and checkpoint ID handling, indicating involvement in optimizing and refining Spark's stateful streaming capabilities.
analyticspythondata-processingsqlapache
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.