Sarah Yurick is a Senior Software Engineer based in San Francisco with six years of experience building high-performance data and ML infrastructure, now contributing at NVIDIA. She combines back-end systems expertise with data-science sensibilities, demonstrated by significant open-source contributions to projects like Apache's Rust SQL parser and NVIDIA's cuDF GPU DataFrame library where she implemented new SQL constructs, datetime features, and dtype compatibility fixes. Her background includes machine learning infrastructure work integrating RAPIDS with AWS SageMaker and practical data engineering and research roles across industry and academia. Comfortable in both low-level parsing and higher-level data workflows, she brings a pragmatic focus on correctness and interoperability. An atypical detail: she once ran a small mouse-breeding microbusiness during college, reflecting resourcefulness and hands-on care outside of engineering.
6 years of coding experience
2 years of employment as a software developer
Master of Science - MS Computer Science, Master of Science - MS Computer Science at Case Western Reserve University
Contributions:8 reviews, 6 commits, 5 PRs in 3 months
Contributions summary:Sarah primarily contributed to the `apache/datafusion-sqlparser-rs` repository by implementing and refining SQL parsing logic. Their commits focused on expanding the parser's capabilities by adding support for new keywords, expressions (like `IS UNKNOWN` and `CEIL`/`FLOOR` with `DateTimeField`), and data types (`DATE`). They also addressed parsing issues and corrected inconsistencies, contributing to the overall robustness and completeness of the SQL parser. The changes involved modifications to the lexer, parser, and related data structures.
Contributions:20 reviews, 7 commits, 11 PRs in 15 days
Contributions summary:Sarah primarily contributed to the `cudf` library, focusing on enhancing its DataFrame functionalities and improving compatibility with Pandas. Their work included fixing bugs related to data type handling in `select_dtypes`, implementing the `is_month_start` feature for datetime series, and enabling the use of Pandas dtype aliases. They also addressed issues with inserting `cudf.NA` values and improved the `where()` function. These changes demonstrate a focus on improving the usability and functionality of the library, particularly its data manipulation capabilities.
cudadataframe-librarydata-analysiscppcudf
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.