Zijie Poh is a Staff Machine Learning Engineer in the San Francisco Bay Area with eight years of experience applying research-grade ML to large-scale production problems across fintech and autonomous driving. With a PhD in particle physics earned in an accelerated four-year path, he blends rigorous statistical thinking with practical ML engineering using Python, Scala, Spark, and cloud platforms like AWS EMR. He has led and scaled prediction teams at Cruise and PayPal, translating research papers into robust, interpretable models and pipelines. An active open-source contributor, Zijie has improved core scientific libraries such as NumPy and scikit-learn and enhanced model-interpretability tooling in Yellowbrick and PyJanitor’s PySpark integration. He’s known for fixing numerical stability edge cases and adding nuanced parsing and diagnostics—work that quietly improves reliability for many downstream users. Colleagues describe him as a fast learner, collaborative manager, and hands-on implementer who bridges research and production.
8 years of coding experience
10 years of employment as a software developer
Bachelor of Arts (BA) - magna cum laude Physics Mathematics, Bachelor of Arts (BA) - magna cum laude Physics Mathematics at Ohio Wesleyan University
Doctor of Philosophy - PhD Physics, Doctor of Philosophy - PhD Physics at The Ohio State University
Visual analysis and diagnostic tools to facilitate machine learning model selection.
Role in this project:
Data Scientist
Contributions:7 commits, 6 PRs, 61 comments in 9 months
Contributions summary:Zijie contributed significantly to the yellowbrick library, enhancing its visualization capabilities for machine learning. They added a timer utility and integrated it into the Manifold visualizer for performance analysis. Further contributions included enhancing the FeatureImportances visualizer to support multi-dimensional coefficients and the development of a new FeatureCorrelation visualizer, demonstrating a focus on model interpretability and feature analysis. Additionally, the user fixed a bug in the PrecisionRecallCurve visualizer related to multi-class labels.
Contributions:12 commits, 10 PRs, 47 comments in 6 months
Contributions summary:Zijie primarily contributed to improving the scikit-learn library. Their work involved refactoring and optimization within various modules, including preprocessing, neighbor algorithms, metrics, and kernel implementations. A significant portion of their commits focused on fixing numerical stability issues, particularly those related to `np.full`, and ensuring robust performance across different data and parameter configurations. Additionally, the user addressed documentation inconsistencies and implemented improvements to error messages to enhance usability.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.