Andrew White is a cofounder and CTO with 14 years of technical experience who blends academic rigor as an associate professor of chemical engineering with hands-on startup leadership in San Francisco. He leads technology and science at early-stage ventures (Edison Scientific and FutureHouse) while maintaining an active research and engineering footprint, particularly in ML-enabled document QA and molecular modeling tools. His open-source contributions span high-accuracy retrieval-augmented generation for scientific papers and performance and algorithmic improvements to molecular simulation and representation libraries (plumed2, SELFIES), showing fluency across backend systems, testing, and scientific software. Comfortable translating deep domain knowledge into production-ready code, he combines pedagogy, research, and product instincts to ship reproducible, well-tested scientific tooling.
14 years of coding experience
5 years of employment as a software developer
Doctor of Philosophy (PhD) Chemical Engineering, Doctor of Philosophy (PhD) Chemical Engineering at University of Washington
Chemical Engineering, Chemical Engineering at Otto-von-Guericke University Magdeburg
Bachelor of Science (BS) Chemical Engineering, Bachelor of Science (BS) Chemical Engineering at Rose-Hulman Institute of Technology
High accuracy RAG for answering questions from scientific documents with citations
Role in this project:
Back-end Developer & ML Engineer
Contributions:107 releases, 247 reviews, 84 commits in 1 month
Contributions summary:Andrew focused on implementing core functionality for the paper-qa project, which involves answering questions from scientific documents. They wrote initial drafts and modified the `docs.py`, `qaprompts.py`, and `tests/test_paperqa.py` files, indicating development of the core question-answering logic and testing infrastructure. The contributions suggest involvement in both back-end development and machine-learning related aspects, given the project's focus on "High accuracy RAG" and use of gpt-index (now replaced with FAISS).
Robust representation of semantically constrained graphs, in particular for molecules in chemistry
Role in this project:
Back-end Developer
Contributions:20 commits, 4 PRs, 21 comments in 11 months
Contributions summary:Andrew primarily focused on enhancing the `selfies` library, which is designed for representing molecules. Their contributions included adding caching mechanisms for the `get_semantic_robust_alphabet` function to improve performance. The user also implemented and enabled unit tests to ensure the library's functionality, specifically testing the cache-clearing mechanism and decoder attribution. Further work involved integrating SELFIES tokens with graph attribution for enhanced molecular representation.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.