Ryan Chi is a research scientist focused on evaluating and improving large language models, currently on the research team at OpenAI with prior research roles at DeepMind, Netflix, and Citadel. Over nine years he has blended academic rigor—leading Stanford NLP’s Alexa Prize team to a Science Innovation award and authoring work on adversarial robustness and bias—with industry-facing ML research and quant experience. He contributes to influential open-source evaluation efforts like HELM, where he implemented bias-detection scenarios and dataset-driven evaluations. A strong educator and organizer, he won Stanford’s Centennial TA Award and ran large-course logistics while mentoring new PhD students. Beyond NLP, his background spans probabilistic methods for synthetic healthcare data and applied quantitative work in trading, reflecting a rare mix of applied engineering, rigorous evaluation, and cross-domain curiosity.
9 years of coding experience
6 years of employment as a software developer
BS, Computer Science with Distinction (music minor), BS, Computer Science with Distinction (music minor) at Stanford University
Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.
Role in this project:
Backend Developer & Data Scientist
Contributions:86 commits in 5 months
Contributions summary:Ryan primarily contributed to the `helm` repository by implementing scenarios related to language models and bias analysis. Their work involved generating and evaluating language model outputs with a focus on the BBQ dataset. Key contributions include the development of a Dyck language scenario and modifications to existing BBQ and Civil Comments scenarios. These changes likely facilitate bias detection and model performance evaluation in the context of natural language processing.
Contributions:15 pushes, 1 branch in 3 years 6 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.