Michi Yasunaga is an AI researcher and Stanford PhD with nine years of experience focused on large language models, multimodality, retrieval, and reasoning. Their work spans top research labs including Meta and DeepMind and includes leading post-training and multimodal LLM research as well as developing novel reasoning models. Michi contributes to influential open-source evaluation tooling—helping improve the HELM framework used for holistic LLM and vision-language benchmarking—and has integrated language modeling datasets into the WILDS benchmark. Based in Palo Alto, they combine rigorous academic training with hands-on engineering, making experimental choices that prioritize reproducibility and interpretability. An under-the-radar strength is their attention to evaluation details (e.g., validation/test distinctions and prediction sorting) that materially improve replicability of benchmarks.
9 years of coding experience
8 years of employment as a software developer
Bachelor's degree, Computer Scinece, Mathematics, Bachelor's degree, Computer Scinece, Mathematics at Yale University
Doctor of Philosophy - PhD, Computer Science, Doctor of Philosophy - PhD, Computer Science at Stanford University
Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.
Role in this project:
ML Engineer
Contributions:2 reviews, 21 commits, 6 PRs in 1 year 2 months
Contributions summary:Michi contributed to the HELM framework, focusing on the evaluation of language models. Their work involved modifying metrics to differentiate between validation and test sets, which is crucial for replicating MMLU results. They also addressed issues related to the integration of the WikiFact scenario, including code cleanup and making the scenario more interpretable. Finally, the user implemented a fix for sorting predictions.
A machine learning benchmark of in-the-wild distribution shifts, with data loaders, evaluators, and default models.
Role in this project:
ML Engineer
Contributions:12 commits, 3 PRs in 8 months
Contributions summary:Michi primarily contributed to the Py150 dataset implementation within the wilds repository. Their work involved modifying the dataset class, including updates for token type metadata, evaluation metrics, and split handling. They also incorporated changes related to loss functions and configurations for the Py150 dataset within the example training scripts and configuration files. These commits demonstrate a focus on integrating a new dataset and adapting existing code for language modeling tasks.
dataloadermachine-learning
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.