Leo Gao is a software engineer and applied mathematics student at Brandeis with six years of hands-on experience building scalable infrastructure, ML tooling, and backend systems. He contributes to high-profile open-source AI projects—optimizing attention mechanisms in EleutherAI's GPT-Neo and improving deployment and data pipelines for GPT-NeoX and The Pile—bridging model-level ML work with production DevOps. At Datadog he works on cloud-native stacks (Kubernetes, Terraform, Go, Postgres, Typescript), and he has a track record of shipping robust evaluation tooling for language models. Prior roles span product engineering, data pipelines, and operational leadership, from architecting React/TypeScript component libraries to scaling an e-commerce business and managing a busy restaurant. Comfortable switching between low-level ML optimizations and large-scale system automation, he brings pragmatic engineering instincts and an uncommon combination of applied math rigor and operational experience.
6 years of coding experience
6 years of employment as a software developer
Amity Regional High School
Mathematics and Computer Science, Mathematics and Computer Science at Brandeis University
A framework for few-shot evaluation of language models.
Role in this project:
Back-end Developer & ML Engineer
Contributions:2 releases, 159 reviews, 516 commits in 2 years 3 months
Contributions summary:Leo primarily contributed to the development of language model evaluation methods within the framework. Their work involved adding and modifying core methods to enable model evaluation, including an `evaluate` method and implementing loglikelihood calculations for the GPT2 language model. They also integrated datasets, such as the BoolQ dataset, which highlights their focus on enabling the evaluation of models on diverse benchmarks. Furthermore, the user implemented the gpt2 loglikelihood functionality.
Contributions:7 reviews, 162 commits, 10 PRs in 9 months
Contributions summary:Leo implemented crucial datasets and data processing pipelines for a large language model project. They added several datasets, including Deepmind Math, Enron Emails, and Literotica, expanding the scope and diversity of the training data. The user also integrated tools like checksums to ensure data integrity and included preprocessing steps and utility functions to handle and combine datasets.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.