Leo Gao

Software Engineer at Datadog

United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Leo Gao is a software engineer and applied mathematics student at Brandeis with six years of hands-on experience building scalable infrastructure, ML tooling, and backend systems. He contributes to high-profile open-source AI projects—optimizing attention mechanisms in EleutherAI's GPT-Neo and improving deployment and data pipelines for GPT-NeoX and The Pile—bridging model-level ML work with production DevOps. At Datadog he works on cloud-native stacks (Kubernetes, Terraform, Go, Postgres, Typescript), and he has a track record of shipping robust evaluation tooling for language models. Prior roles span product engineering, data pipelines, and operational leadership, from architecting React/TypeScript component libraries to scaling an e-commerce business and managing a busy restaurant. Comfortable switching between low-level ML optimizations and large-scale system automation, he brings pragmatic engineering instincts and an uncommon combination of applied math rigor and operational experience.
code6 years of coding experience
job6 years of employment as a software developer
bookAmity Regional High School
bookMathematics and Computer Science, Mathematics and Computer Science at Brandeis University
github-logo-circle

Github Skills (32)

kubernetes10
docker10
language-model10
data-pipelines10
python10
scripting10
data-engineering10
attention-mechanism10
machine-learning10
dockers10
data-preprocessing10
script10
tensorflow10
natural-language-processing10
language-modeling10

Programming languages (7)

C++ShellJavaScriptHTMLJupyter NotebookPythonClojure

Github contributions (5)

github-logo-circle
A framework for few-shot evaluation of language models.
Role in this project:
userBack-end Developer & ML Engineer
Contributions:2 releases, 159 reviews, 516 commits in 2 years 3 months
Contributions summary:Leo primarily contributed to the development of language model evaluation methods within the framework. Their work involved adding and modifying core methods to enable model evaluation, including an `evaluate` method and implementing loglikelihood calculations for the GPT2 language model. They also integrated datasets, such as the BoolQ dataset, which highlights their focus on enabling the evaluation of models on diverse benchmarks. Furthermore, the user implemented the gpt2 loglikelihood functionality.
pytorchnlplarge-language-modelslanguage-modeldeep-learning
EleutherAI/the-pile

Sep 2020 - Jun 2021

Role in this project:
userBack-end Developer & Data Engineer
Contributions:7 reviews, 162 commits, 10 PRs in 9 months
Contributions summary:Leo implemented crucial datasets and data processing pipelines for a large language model project. They added several datasets, including Deepmind Math, Enron Emails, and Literotica, expanding the scope and diversity of the training data. The user also integrated tools like checksums to ensure data integrity and included preprocessing steps and utility functions to handle and combine datasets.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial