Matthew Hayes

Senior Staff Software Engineer at Databricks

San Francisco, California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Matthew Hayes is a Senior Staff Software Engineer in San Francisco with 15 years building scalable ML and data platforms across companies from Microsoft and LinkedIn to Databricks. He combines deep engineering chops in distributed systems with hands-on machine learning work—most recently contributing to Databricks’ Dolly LLM training and generation pipelines to improve tokenization and usability. Previously a Principal ML Engineer at Workday and long-time Apache DataFu vice president, he blends open-source leadership with product-focused delivery. Matthew has a track record of moving research-grade models into production and mentoring teams through architectural transitions. He pairs an electrical engineering foundation from UCLA with practical experience across startups and large enterprises, and often surfaces optimizations that make ML workflows measurably more robust and reproducible.
code15 years of coding experience
job21 years of employment as a software developer
bookUniversity of California, Los Angeles
github-logo-circle

Github Skills (16)

transformers10
machine-learning10
pytorch10
nlp10
python10
natural-language-processing10
fine-tuning10
datasets9
data-set9
load-data9
tokenizer9
gpt9
text-generation9
data-loading9
chatbot8

Programming languages (9)

TypeScriptJavaShellCRustJavaScriptHTMLPython

Github contributions (5)

github-logo-circle
databrickslabs/dolly

Mar 2023 - Jun 2023

Databricks’ Dolly, a large language model trained on the Databricks Machine Learning Platform
Role in this project:
userML Engineer
Contributions:17 reviews, 27 PRs, 27 pushes in 3 months
Contributions summary:Matthew primarily contributed to the training and generation components of the Dolly large language model. They modified the trainer script, including dataset loading, preprocessing, and training arguments. The user also introduced text generation capabilities with the `InstructionTextGenerationPipeline` and improved the tokenization process. These changes suggest a focus on fine-tuning and improving the model's usability.
databrickslarge-language-modelsmachine-learninggptchatbot
apache/datafu

Apr 2014 - Apr 2020

Mirror of Apache DataFu
Contributions:1 review, 101 commits, 5 PRs in 6 years
datafuapachebig-datadatastreamjava
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial