Matthew Hayes is a Senior Staff Software Engineer in San Francisco with 15 years building scalable ML and data platforms across companies from Microsoft and LinkedIn to Databricks. He combines deep engineering chops in distributed systems with hands-on machine learning work—most recently contributing to Databricks’ Dolly LLM training and generation pipelines to improve tokenization and usability. Previously a Principal ML Engineer at Workday and long-time Apache DataFu vice president, he blends open-source leadership with product-focused delivery. Matthew has a track record of moving research-grade models into production and mentoring teams through architectural transitions. He pairs an electrical engineering foundation from UCLA with practical experience across startups and large enterprises, and often surfaces optimizations that make ML workflows measurably more robust and reproducible.
Databricks’ Dolly, a large language model trained on the Databricks Machine Learning Platform
Role in this project:
ML Engineer
Contributions:17 reviews, 27 PRs, 27 pushes in 3 months
Contributions summary:Matthew primarily contributed to the training and generation components of the Dolly large language model. They modified the trainer script, including dataset loading, preprocessing, and training arguments. The user also introduced text generation capabilities with the `InstructionTextGenerationPipeline` and improved the tokenization process. These changes suggest a focus on fine-tuning and improving the model's usability.
Contributions:1 review, 101 commits, 5 PRs in 6 years
datafuapachebig-datadatastreamjava
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.