Drew Hudson is a Senior Research Scientist at DeepMind with nine years of experience bridging academic rigor and scalable ML systems, having completed a Ph.D. in Computer Science at Stanford's SAIL group. He specializes in representation learning at the intersection of vision and language, with notable contributions to multi-step visual reasoning (e.g., MAC, Neural State Machine) and compositional scene generation (Generative Adversarial Transformer). Drew has driven dataset and evaluation efforts—contributing to HELM’s QA scenarios and GQA—reflecting a commitment to robust benchmarks as well as models. His work combines a deep interest in compositionality, knowledge and reasoning with practical engineering (TFRecord pipelines and dataset tooling) to support large-scale training. Based in London, he pairs research leadership at DeepMind with hands-on open-source contributions that improve evaluation and data infrastructure for language and vision models.
9 years of coding experience
5 years of employment as a software developer
Transfer student Electrical Engineering and Computer Science (EECS), Transfer student Electrical Engineering and Computer Science (EECS) at Massachusetts Institute of Technology
Bachelor of Science (BSc) Computer Science, Bachelor of Science (BSc) Computer Science at Technion Institute of Technology
Doctor of Philosophy (Ph.D.) Computer Science, Doctor of Philosophy (Ph.D.) Computer Science at Stanford University
Contributions:2 releases, 320 commits, 6 PRs in 1 year 3 months
Contributions summary:Drew's commits primarily focused on the development of a tool for creating TFRecords datasets, essential for training machine learning models. They implemented functions to add images and labels, along with methods to handle dataset shuffling and prefetching. The code also includes utilities for managing threading and concurrent processing of image datasets, highlighting the importance of efficiency in data preparation for large-scale machine learning tasks.
Holistic Evaluation of Language Models (HELM), a framework to increase the transparency of language models (https://arxiv.org/abs/2211.09110). This framework is also used to evaluate text-to-image models in HEIM (https://arxiv.org/abs/2311.04287) and vision-language models in VHELM (https://arxiv.org/abs/2410.07112).
Role in this project:
ML Engineer
Contributions:44 commits in 6 months
Contributions summary:Drew contributed to the development and enhancement of the bAbI QA scenario, a question-answering dataset within the HELM framework. Their work focused on adapting the existing bAbI QA scenario by fixing answer formats, updating the prompts, and extending support for new question answering tasks. The changes also include adding the LSAT QA scenario, suggesting a broader focus on question-answering benchmarks. Overall, the contributions centered on improving the framework's capabilities in evaluating language models on question-answering tasks.
nlparxivabsberthelm
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.