Tom Hepworth

Lead Data Scientist

London, England, United Kingdom
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Tom Hepworth is a Lead Data Scientist based in London with five years of hands-on experience building large-scale data engineering and linkage systems for the UK public sector. He specialises in probabilistic record linkage and UK address matching using Splink, ensemble techniques, Python-first engineering and AWS compute to deliver performant, production-grade pipelines. At the Ministry of Justice he modernised linkage workflows—migrating from Glue to DuckDB and implementing a SQL connected-components algorithm—to accelerate pipelines 2–5x and shrink migration timeframes dramatically. A core contributor to the open-source Splink library, he blends backend development, rigorous testing and benchmarking to make scalable linkage accessible to governments and researchers. Tom pairs an econometrics background with strong software engineering practices (TDD, CI/CD, containerisation), and he actively shares knowledge through talks and community hackathons.
code5 years of coding experience
job6 years of employment as a software developer
bookBSc Economics with Econometrics, Economics, First Class Honours, BSc Economics with Econometrics, Economics, First Class Honours at University of Kent
github-logo-circle

Github Skills (7)

python10
data-science10
testing9
duckdb9
deduplication8
entity-resolution8
sql8

Programming languages (6)

RShellRustJavaScriptHTMLPython

Github contributions (5)

github-logo-circle
Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends
Role in this project:
userBack-end Developer & Data Scientist
Contributions:11 releases, 496 reviews, 446 commits in 11 months
Contributions summary:Tom primarily focused on fixing pathing issues and implementing new functionality within the `splink` library, specifically addressing testing and benchmarking. They made code modifications across various modules, including those related to comparison level testing, linking with different backends, and settings. The commits indicate involvement in both data analysis and core library enhancements.
data-matchingrecord-linkagelinkagescalableduckdb
Fast, accurate and scalable probabilistic data linkage using your choice of SQL backend
Contributions:35 PRs, 73 pushes, 11 branches in 10 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial