Elizabeth T

Senior ML Data Engineer

Boston, Massachusetts, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

👤
Senior
🎓
Top School
Elizabeth T is a Senior ML Data Engineer with 12 years of experience blending data engineering, MLOps, and applied research, currently building production ML data pipelines at SimpliSafe. She holds a Ph.D. in Biology and transitioned from postdoctoral work in microbiology and genetics into industry roles that span startups and enterprise teams, bringing strong experimental rigor to data-centric systems. Elizabeth has led data science and MLOps efforts at companies including iRobot and Healthie, and contributed to high-impact open-source work on the Hugging Face datasets hub by adding low-resource Filipino NLP datasets for hate speech, fake news, and entailment tasks. A certified FIDE School chess instructor and long-time freeCodeCamp top contributor, she combines teaching, community engagement, and a knack for making complex ML datasets usable in production.
code12 years of coding experience
job12 years of employment as a software developer
bookDoctor of Philosophy Biology General, Doctor of Philosophy Biology General at Kobe University
github-logo-circle

Github Skills (9)

machine-learning10
nlp10
python10
natural-language-processing10
data-science10
datasets10
pandas8
pytorch6
tensorflow6

Programming languages (4)

ScalaJavaScriptJupyter NotebookPython

Github contributions (5)

github-logo-circle
huggingface/datasets

Dec 2020 - Dec 2020

🤗 The largest hub of ready-to-use datasets for ML models with fast, easy-to-use and efficient data manipulation tools
Role in this project:
userData Scientist & ML Engineer
Contributions:5 commits, 6 PRs, 4 comments in 3 days
Contributions summary:Elizabeth primarily contributed to adding and integrating new datasets focused on Filipino language processing, specifically for tasks like hate speech detection, fake news classification, and sentence entailment. They created dataset scripts, including data loading, feature engineering, and metadata, which is critical for the usage of the datasets in downstream ML tasks. The user's contributions demonstrate a focus on expanding the repository's capabilities to support low-resource language research and the application of NLP to social media data.
ml-modelstensorflownatural-language-processingmanipulationdata-science
Contributions:61 pushes, 2 branches in 2 years 6 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial