Hannes Hapke is an experienced ML leader and open source advocate with 12+ years building production-grade AI systems and developer tooling from startup to enterprise. Currently Director of Open Source at Dataiku and a Google Developer Expert, he has authored O'Reilly and Manning books and contributed to the widely used "Building Machine Learning Pipelines" codebase. He architected end-to-end ML infrastructure at Digits—delivering low-latency similarity models, a 98% accurate invoice extraction pipeline, and internal LLM hosting—and has deep expertise deploying Transformer models with TFX and TensorFlow Serving. A hands-on engineer who moves research to production, Hannes also improves web stacks (Django mixins) and workshop materials for TensorFlow, showing breadth across backend, MLOps, and NLP. Based in Portland, he combines academic rigor (MS in Electrical Engineering) with startup grit, and quietly influences the ecosystem through books, talks, and practical open-source contributions.
12 years of coding experience
15 years of employment as a software developer
Master of Science, Electrical Engineering, Master of Science, Electrical Engineering at Oregon State University
Electrical Engineering, Electrical Engineering at University of Stuttgart
Code repository for the O'Reilly publication "Building Machine Learning Pipelines" by Hannes Hapke & Catherine Nelson
Role in this project:
Data Scientist
Contributions:2 releases, 3 reviews, 110 commits in 1 year 10 months
Contributions summary:Hannes contributed to the development of a machine learning pipeline, as evidenced by the addition of utility scripts for data splitting, and the creation of a Keras-based model experiment notebook. The primary focus of the work appears to be around data ingestion, preprocessing and model building. The user also focused on visualizing the model.
Contributions:19 commits, 5 PRs, 10 comments in 8 months
Contributions summary:Hannes primarily contributed to a TFX pipeline designed for processing and training sentiment analysis models using BERT. Their work involved updating the notebook to include the setup of an ALBERT model and integrating it with the existing pipeline. The user also focused on correcting and refining the model architecture and data preparation steps, including handling the input data structure for the ALBERT model. Furthermore, they adjusted pipeline configurations and package installations to ensure compatibility with the updated TF and TFX versions.
javaevents
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.