Pierre Marcenac is a research engineer based in Zurich with 10 years of experience building data-centric ML systems and production-ready engineering at companies from startups to Google DeepMind. He has led engineering and data teams (as first employee and lead at Kili Technology) to help organizations move AI projects from prototype to production, and later worked as a senior software engineer at Google before joining DeepMind. His pragmatic blend of back-end engineering and data science is reflected in open-source contributions to tensorflow/datasets—fixing parsing bugs and integrating Hugging Face datasets to make large-scale datasets easier to consume. Trained in engineering and machine learning at CentraleSupélec and TU Berlin, he pairs rigorous academic foundations with hands-on product delivery across annotation platforms and research code. Colleagues would note his knack for improving data reliability in complex pipelines, a detail that repeatedly surfaces across his roles.
10 years of coding experience
6 years of employment as a software developer
MS General Engineering Computer Science, MS General Engineering Computer Science at CentraleSupélec
MS Industrial Engineering Machine Learning, MS Industrial Engineering Machine Learning at Technische Universität Berlin
BS Mathematics and Physics, BS Mathematics and Physics at Lycée Janson-de-Sailly (Paris)
TFDS is a collection of datasets ready to use with TensorFlow, Jax, ...
Role in this project:
Back-end Developer & Data Scientist
Contributions:1 release, 5 reviews, 25 commits in 3 months
Contributions summary:Pierre primarily contributed to the "tensorflow/datasets" repository by fixing data parsing issues within the "criteo" dataset, specifically addressing the incorrect parsing of critical fields. They also integrated Hugging Face datasets into TFDS, enabling users to download and prepare those datasets within the TFDS framework, showcasing proficiency in data integration. Furthermore, the user's work involved addressing the correct handling of data types, ensuring accurate parsing of boolean values.
Croissant is a high-level format for machine learning datasets that brings together four rich layers.
Contributions:7 releases, 519 reviews, 508 PRs in 1 year 11 months
datasetsjson-ldmachine-learningschema-org
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.