Romain Beaumont is an experienced engineering leader based in Paris with 13 years blending technical management and hands-on software development. As Responsable de Groupe Formation et Médiatisation at LGM he combines team leadership with media and training responsibilities, having previously led médiatisation efforts since 2013. He is an active open-source contributor across ML and Node.js ecosystems, with notable work on high-traffic projects like PrismarineJS (Minecraft protocol, bots and data), efficient image dataset tooling (img2dataset) and CLIP/DALLE PyTorch implementations. His contributions span backend protocol parsing, front-end web client features, CI/DevOps for large ML repos, and scalable data processing—demonstrating a rare mix of production ML engineering and low-level systems work. Colleagues value his practical problem-solving: he automates tedious data tasks (wiki extraction, schema validation) and optimizes training/build pipelines for measurable performance gains. Comfortable bridging creative media workflows and rigorous engineering, he brings multidisciplinary experience from technical illustration to machine learning infrastructure.
13 years of coding experience
BEP / Bac Pro MVA Technologie / technicien de mécanique automobile, BEP / Bac Pro MVA Technologie / technicien de mécanique automobile at CFA Jean Claude Andrieu
Language independent module providing minecraft data for minecraft clients, servers and libraries.
Role in this project:
Back-end Developer
Contributions:131 reviews, 824 commits, 505 PRs in 7 years 11 months
Contributions summary:Romain primarily contributed to the development of the `minecraft-data` module. Their work focused on enhancing the data structure and validation processes. This involved adding and improving JSON schemas, which automatically validated the data, and fixing data inconsistencies like stack sizes. The user also implemented wiki extraction scripts to automatically retrieve and parse information from the Minecraft wiki.
Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
Role in this project:
Back-end Developer
Contributions:47 releases, 52 reviews, 262 commits in 1 year 5 months
Contributions summary:Romain implemented the core functionality of the `img2dataset` project, including the image downloading and resizing processes. The initial commit introduces the primary Python script (`img2dataset.py`) which defines image processing functions. The user also added features for resizing images using different modes, and incorporated the use of threading to download images. They also contributed to setup files and notebook examples to guide the user.
image-urldeep-learningdatasetbig-dataimage
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.