Luca Foppiano is a Text and Data Mining engineer and founder with 12 years of experience building ML-driven document and knowledge extraction systems across research institutes, public organizations, and startups. He has led production integrations of Grobid at scale for the European Patent Office, contributed back-end improvements to the popular Grobid project, and architected large patent and scientific text pipelines using document databases and RAG approaches. His recent work spans LLM evaluation and fine-tuning for materials informatics, automatic construction of materials-property databases, and practical deployments of TDM in both research and industry settings. Comfortable bridging research and engineering, he has delivered seminars and taught TDM/ML topics while running R&D projects at Inria and the National Institute for Materials Science. Based in Aveiro, Portugal, he blends deep academic training (PhD-level work) with hands-on engineering—an unusual mix that fuels both reproducible research and production-ready tooling. Outside core development, he has a history of open-source experimentation (e.g., Grobid contributions) and a pattern of turning research prototypes into scalable systems.
12 years of coding experience
16 years of employment as a software developer
Vienna University of Technology
Master Computer Engineering, Master Computer Engineering at Università di Pavia
Doctor of Philosophy - PhD Computer Science, Doctor of Philosophy - PhD Computer Science at University of Tsukuba
Communication General, Communication General at Toastmaster International
A machine learning software for extracting information from scholarly documents
Role in this project:
Back-end Developer
Contributions:4 releases, 36 reviews, 687 commits in 6 years 10 months
Contributions summary:Luca contributed to the Grobid project by adding UNIT model names, tagging libraries, and creating a model directory if it doesn't exist during trainer execution. They modified the core Java files, including the GrobidModels and TaggingLabel, indicating work related to model definition and annotation. Further contributions include improving feedback for backend errors by adding error codes for specific failures.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.