Guillem Subies is a Senior Data Scientist with eight years of experience specializing in NLP and language models, currently working in R&D at Instituto de Ingeniería del Conocimiento while pursuing a PhD in Computer Science at UC3M. He holds dual B.Sc. degrees in Mathematics and Computer Science and an M.Sc. in Artificial Intelligence from UPM, with top honors for research on transformers and recurrent language models. A pragmatic Python developer and clean-code advocate, he contributes to major open-source projects such as Hugging Face Transformers (tokenization improvements) and the textstat readability library, focusing on robust preprocessing and performance. His work bridges academic rigor and production readiness, from implementing text normalization using ftfy/spacy to hardening readability algorithms against edge cases. Based in the Madrid metro area, he combines strong theoretical foundations with hands-on engineering to deliver reliable NLP solutions.
8 years of coding experience
Master Universitario en Inteligencia Artificial, 8.68/10, Master Universitario en Inteligencia Artificial, 8.68/10 at Universidad Politécnica de Madrid
Charles III University of Madrid (Universidad Carlos III de Madrid)
Bachillerato, Ciencias, Matrícula de Honor, Bachillerato, Ciencias, Matrícula de Honor at Col·legi Santa María
Grado en Ingeniería Informática, 7,51/10, Grado en Ingeniería Informática, 7,51/10 at Universidad Rey Juan Carlos
:memo: python package to calculate readability statistics of a text object - paragraphs, sentences, articles.
Role in this project:
Back-end Developer
Contributions:1 review, 9 commits, 3 PRs in 1 year 10 months
Contributions summary:Guillem contributed to the textstat Python package, focusing on improving the code's robustness and performance. They updated the `setup.py` file and made several changes to `textstat.py`, primarily implementing caching and addressing potential `ZeroDivisionError` issues within readability calculations. Furthermore, the user updated testing files to ensure proper function. The user demonstrated a strong understanding of Python and the specific readability algorithms used in the project.
🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
Role in this project:
ML Engineer
Contributions:6 commits, 7 PRs, 24 comments in 1 year 2 months
Contributions summary:Guillem primarily contributed to the tokenization components within the `pytorch_transformers` directory. Their changes involved updating `tokenization_xlm.py` and `tokenization_openai.py` files, specifically by integrating `ftfy` and `spacy` libraries for text normalization. Additionally, the user updated documentation related to hyperparameter search functionality in the `Trainer` class. This work indicates a focus on improving model preprocessing and training procedures.
pythonbertspeech-recognitionstate-of-the-artflax
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.