Guillem Subies

Senior Data Scientist

Greater Madrid Metropolitan Area Spain
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

👤
Senior
🎓
Top School
Guillem Subies is a Senior Data Scientist with eight years of experience specializing in NLP and language models, currently working in R&D at Instituto de Ingeniería del Conocimiento while pursuing a PhD in Computer Science at UC3M. He holds dual B.Sc. degrees in Mathematics and Computer Science and an M.Sc. in Artificial Intelligence from UPM, with top honors for research on transformers and recurrent language models. A pragmatic Python developer and clean-code advocate, he contributes to major open-source projects such as Hugging Face Transformers (tokenization improvements) and the textstat readability library, focusing on robust preprocessing and performance. His work bridges academic rigor and production readiness, from implementing text normalization using ftfy/spacy to hardening readability algorithms against edge cases. Based in the Madrid metro area, he combines strong theoretical foundations with hands-on engineering to deliver reliable NLP solutions.
code8 years of coding experience
bookMaster Universitario en Inteligencia Artificial, 8.68/10, Master Universitario en Inteligencia Artificial, 8.68/10 at Universidad Politécnica de Madrid
bookCharles III University of Madrid (Universidad Carlos III de Madrid)
bookBachillerato, Ciencias, Matrícula de Honor, Bachillerato, Ciencias, Matrícula de Honor at Col·legi Santa María
bookGrado en Ingeniería Informática, 7,51/10, Grado en Ingeniería Informática, 7,51/10 at Universidad Rey Juan Carlos
languagesSpanish, Catalan, English
stackoverflow-logo

Stackoverflow

Stats
154reputation
41kreached
4answers
6questions
github-logo-circle

Github Skills (17)

transformers10
pytorch10
language-model10
python10
readability10
nlp10
pre-trained-model9
caching9
machine-learning9
testing8
keras6
deep-learning6
tensorflow6
lemmatization6
numpy6

Programming languages (4)

C++GoJupyter NotebookPython

Github contributions (5)

github-logo-circle
textstat/textstat

Aug 2019 - Jun 2021

:memo: python package to calculate readability statistics of a text object - paragraphs, sentences, articles.
Role in this project:
userBack-end Developer
Contributions:1 review, 9 commits, 3 PRs in 1 year 10 months
Contributions summary:Guillem contributed to the textstat Python package, focusing on improving the code's robustness and performance. They updated the `setup.py` file and made several changes to `textstat.py`, primarily implementing caching and addressing potential `ZeroDivisionError` issues within readability calculations. Furthermore, the user updated testing files to ensure proper function. The user demonstrated a strong understanding of Python and the specific readability algorithms used in the project.
nlpstatisticspythonsentencessmog
huggingface/transformers

Aug 2019 - Nov 2020

🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
Role in this project:
userML Engineer
Contributions:6 commits, 7 PRs, 24 comments in 1 year 2 months
Contributions summary:Guillem primarily contributed to the tokenization components within the `pytorch_transformers` directory. Their changes involved updating `tokenization_xlm.py` and `tokenization_openai.py` files, specifically by integrating `ftfy` and `spacy` libraries for text normalization. Additionally, the user updated documentation related to hyperparameter search functionality in the `Trainer` class. This work indicates a focus on improving model preprocessing and training procedures.
pythonbertspeech-recognitionstate-of-the-artflax
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial