Jose Dianes is a Principal Data Scientist with 14 years of experience blending computational science, statistical modelling and machine learning to solve problems across life sciences, ambient sensing and real-time simulation. Based in Cambridge, he has worked at the intersection of academia and industry—from PhD research to senior engineering roles at EMBL-EBI and applied data science leadership at Chronomics and Mosaic Therapeutics. He brings hands-on expertise in big-data tooling (notably PySpark) and production ML—demonstrated by open-source Spark notebooks and a MovieLens recommender that bridges research and web deployment. Comfortable leading teams of varying sizes, Jose pairs technical autonomy with practical delivery, often translating complex simulations and bioinformatics challenges into reproducible analytics pipelines. A detail that sets him apart is his long arc from research engineer to principal scientist, giving him deep domain intuition alongside production engineering skills.
14 years of coding experience
15 years of employment as a software developer
Bachelor's degree, MSc, Bachelor's degree, MSc at Universidad de Málaga
Ways of doing Data Science Engineering and Machine Learning in R and Python
Role in this project:
Data Scientist
Contributions:121 commits, 15 PRs, 69 pushes in 5 years 10 months
Contributions summary:Jose appears to be primarily focused on data analysis and model implementation within the repository. The commits showcase the user working on creating and preparing dataframes, likely for a data science project. The user then implemented methods for indexing and data selection, and some initial exploratory data analysis tasks. The user also worked on a Python script for a sentiment analysis and also started a Django app.
Apache Spark & Python (pySpark) tutorials for Big Data Analysis and Machine Learning as IPython / Jupyter notebooks
Role in this project:
Back-end Developer & Data Scientist
Contributions:154 commits, 5 PRs, 146 pushes in 1 year 1 month
Contributions summary:Jose's contributions center around setting up the project structure and creating initial notebooks for Apache Spark and Python (pySpark) tutorials focused on big data analysis and machine learning. The user is establishing a base of knowledge using pySpark. They created an initial set of Jupyter notebooks. The primary focus appears to be setting up the foundation and illustrating basic Spark concepts for big data analysis and machine learning.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.