Summary
Michal Valko is a founding researcher and tenured scientist who builds algorithms to minimize human supervision, spanning deep RL, bandits, self-supervised learning and scalable alignment for large language models. With a PhD from Pittsburgh and an HdR in bandit theory, he co-created DeepMind Paris, led BYOL and world-model work at DeepMind, and served as principal Llama scientist at Meta on Llama 3 post-training and online RL stacks. His research blends theoretical guarantees with practical systems—examples include novel Monte-Carlo tree search methods, online kernel and graph sparsification results, and scalable fine-tuning techniques for LMMs. He teaches advanced graph-based ML at ENS Paris-Saclay and now applies decades of sequential and representation-learning expertise to steer a stealth AI startup out of London. An uncommon thread in his career is repeatedly closing the loop between provable algorithms and large-scale production ML, from hospital anomaly detection to state-of-the-art model alignment.
8 years of coding experience
14 years of employment as a software developer
Habilitation à Diriger des Recherches (HdR), Bandit theory, Habilitation à Diriger des Recherches (HdR), Bandit theory at ENS Paris-Saclay
MSc., Computer Science, MSc., Computer Science at Univerzita Komenského v Bratislave
Erasmus Exchange, Machine Learning, Erasmus Exchange, Machine Learning at Universidade Nova de Lisboa
Gymnasium Diploma, Mathematics, Gymnasium Diploma, Mathematics at Alejová High School
PhD., Machine Learning, PhD., Machine Learning at University of Pittsburgh
Visiting Student (Cross-registered from University of Pittsburgh), Machine Learning, Computer Science, Visiting Student (Cross-registered from University of Pittsburgh), Machine Learning, Computer Science at Carnegie Mellon University
English, French, Slovak, Czech, Portuguese