Summary
Florian Strub is a research-driven engineering leader with over a decade of experience building and scaling post-training stacks for large language models, currently heading RLVR and post-training engineering at Cohere from Paris. He leads a 15-person European RL team, co-orchestrates company-wide fine-tuning strategies (notably Command A) and harmonized tooling and evaluations to enable weekly model improvements across 7B–100B families. His background spans DeepMind research on multi-agent RL and multimodal self-supervision (co‑author of BYOL, DeepNash work) through a PhD that produced enduring multimodal datasets and methods like FiLM, blending deep scientific impact with production-grade engineering. Florian is as comfortable debugging TPU multi-device issues and refactoring large codebases as he is designing novel RL algorithms (RCfD, CoPG) that stabilized RLHF and enabled multi-reward optimization. He has a track record of growing teams and decentralizing organization design to preserve algorithmic freedom while industrializing R&D. An uncommon strength is his ability to move seamlessly from low-level performance fixes to shaping company-wide training strategy and academic collaborations.
10 years of coding experience
7 years of employment as a software developer
Double Master in Engineering and Management, Double Master in Engineering and Management at École des Mines de Saint-Étienne
Msc in Advanced Computing Mathematics and Computer Science, Msc in Advanced Computing Mathematics and Computer Science at Imperial College London
Bachelor of Applied Science (B.A.Sc.) Mathematics Physics Chemistry, Bachelor of Applied Science (B.A.Sc.) Mathematics Physics Chemistry at Lycee Janson de Sailly
French, Spanish, English