Alexandre Défossez is a research-driven AI leader and founding member of Kyutai’s Paris lab, now serving as Chief Science Officer with over a decade of experience building state-of-the-art multi-modal and speech models. He led music generation research at Meta/FAIR, co-developed the widely used AudioCraft framework (including MusicGen and EnCodec) and created Demucs during his PhD, a benchmark source-separation model. His hands-on contributions span backend engineering and model optimization—recently improving transformer internals, KV cache, and RVQ training for Kyutai’s Moshi speech-text foundation model—bridging research and production. Trained in mathematics and physics at École Normale Supérieure with an MVA master, he blends rigorous theory with practical open-source engineering and a persistent focus on reproducible, demo-ready tools. An intriguing thread through his work is applying deep learning to both creative audio generation and brain activity decoding, reflecting a rare mix of artistic and neuroscientific interests.
12 years of coding experience
5 years of employment as a software developer
Doctor of Philosophy - PhD Mathématiques, Doctor of Philosophy - PhD Mathématiques at Ecole normale supérieure
Real Time Speech Enhancement in the Waveform Domain (Interspeech 2020)We provide a PyTorch implementation of the paper Real Time Speech Enhancement in the Waveform Domain. In which, we present a causal speech enhancement model working on the raw waveform that runs in real-time on a laptop CPU. The proposed model is based on an encoder-decoder architecture with skip-connections. It is optimized on both time and frequency domains, using multiple loss functions. Empirical evidence shows that it is capable of removing various kinds of background noise including stationary and non-stationary noises, as well as room reverb. Additionally, we suggest a set of data augmentation techniques applied directly on the raw waveform which further improve model performance and its generalization abilities.
Role in this project:
ML Engineer
Contributions:2 reviews, 55 commits, 11 PRs in 2 years 3 months
Contributions summary:Alexandre primarily contributed to the core functionality and stability of the real-time speech enhancement model. Their work involved refactoring the evaluation process, making it more robust to missing dependencies. They updated dependencies, and introduced support for processing more frames at once, indicating a focus on performance and real-time processing. Furthermore, the user made enhancements related to the pre-trained models and configurations.
State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.
Role in this project:
ML Engineer
Contributions:3 reviews, 20 commits, 4 PRs in 4 months
Contributions summary:Alexandre primarily focused on maintaining and improving the EnCodec neural audio codec, making several version updates and addressing bugs. They implemented a loss balancer and added a warning for training with RVQ. Furthermore, they made changes related to audio conversion and Windows path fixes, indicating a focus on code stability, usability, and addressing platform-specific issues.
audiocodecdeep-learning
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.