Alexandre Défossez

Chief Science Officer at Kyutai

Paris, Ile-de-France
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Alexandre Défossez is a research-driven AI leader and founding member of Kyutai’s Paris lab, now serving as Chief Science Officer with over a decade of experience building state-of-the-art multi-modal and speech models. He led music generation research at Meta/FAIR, co-developed the widely used AudioCraft framework (including MusicGen and EnCodec) and created Demucs during his PhD, a benchmark source-separation model. His hands-on contributions span backend engineering and model optimization—recently improving transformer internals, KV cache, and RVQ training for Kyutai’s Moshi speech-text foundation model—bridging research and production. Trained in mathematics and physics at École Normale Supérieure with an MVA master, he blends rigorous theory with practical open-source engineering and a persistent focus on reproducible, demo-ready tools. An intriguing thread through his work is applying deep learning to both creative audio generation and brain activity decoding, reflecting a rare mix of artistic and neuroscientific interests.
code12 years of coding experience
job5 years of employment as a software developer
bookDoctor of Philosophy - PhD Mathématiques, Doctor of Philosophy - PhD Mathématiques at Ecole normale supérieure
bookM2 MVA Mathématiques appliquées, M2 MVA Mathématiques appliquées at ENS Cachan
bookMathématiques, Mathématiques at Lycée privé sainte geneviève
github-logo-circle

Github Skills (31)

music-generation10
pytorch10
python10
speech-enhancement10
machine-learning10
audio-processing10
vector10
transformer-models10
gradio10
cuda10
quantization10
modeling10
ffmpeg9
model-driven9
continuous-deployment9

Programming languages (8)

C#CRustMakefileTeXHTMLJupyter NotebookPython

Github contributions (5)

github-logo-circle
facebookresearch/denoiser

Sep 2020 - Dec 2022

Real Time Speech Enhancement in the Waveform Domain (Interspeech 2020)We provide a PyTorch implementation of the paper Real Time Speech Enhancement in the Waveform Domain. In which, we present a causal speech enhancement model working on the raw waveform that runs in real-time on a laptop CPU. The proposed model is based on an encoder-decoder architecture with skip-connections. It is optimized on both time and frequency domains, using multiple loss functions. Empirical evidence shows that it is capable of removing various kinds of background noise including stationary and non-stationary noises, as well as room reverb. Additionally, we suggest a set of data augmentation techniques applied directly on the raw waveform which further improve model performance and its generalization abilities.
Role in this project:
userML Engineer
Contributions:2 reviews, 55 commits, 11 PRs in 2 years 3 months
Contributions summary:Alexandre primarily contributed to the core functionality and stability of the real-time speech enhancement model. Their work involved refactoring the evaluation process, making it more robust to missing dependencies. They updated dependencies, and introduced support for processing more frames at once, indicating a focus on performance and real-time processing. Furthermore, the user made enhancements related to the pre-trained models and configurations.
data-augmentationencoder-decoderfrequency-domainmultiple-loss-functionpytorch
facebookresearch/encodec

Oct 2022 - Mar 2023

State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.
Role in this project:
userML Engineer
Contributions:3 reviews, 20 commits, 4 PRs in 4 months
Contributions summary:Alexandre primarily focused on maintaining and improving the EnCodec neural audio codec, making several version updates and addressing bugs. They implemented a loss balancer and added a warning for training with RVQ. Furthermore, they made changes related to audio conversion and Windows path fixes, indicating a focus on code stability, usability, and addressing platform-specific issues.
audiocodecdeep-learning
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial