Félix Marty is a Senior Software Engineer with 10 years of experience specializing in deep learning deployment, model optimization, and hardware-aware MLOps. Based in Paris, he has driven production-ready enhancements at Hugging Face and now works at AMD, contributing notably to Hugging Face projects like transformers, accelerate, optimum and text-generation-inference to expand GPU support (ROCm/AMD MI2xx/MI3xx), ONNXRuntime integrations, quantization and GPTQ on alternative hardware. He blends strong mathematical training from ENS Paris-Saclay and MINES ParisTech with hands-on systems work, often fixing subtle device and control-flow bugs to make large models export and run reliably. Félix’s contributions reveal an unusual depth in bridging model research and low-level GPU kernels, making him effective at squeezing performance out of cutting-edge accelerators.
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Role in this project:
Back-end Developer
Contributions:21 releases, 1081 reviews, 109 commits in 9 months
Contributions summary:Félix significantly contributed to the `optimum` repository, focusing on enhancing ONNX Runtime integration. The user added the ability to specify ONNX runtime execution providers and incorporated support for custom input shapes. Their contributions included implementing features related to quantization and model comparisons. The user's work streamlined the utilization of ONNX Runtime with the goal of optimizing the performance and efficiency of Transformers and other models.
Contributions:55 reviews, 31 PRs, 153 pushes in 1 year 3 months
Contributions summary:Félix primarily contributed to the project by adding support for AMD Instinct MI210 & MI250 GPUs, including ROCm integration and support for custom kernels. They also implemented GPTQ support on ROCm, improving the model's efficiency. Furthermore, the user made several improvements to the code base and added GPU support to address performance challenges and broaden hardware compatibility. The user also contributed to the implementation of MI300 compatibility including support for PyTorch TunableOp and flash attention.
nlppytorchlanguage-modelbloombert
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.