Fadi Arafeh

Graduate Software Engineer - ML Libraries

London, England, United Kingdom
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Fadi Arafeh is a Graduate Software Engineer at Arm with three years of hands-on experience optimizing ML libraries for AArch64 architectures. He specializes in accelerating quantized neural operations, having contributed upstream to PyTorch by integrating and leveraging the Arm Compute Library for faster Qint8/QUint8 paths. Based in London, he brings practical firmware-to-framework insight from internships and research roles at Arm, The University of Manchester, and Rapita Systems. His background in AI (first-class honours) and repository-mining tooling shows a blend of applied research and production software engineering. Colleagues know him for pragmatic performance work that makes “ML models go brrrrr” in real hardware environments.
code3 years of coding experience
job3 years of employment as a software developer
bookBachelor's degree, Artificial Intelligence, 86.8, Bachelor's degree, Artificial Intelligence, 86.8 at The University of Manchester
github-logo-circle

Github Skills (10)

quantization10
arm10
pytorch10
machine-learning10
deep-learning10
gpu9
cprogramming-language9
c-language9
tensor9
autograd7

Programming languages (3)

C++CPython

Github contributions (5)

github-logo-circle
pytorch/pytorch

Aug 2024 - Apr 2025

Tensors and Dynamic neural networks in Python with strong GPU acceleration
Role in this project:
userML Engineer
Contributions:41 reviews, 18 PRs, 38 pushes in 7 months
Contributions summary:Fadi contributed to the optimization and direct integration of the Arm Compute Library (ACL) within the PyTorch framework, specifically targeting AArch64 architecture. Their work focused on enhancing performance through direct ACL usage in quantized linear operations, including static and dynamic quantization paths. They addressed performance bottlenecks in existing implementations by enabling the direct use of ACL for fast quantized operations and enabling support for Qint8 and QUint8 add operations. This work included updating the ACL version and incorporating new ACL features within the PyTorch codebase.
pythongpu-accelerationdeep-learninggpunumpy
fadara01/oneDNN

Nov 2022 - Jan 2025

oneAPI Deep Neural Network Library (oneDNN)
Contributions:6 pushes, 13 branches in 2 years 2 months
caffe2deep-learningdeep-neural-networkoneapineural-network
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial