Michael Carilli

Senior Applied Cryptography Engineer at Matter Labs

Albuquerque, New Mexico, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Michael Carilli is a Senior Applied Cryptography Engineer with 11 years of high-performance computing and GPU expertise, currently building GPU-accelerated zero-knowledge provers for zkSync at Matter Labs to help Ethereum scale securely. He previously spent five years on the PyTorch team at NVIDIA, improving mixed-precision, multi-GPU training and contributing CUDA kernels and distributed training optimizations to the widely used apex repository. His background as a computational scientist and PhD-trained physicist informs a knack for squeezing large speedups from heterogeneous hardware (32x in one CFD project) and adapting numerical methods across GPUs, Xeon Phis, and multicore CPUs. Comfortable moving between low-level CUDA kernels, C++/Fortran interoperability, and cryptographic primitives like multi-scalar multiplication and NTTs, he blends rigorous academic training with practical production engineering. Based in Albuquerque, he pairs deep performance tuning chops with an uncommon focus on cryptography-for-scale, making him effective at turning compute-bound research into deployable blockchain infrastructure.
code11 years of coding experience
job7 years of employment as a software developer
bookBS, Physics, 4.0 GPA in major, BS, Physics, 4.0 GPA in major at University of Notre Dame
bookDoctor of Philosophy (PhD), Physics, 4.0 GPA, Doctor of Philosophy (PhD), Physics, 4.0 GPA at University of California, Santa Barbara
languagesEnglish, German, Spanish
stackoverflow-logo

Stackoverflow

Stats
381reputation
11kreached
0answers
14questions
github-logo-circle

Github Skills (14)

multiprecision10
cuda10
pytorch10
deep-q-learning10
deep-learning10
distributed-training10
optimization10
machine-learning9
kernel9
caching6
cpython6
avx6
filesystem6
python-c-api6

Programming languages (5)

C++RustJupyter NotebookPythonCuda

Github contributions (5)

github-logo-circle
NVIDIA/apex

May 2018 - Aug 2020

A PyTorch Extension: Tools for easy mixed precision and distributed training in Pytorch
Role in this project:
userML Engineer
Contributions:2 reviews, 94 commits, 165 PRs in 2 years 3 months
Contributions summary:Michael primarily contributed to the `apex` repository, a PyTorch extension for mixed precision and distributed training. Their contributions include modifications to the `DistributedDataParallel` module to improve parameter handling, and efficient bucketing and allreduce operations for optimized distributed training. Furthermore, the user made changes to the fused Adam optimizer and related CUDA kernels, indicating expertise in optimizing deep learning training.
distributed-trainingpytorch
mcarilli/pytorch

Oct 2018 - Jun 2022

Tensors and Dynamic neural networks in Python with strong GPU acceleration
Contributions:537 pushes, 115 branches in 3 years 8 months
pythongpu-accelerationdeep-learninggpuacceleration
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial