Dilawar Mahmood is a Founding Research Engineer with eight years of experience building high-performance ML training and inference systems, currently focused on GPU kernels, scheduling, and distributed training at ZeroEntropy in San Francisco. He blends a strong math and competitive programming background with hands-on work optimizing throughput vs. goodput, end-to-end debugging of training runs, and synthetic data pipelines. Previously at Apple he led large-scale pre-training data research and reliable batch inference across GPU clusters, and he has a track record of shipping production ML systems as an early engineer and founder. Comfortable with low-level GPU work (CUDA, Triton) as well as functional programming and theorem proving, he brings both systems-level rigor and experimental research instincts. Notably, he’s applied production-grade engineering to novel domains—from implementing AlphaZero in C++ to building full-stack ML apps—showing a bias for shipping complex, high-performance systems.
Contributions:34 commits, 5 PRs, 30 pushes in 6 months
pythondjangoface-recognitionrecognitionfacial
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.