Hanming Lu is an AI research scientist with five years of experience specializing in inference and LLM performance engineering across startups and large tech, currently working on inference at Meta. He previously contributed to xAI as Member of Technical Staff and built CUDA kernels and performance optimizations for large language models at Anyscale, bridging research and production needs. Trained at UC Berkeley (MS/PhD) with foundational work from the University of Waterloo, he combines rigorous academic research with hands-on systems engineering. Hanming’s strengths are squeezing latency and cost out of model inference pipelines and translating novel model ideas into deployable, high-performance code. Colleagues appreciate that he moves comfortably between low-level GPU work and higher-level model evaluation, making him effective at both prototyping and production hardening. Based in San Francisco, he brings a practical researcher’s mindset to real-world ML systems challenges.
5 years of coding experience
3 years of employment as a software developer
Doctor of Philosophy - PhD, Computer Science, Doctor of Philosophy - PhD, Computer Science at University of California, Berkeley
Bachelor’s Degree, Computer Science, Bachelor’s Degree, Computer Science at University of Waterloo
An open source framework that provides a simple, universal API for building distributed applications. Ray is packaged with RLlib, a scalable reinforcement learning library, and Tune, a scalable hyperparameter tuning library.
Contributions:18 pushes, 5 branches in 1 year 7 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.