Summary
Lang Zhao is a Senior Software Engineer with 11 years of experience specializing in turning cutting-edge LLMs into production-grade, low-latency inference systems. Based in San Jose, he leads GPU- and system-level performance work—optimizing KV cache reuse, quantization, and continuous batching with TensorRT-LLM and Triton to deliver high throughput under strict SLAs. He designs multi-tenant serving architectures with intelligent load balancing and resource-aware scheduling, and has a background in compiler-level MLIR optimizations for graph fusion and execution efficiency. Lang moves quickly to evaluate new inference paradigms and emerging GPU hardware, blending deep performance tuning with pragmatic system design. Despite a modest GitHub bio, his track record spans enterprise AI at Galileo (now part of Cisco), applied ML at ByteDance, and research-rooted work from Purdue, reflecting both production impact and academic rigor.
11 years of coding experience
4 years of employment as a software developer
Bachelor's degree, Flying Vehicle Power Engineering, Bachelor's degree, Flying Vehicle Power Engineering at Beihang University
Master of Science - MS, Machine Learning, Master of Science - MS, Machine Learning at Purdue University - Office of the Vice Provost for Graduate Students and Postdoctoral Scholars