Summary
Calvin Fei is a Machine Learning Engineer II at AWS Annapurna Labs with nine years of experience building and optimizing large-scale LLM inference systems for specialized accelerators like Trainium and Inferentia. He combines a Cornell CS master’s background with hands-on kernel and accelerator work—CUDA, Triton, C/C++—to squeeze latency, throughput, and memory efficiency out of production workloads. Calvin partners with strategic customers to port PyTorch models, prototype benchmarking pipelines, and collaborate across compiler and hardware teams to diagnose real-world bottlenecks. His expertise spans the full LLM stack from KV-cache and speculative decoding strategies to tensor-core-aware kernel tuning and distributed inference orchestration. Prior to AWS he led LLM research and agentic workflow efforts at InfiniteScene.ai, giving him a rare mix of model-research intuition and systems-first optimization skills. Based in Cupertino, he’s quietly focused on bridging model architectures, ML systems software, and next-generation AI compute platforms to make high-performance inference practical at scale.
9 years of coding experience
Master's degree, Computer Science, Master's degree, Computer Science at Cornell University
Chinese, English