Summary
Abhinav Venkataraman is a machine learning engineer with 12 years of experience, currently driving inference and optimization for foundation models at Amazon Special Projects in Seattle. He specializes in 0-to-1 product development for LLM pipelines, having co-invented Amazon’s first cross-model prompt translation and optimization framework and led infrastructure work that cut manual adaptation time from 60 days to 3 and boosted catalog accuracy by 10%. His work blends research and production: building iterative evaluation and prompt-refinement pipelines, developing benchmark datasets, and experimenting with advanced optimization techniques like lightweight RL and genetic Pareto methods. He’s delivered large-scale inference systems—optimizing LLMs on multi-GPU P5 instances, achieving 7x latency reductions and ~90% GPU utilization while saving ~40% compute cost—and built high-throughput transformer deployments on AWS Inferentia and EKS. Comfortable in ambiguous, multidisciplinary settings, he pairs deep algorithmic roots from academic research with pragmatic engineering, having also taught algorithms at the University of Florida. Notably, he has driven adoption of Kubernetes for hardware-accelerated ML at Amazon and mentors teams to reduce model deployment and experimentation cycles from weeks to hours.
13 years of coding experience
Bachelor of Technology (B.Tech.), Computer Science and Engineering, Bachelor of Technology (B.Tech.), Computer Science and Engineering at Shanmugha Arts, Science, Technology and Research Academy
P.S.Senior Secondary School
Master of Science (M.S.), Computer Science, Master of Science (M.S.), Computer Science at University of Florida
Tamil, English, संस्कृतम्(sanskrit), español (spanish)