Vamsi Sripathi is a performance-focused software engineer with 15 years of experience in x86 code optimization and thread parallelism, currently a Member of Technical Staff at Microsoft based in Mountain View. He spent over a decade at Intel driving HPC and ML optimizations—delivering measurable speedups across production workloads (e.g., 1.4x for HotQCD, 1.25x for MPAS-A, and 1.3x for an ECMWF weather app) and contributing to silicon design wins. His work spans low-level kernel development (MKL/BLAS), AVX/AVX512 and VNNI micro-optimizations, memory bandwidth analysis including HBM issues, and tooling such as TACKLE for thread/NUMA tuning. Vamsi has also upstreamed profiling features and MKL-DNN improvements and optimized TensorFlow Serving to leverage Intel MKL for better CPU ML performance. He combines hands-on assembly-level tuning with pragmatic engineering to unlock performance on modern CPUs and accelerators, and has a strong background in scaling scientific codes on supercomputers from his PhD-era research.
9 years of coding experience
12 years of employment as a software developer
M.S. Computer Science, M.S. Computer Science at North Carolina State University
A flexible, high-performance serving system for machine learning models
Role in this project:
ML Engineer
Contributions:5 commits, 1 PR, 2 comments in 5 months
Contributions summary:Vamsi focused on optimizing the TensorFlow serving system for improved CPU performance, specifically by enabling the use of Intel MKL. Their commits modified the `tools/bazel.rc` file to include MKL build configurations and rearranged existing options. They also added an option for an open-source MKL-DNN build. Finally, they merged updates from the upstream master branch.
Contributions:22 commits, 23 pushes, 1 branch in 1 month
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.