Tanmay Verma is a Senior System Software Engineer with 9 years of experience focused on system-level inference infrastructure and heterogeneous compute, currently contributing to NVIDIA’s Triton Inference Server. He combines deep expertise in computer architecture, embedded systems, and ML deployment—having improved Triton’s Python and C++ backends, refactored perf_client, and enabled better OpenCV4 and TensorFlow compatibility for production inference. A strong believer in first-principles learning, he has a track record of solving tricky performance and compatibility issues across backend, DevOps, and runtime layers. His background spans diagnostics firmware and power testing at Oracle, sensor-emulation for autonomous systems, and FPGA ML work, giving him unusual breadth from silicon-adjacent diagnostics to cloud/edge ML serving. Based in California, he pairs rigorous academic credentials (M.S. Computer Engineering, Texas A&M) with practical open-source contributions to a widely used inference platform.
9 years of coding experience
3 years of employment as a software developer
Masters, Computer Engineering, 3.96/4.00, Masters, Computer Engineering, 3.96/4.00 at Texas A&M University
Bachelor of Technology (B.Tech.), Electronics Engineering, 8.08, Bachelor of Technology (B.Tech.), Electronics Engineering, 8.08 at Indian Institute of Technology (Banaras Hindu University), Varanasi
Indian Certificate of Secondary Education, Science and Computer Applications, 96.4/100, Indian Certificate of Secondary Education, Science and Computer Applications, 96.4/100 at The Modern School
Triton Python, C++ and Java client libraries, and GRPC-generated client examples for go, java and scala.
Role in this project:
Back-end Developer
Contributions:224 reviews, 166 commits, 47 PRs in 3 years 7 months
Contributions summary:Tanmay primarily focused on enhancing the functionality and compatibility of the client library. Their contributions included fixing issues related to OpenCV4 compatibility, implementing complete concurrency printing, and refactoring the performance client. The code changes modified core components like image clients and the performance client, indicating a focus on refining the performance and stability of the client library. This included updates, refactoring and addressing the comments.
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Role in this project:
Back-end Developer
Contributions:1 release, 1148 reviews, 353 commits in 3 years 6 months
Contributions summary:Tanmay's contributions primarily focused on enhancing the functionality and compatibility of the Triton Inference Server. They implemented fixes for OpenCV4 compatibility within the image client, addressed concurrency issues by printing complete information on Ctrl+C, and validated compute capabilities before loading models. Furthermore, the user refactored the perf_client and introduced features, such as allowing GraphDef to dictate model placement and the ability to include the statistics of ensemble models.
nvidia-dockernvidiadeep-learninggpuinference
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.