Hemant Jain is a Staff Software Engineer with 11 years of experience building high-performance ML inference and serving systems, currently working on multi-modal AI inference at Meta Superintelligence Labs. He has deep hands-on expertise in model serving stacks (Triton, vLLM, TensorRT) and has driven production migrations and optimizations that scaled Cohere’s platform by 100x–1000x for conversational and embedding models. A contributor to NVIDIA’s widely used Triton Inference Server, Hemant has implemented nuanced features around gRPC/HTTP porting and tensor format handling that improve robustness in cloud and edge deployments. His background blends a UW MS in Data Science and academic research in computer vision with practical backend engineering across Go, C++, and Python. Known for shipping cost-efficient solutions to run thousands of fine-tuned LLMs concurrently, he pairs systems-level performance tuning with an appreciation for model-centric needs. Based in the Bay Area, he brings both deep platform engineering and ML research experience to productionizing cutting-edge AI.
11 years of coding experience
8 years of employment as a software developer
Master of Science - MS Data Science, Master of Science - MS Data Science at University of Washington
Delhi Public School Ruby Park, Kolkata
Bachelor of Technology (B.Tech.) Computer Science and Engineering, Bachelor of Technology (B.Tech.) Computer Science and Engineering at Vellore Institute of Technology
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Role in this project:
Back-end Developer
Contributions:1047 reviews, 412 commits, 488 PRs in 3 years 2 months
Contributions summary:Hemant's commits focus on implementing features for the Triton Inference Server, specifically related to the configuration and handling of different gRPC and HTTP API ports for status, health, and other API endpoints. This includes modifying the server's main file to allow for unique ports for HTTP and gRPC services, addressing potential port collisions, and adding support for different input and output tensor formats. These changes enhance the server's capabilities and expand its configuration options.
Contributions:28 reviews, 21 commits, 30 PRs in 1 year 7 months
pythonbackendtensorflowtritontensorflow-2
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.