Hemant Jain

Staff Software Engineer at Meta

San Francisco Bay Area United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Hemant Jain is a Staff Software Engineer with 11 years of experience building high-performance ML inference and serving systems, currently working on multi-modal AI inference at Meta Superintelligence Labs. He has deep hands-on expertise in model serving stacks (Triton, vLLM, TensorRT) and has driven production migrations and optimizations that scaled Cohere’s platform by 100x–1000x for conversational and embedding models. A contributor to NVIDIA’s widely used Triton Inference Server, Hemant has implemented nuanced features around gRPC/HTTP porting and tensor format handling that improve robustness in cloud and edge deployments. His background blends a UW MS in Data Science and academic research in computer vision with practical backend engineering across Go, C++, and Python. Known for shipping cost-efficient solutions to run thousands of fine-tuned LLMs concurrently, he pairs systems-level performance tuning with an appreciation for model-centric needs. Based in the Bay Area, he brings both deep platform engineering and ML research experience to productionizing cutting-edge AI.
code11 years of coding experience
job8 years of employment as a software developer
bookMaster of Science - MS Data Science, Master of Science - MS Data Science at University of Washington
bookDelhi Public School Ruby Park, Kolkata
bookBachelor of Technology (B.Tech.) Computer Science and Engineering, Bachelor of Technology (B.Tech.) Computer Science and Engineering at Vellore Institute of Technology
bookSt. James School, Kolkata
languagesBengali, Hindi, English
github-logo-circle

Github Skills (7)

http10
c-language10
cprogramming-language10
grpc10
api-design9
inference9
multithreading8

Programming languages (10)

TypeScriptJavaC++CJavaScriptGoPHPHTML

Github contributions (5)

github-logo-circle
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Role in this project:
userBack-end Developer
Contributions:1047 reviews, 412 commits, 488 PRs in 3 years 2 months
Contributions summary:Hemant's commits focus on implementing features for the Triton Inference Server, specifically related to the configuration and handling of different gRPC and HTTP API ports for status, health, and other API endpoints. This includes modifying the server's main file to allow for unique ports for HTTP and gRPC services, addressing potential port collisions, and adding support for different input and output tensor formats. These changes enhance the server's capabilities and expand its configuration options.
nvidia-dockernvidiadeep-learninggpuinference
The Triton backend for TensorFlow.
Contributions:28 reviews, 21 commits, 30 PRs in 1 year 7 months
pythonbackendtensorflowtritontensorflow-2
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial