Mohit Sharma

Machine Learning Engineer at Hugging Face, Inc.

email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
Mohit Sharma is a Machine Learning Engineer with nine years of experience building and productionizing deep learning systems across CV, NLP, LLMs, Graph ML and GenAI, now working on Search and Content AI at Adobe. He blends hands-on model development in PyTorch and Transformers with low-level optimization and runtime engineering in C++—notably contributing ONNX export and ONNX Runtime support to Hugging Face transformers and optimum. Mohit has a strong track record of accelerating inference on cloud and edge (TensorRT, ROCm, custom kernels, FP8) and has implemented execution providers and graph-level tooling at Qualcomm. Comfortable across the ML lifecycle, he also designs reusable libraries and pipelines for reproducible training and low-latency deployment. Colleagues know him for bridging research-grade models with pragmatic engineering constraints to deliver scalable inference in real-world systems.
code9 years of coding experience
github-logo-circle

Github Skills (28)

transformers10
pytorch10
system-configuration10
python10
attention-mechanism10
machine-learning10
inference10
onnx10
llm10
kernel10
custom-configuration10
roc10
deep-learning10
onnxruntime10
performance-optimization10

Programming languages (5)

C++RustJupyter NotebookPythonCuda

Github contributions (5)

github-logo-circle
huggingface/optimum

Nov 2022 - Jan 2023

🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Role in this project:
userMLOps Engineer
Contributions:180 reviews, 9 commits, 60 PRs in 2 months
Contributions summary:Mohit primarily contributed to integrating and supporting ONNX Runtime (ORT) within the `optimum` library, focusing on accelerating Transformers and other machine learning models. Their work involved adding support for specific models like Whisper, integrating IO binding, and configuring export processes to facilitate efficient inference using ORT. Furthermore, the user demonstrated proficiency in modifying model configurations, updating tests, and refining the export mechanisms for various encoder-decoder architectures.
diffusershardwareinferencesentence-transformerstransformers
Large Language Model Text Generation Inference
Role in this project:
userMLOps Engineer
Contributions:66 reviews, 28 PRs, 86 pushes in 1 year 1 month
Contributions summary:Mohit focused on improving ROCm (AMD's implementation of CUDA) support for the text generation inference framework. They implemented and optimized custom ROCm kernels for attention mechanisms, including FP8 (float8) support. The user's contributions included updating existing kernels from VLLM and adding new kernels for flash decoding and other optimizations. They also integrated various kernel repositories and improved model support for ROCm.
inferencelarge-language-modelstext-generationbloomnlp
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial