Mohit Sharma is a Machine Learning Engineer with nine years of experience building and productionizing deep learning systems across CV, NLP, LLMs, Graph ML and GenAI, now working on Search and Content AI at Adobe. He blends hands-on model development in PyTorch and Transformers with low-level optimization and runtime engineering in C++—notably contributing ONNX export and ONNX Runtime support to Hugging Face transformers and optimum. Mohit has a strong track record of accelerating inference on cloud and edge (TensorRT, ROCm, custom kernels, FP8) and has implemented execution providers and graph-level tooling at Qualcomm. Comfortable across the ML lifecycle, he also designs reusable libraries and pipelines for reproducible training and low-latency deployment. Colleagues know him for bridging research-grade models with pragmatic engineering constraints to deliver scalable inference in real-world systems.
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Role in this project:
MLOps Engineer
Contributions:180 reviews, 9 commits, 60 PRs in 2 months
Contributions summary:Mohit primarily contributed to integrating and supporting ONNX Runtime (ORT) within the `optimum` library, focusing on accelerating Transformers and other machine learning models. Their work involved adding support for specific models like Whisper, integrating IO binding, and configuring export processes to facilitate efficient inference using ORT. Furthermore, the user demonstrated proficiency in modifying model configurations, updating tests, and refining the export mechanisms for various encoder-decoder architectures.
Contributions:66 reviews, 28 PRs, 86 pushes in 1 year 1 month
Contributions summary:Mohit focused on improving ROCm (AMD's implementation of CUDA) support for the text generation inference framework. They implemented and optimized custom ROCm kernels for attention mechanisms, including FP8 (float8) support. The user's contributions included updating existing kernels from VLLM and adding new kernels for flash decoding and other optimizations. They also integrated various kernel repositories and improved model support for ROCm.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.