Chirayu Garg

Engineering Manager, Developer Technology at NVIDIA

California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Chirayu Garg is a Senior Engineering Manager at NVIDIA with 7 years of experience specializing in deep learning, computer vision, HPC and system software. He has progressed from system software and video-algorithms roles into developer technology leadership, building teams that optimize AI models for NVIDIA hardware and drive recommender-system performance for key partners and MLPerf. Hands-on contributions to high-profile open-source projects such as RAPIDS cuML and NVIDIA HugeCTR demonstrate his expertise in GPU performance engineering—adding asynchronous stream support and optimizing Parquet readers for multi-node training. He combines academic rigor from UW–Madison with practical driver-to-application engineering, having implemented GPU-parallel image processing, CNN-based quality models, and driver features for shipping products. Colleagues describe him as a pragmatic leader who still codes complex backend performance fixes, bridging low-level CUDA/driver work with large-scale ML system design.
code7 years of coding experience
job8 years of employment as a software developer
bookBITS Pilani, Birla Institute of Technology and Science
bookMaster of Science - MS Computer Science, Master of Science - MS Computer Science at University of Wisconsin-Madison
bookBirla Public School,Pilani
github-logo-circle

Github Skills (17)

c-language10
cublas10
gpu-programming10
cudf10
cusolver10
deeplearning-ai10
deep-learning10
parquet10
cuda10
cpp10
gpu-acceleration10
cprogramming-language10
linear-algebra10
recommendation-system9
recommender-system9

Programming languages (5)

C++HTMLJupyter NotebookPythonCuda

Github contributions (5)

github-logo-circle
rapidsai/cuml

Mar 2019 - May 2019

cuML - RAPIDS Machine Learning Library
Role in this project:
userBack-end Developer & ML Engineer
Contributions:53 commits, 5 PRs, 26 comments in 2 months
Contributions summary:Chirayu primarily contributed to adding stream parameters to various cusolver and cublas functions. This was likely done to enhance GPU resource utilization and performance within the cuML library. The changes involved modifying header files to incorporate the stream parameter, enabling asynchronous execution. The user's work focused on integrating these features into core linear algebra functions used by the cuML project.
cudacumlnvidiadata-sciencegpu
NVIDIA-Merlin/HugeCTR

Aug 2020 - Nov 2020

HugeCTR is a high efficiency GPU framework designed for Click-Through-Rate (CTR) estimating training
Role in this project:
userBack-end Developer & Performance Engineer
Contributions:12 commits, 2 comments in 3 months
Contributions summary:Chirayu primarily focused on optimizing the Parquet data reader within the HugeCTR framework, as indicated by the code changes related to row group caching and stream synchronization. They added functionalities like row group caching to the parquet reader, and improved performance by addressing blocking issues. Additionally, the commits involved integrating cudf and rmm as submodules and building scripts to build librmm and libcudf. Furthermore, the user added multi-node parquet data reader support.
cudapytorchcppgpu-accelerationdeep-learning
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial