Avnish Narayan is a research engineer with 8 years of experience building distributed ML and reinforcement learning systems, currently working on humanoid robotics and large-scale RL at NVIDIA GEAR Lab. Previously a Ray and RLlib maintainer at Anyscale, he helped scale RL tooling and shipped features for serving open-source LLMs, including function-calling and constrained JSON generation in Anyscale Endpoints. He has a strong open-source track record—contributing deterministic training, distributed samplers, and multi-task robotics benchmark fixes to prominent projects like Ray, Garage, and Metaworld. Comfortable across research and production, he combines deep RL algorithm work with MLOps and backend engineering to make experiments reproducible at cluster scale. Notably, his efforts have enabled reproducible, fully deterministic RL training runs and sped up parallel data collection in real-world toolchains.
8 years of coding experience
6 years of employment as a software developer
Master of Science - MS, Computer Science, Master of Science - MS, Computer Science at University of Southern California
Undeclared, Undeclared at University of Washington
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Role in this project:
ML Engineer
Contributions:1 release, 1097 reviews, 104 commits in 1 year 3 months
Contributions summary:Avnish primarily worked on the RLlib library within the Ray project, specifically on code related to Reinforcement Learning. Their contributions focused on the implementation of training error messages for KL penalties in the Distributed Distributional PPO (DDPPO) algorithm, and other improvements to the RLLib library. Additionally, the user addressed an issue by implementing a fully deterministic, repeatable RLlib train run using the "seed" config key.
A toolkit for reproducible reinforcement learning research.
Role in this project:
ML Engineer
Contributions:45 reviews, 34 commits, 70 PRs in 1 year 9 months
Contributions summary:Avnish implemented a distributed Ray sampler for the garage reinforcement learning toolkit, enabling parallel data collection. They addressed performance bottlenecks, particularly when using TensorFlow neural network policies, and optimized the sampler's speed. Further contributions include modifying the codebase to suppress Ray worker output, and a variety of bug fixes, improving the reliability of the project.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.