Summary
Sreyan Ghosh is a research scientist and Ph.D. candidate in computer science specializing in speech, audio, and language processing, with eight years of industry and academic experience spanning Nvidia, Cisco, Google DeepMind, Adobe, and Microsoft. He blends applied ML engineering—delivering production NLP and speech solutions for enterprise customers—with frontier research on multi-modal and audio-centric models such as Audio Flamingo, Nemotron-Audio, and Cosmos World Models. His work bridges resource-conscious self-supervised speech techniques developed at IIT Madras and low-resource ASR and content moderation systems for Indian languages at IIIT Delhi, reflecting a strong focus on practical, inclusive ML. At Nvidia he co-led tool and model efforts (NEMO/RIVA contributions) and at Google DeepMind he now advances audio intelligence for Gemini, demonstrating rapid progression from solutions architect to research lead. He has a track record of publishing at major conferences (AAAI, ACL, Interspeech, ICASSP) and contributing educational resources (Keras examples via GSoC), indicating both scholarly impact and commitment to community-facing tooling. Colleagues describe him as someone who seamlessly moves between building scalable production systems and pushing novel research on multimodal audio understanding.
8 years of coding experience
6 years of employment as a software developer
High School, High School at Garden High School, Kolkata
Doctor of Philosophy - PhD Computer Science, Doctor of Philosophy - PhD Computer Science at University of Maryland
Bachelor of Technology - BTech Computer Science, Bachelor of Technology - BTech Computer Science at Christ University, Bangalore
English, Hindi, Bengali