Siddharth Dalmia is a research scientist with nine years of experience building next-generation voice and multimodal audio systems, currently advancing Audio LLM-driven voice interactions at Meta SuperIntelligence Labs after WaveForms AI’s acquisition. He previously contributed to long-context and multimodal audio work for Gemini at Google DeepMind and holds a Ph.D. in Language Technologies from Carnegie Mellon, where his research emphasized compositional system design—task simplification, reusability, transferability, and data-pooling—for sequence models in speech and language tasks. Siddharth combines deep research rigor with practical engineering, evidenced by hands-on contributions to the widely used ESPnet toolkit (improving compatibility, tooling, and model logging). Based in New York, he bridges academia and industry through internships at Google Brain, AWS, Facebook AI, and Inria, bringing a proven ability to move speech research into scalable, production-ready systems.
9 years of coding experience
9 years of employment as a software developer
BITS Pilani, Birla Institute of Technology and Science
Doctor of Philosophy - PhD Language Technologies Computer Science, Doctor of Philosophy - PhD Language Technologies Computer Science at Carnegie Mellon University
Contributions:26 reviews, 53 commits, 30 PRs in 1 year 4 months
Contributions summary:Siddharth primarily focused on improving the code's compatibility, debugging, and expanding the functionality of the core components. They addressed issues related to PyTorch versions, run script errors, and automated the process of resampling audio files. Additionally, the user contributed to improving model parameter logging and made several modifications to the configuration options of the language model.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.