Siddharth Dalmia

Member Of Technical Staff at WaveForms AI

New York, New York, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Siddharth Dalmia is a research scientist with nine years of experience building next-generation voice and multimodal audio systems, currently advancing Audio LLM-driven voice interactions at Meta SuperIntelligence Labs after WaveForms AI’s acquisition. He previously contributed to long-context and multimodal audio work for Gemini at Google DeepMind and holds a Ph.D. in Language Technologies from Carnegie Mellon, where his research emphasized compositional system design—task simplification, reusability, transferability, and data-pooling—for sequence models in speech and language tasks. Siddharth combines deep research rigor with practical engineering, evidenced by hands-on contributions to the widely used ESPnet toolkit (improving compatibility, tooling, and model logging). Based in New York, he bridges academia and industry through internships at Google Brain, AWS, Facebook AI, and Inria, bringing a proven ability to move speech research into scalable, production-ready systems.
code9 years of coding experience
job9 years of employment as a software developer
bookBITS Pilani, Birla Institute of Technology and Science
bookDoctor of Philosophy - PhD Language Technologies Computer Science, Doctor of Philosophy - PhD Language Technologies Computer Science at Carnegie Mellon University
github-logo-circle

Github Skills (16)

pytorch10
python10
scripting9
machine-learning9
shell9
script9
sh9
deep-learning8
speech-synthesis8
speech-recognition8
text-to-speech8
machine-translation8
docker4
kubernetes4
dockers4

Programming languages (7)

CSSShellC++SCSSJavaScriptPythonCuda

Github contributions (5)

github-logo-circle
espnet/espnet

Oct 2020 - Mar 2022

End-to-End Speech Processing Toolkit
Role in this project:
userBackend & DevOps Engineer
Contributions:26 reviews, 53 commits, 30 PRs in 1 year 4 months
Contributions summary:Siddharth primarily focused on improving the code's compatibility, debugging, and expanding the functionality of the core components. They addressed issues related to PyTorch versions, run script errors, and automated the process of resampling audio files. Additionally, the user contributed to improving model parameter logging and made several modifications to the configuration options of the language model.
speech-recognitionspeech-separationchainerspoken-language-understandingspeech-processing
siddalmia/espnet

Feb 2021 - Aug 2022

End-to-End Speech Processing Toolkit
Contributions:5 PRs, 125 pushes, 24 branches in 1 year 6 months
end-to-endspeech-to-textspeech-recognitionspeech-synthesisspeech
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial