Summary
Sang-gil Lee is an applied deep learning research scientist with a Ph.D. from Seoul National University and a decade of experience advancing generative models for speech and audio. He has led foundational work on waveform synthesis—most notably as lead author of BigVGAN—and has interned at top labs including NVIDIA and Microsoft Research, where he contributed to state-of-the-art diffusion and vocoder models. Currently at NVIDIA, he focuses on building multi-modal large language models centered on audio, while previously developing edge-optimized TTS frameworks at Qualcomm. Sang-gil’s research spans AR, flow, GAN, and diffusion methods applied to sequential data, with practical impact across TTS, voice conversion, music generation, neural audio codecs, and audio language models. His profile reflects a rare combination of rigorous academic work, high-impact conference publications, and production-aware engineering for real-world audio systems.
10 years of coding experience
1 year of employment as a software developer
Doctor of Philosophy - PhD, Electrical and Computer Engineering, Doctor of Philosophy - PhD, Electrical and Computer Engineering at Seoul National University
Bachelor of Engineering - BE, Electrical and Computer Engineering, Cum Laude, Bachelor of Engineering - BE, Electrical and Computer Engineering, Cum Laude at 서울대학교 (Seoul National University)
English, Korean