Jong Kim is a Member of Technical Staff at OpenAI with 16 years of experience building large-scale, distributed systems and production-grade ML infrastructure. He brings deep expertise in audio and music technology—having built real-time polyphonic synthesis at Spotify and contributed to prominent open-source projects like OpenAI's Whisper and CLIP—while also designing recommender and stream-processing platforms at Kakao and Pandora. Jong combines low-latency systems engineering (Scala/Java/C++) with machine learning model work (pitch estimation, speech recognition), enabling end-to-end solutions from model research to scalable deployment. He has a strong academic foundation with a PhD in Music Technology and practical experience optimizing CI/CD, Spark jobs, and fault-tolerant services. Notably, his contributions to Whisper improved decoding and language handling, reflecting an ability to enhance core model behavior as well as system-level performance.
16 years of coding experience
4 years of employment as a software developer
Master of Science, Computer Science and Engineering, 7.76/9.00, Master of Science, Computer Science and Engineering, 7.76/9.00 at University of Michigan
Doctor of Philosophy (Ph.D.), Music Technology, 3.76/4.00, Doctor of Philosophy (Ph.D.), Music Technology, 3.76/4.00 at New York University
Bachelor of Science, Electrical Engineering, 3.83/4.30, Bachelor of Science, Electrical Engineering, 3.83/4.30 at Korea Advanced Institute of Science and Technology
Robust Speech Recognition via Large-Scale Weak Supervision
Role in this project:
Back-end Developer
Contributions:11 reviews, 52 commits, 177 PRs in 4 months
Contributions summary:Jong primarily focused on modifying the `whisper/transcribe.py` file, suggesting a role in the core functionality and logic of the speech recognition system. Their changes involved altering model defaults, renaming variables, and adding features such as the `condition_on_previous_text` flag and improved language handling, reflecting a contribution to the overall performance and configuration options of the transcription process. They also made related adjustments to the decoding and tokenization modules.
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
Role in this project:
ML Engineer
Contributions:31 commits, 45 PRs, 58 pushes in 1 year 4 months
Contributions summary:Jong made significant contributions to the `openai/clip` repository, primarily focused on model implementation and updates. They added a new RN50 checkpoint, implemented a non-JIT model, and incorporated several other model variants (RN101, RN50x4, RN50x16, RN50x64, ViT-B/16, ViT-L/14, and ViT-L/14@336px). Their work also included initializing model parameters and fixing model loading issues, demonstrating a strong understanding of the model architecture and its practical usage. Additionally, the user has also added a prompt engineering notebook for the end-user.
styleganclipcomputer-visioncontrastivepretraining
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.