Yuekai Zhang

解决方案工程师 at 英伟达公司

Shanghai, Shanghai, China
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Yuekai Zhang is a solutions engineer based in Shanghai with eight years' experience focused on end-to-end automatic speech recognition (ASR) and production deployment. At NVIDIA and previously Zoom, he has bridged model development and serving—contributing recipes, data pipelines, and Triton-based inference integrations for prominent open-source toolkits like WeNet, ESPnet, and Sherpa. His work spans ML engineering and DevOps: adding RNN-T recipes, fp16 DDP optimizations, model export tooling, and streaming ASR support to make research models production-ready. A Johns Hopkins ECE master's graduate with an EE bachelor's from Xi’an Jiaotong, he combines rigorous academic grounding with hands-on engineering. Notably, his contributions often focus on reproducible pipelines and deployment scripts that simplify taking complex speech models from training to low-latency serving.
code8 years of coding experience
job1 year of employment as a software developer
bookJohns Hopkins University
bookBachelor's degree, EE, Bachelor's degree, EE at 西安交通大学
github-logo-circle

Github Skills (31)

pytorch10
docker10
python10
scripting10
onnx10
automatic-speech-recognition10
dockers10
cicd10
data-preprocessing10
script10
triton10
dataprep10
asr10
build-automation10
sh10

Programming languages (7)

C++CHTMLJupyter NotebookMATLABPythonCuda

Github contributions (5)

github-logo-circle
k2-fsa/sherpa

May 2022 - Jan 2023

Speech-to-text server framework with next-gen Kaldi
Role in this project:
userML Engineer & DevOps Engineer
Contributions:39 reviews, 11 commits, 33 PRs in 7 months
Contributions summary:Yuekai contributed to the integration of a Triton server for speech-to-text, specifically focusing on deploying a Conformer model. They added configurations and scripts for the Triton inference server, along with necessary modifications to the model export and client-side scripts. Additionally, the user added code to support streaming ASR using the Triton server and made updates to the onnx export process and config files. These contributions suggest a focus on model deployment, serving infrastructure, and potentially performance optimization within the speech recognition framework.
pytorchtransducerkaldicpppython
wenet-e2e/wenet

May 2021 - Jan 2023

Production First and Production Ready End-to-End Speech Recognition Toolkit
Role in this project:
userML Engineer
Contributions:12 reviews, 11 commits, 18 PRs in 1 year 7 months
Contributions summary:Yuekai primarily contributes to the examples directory, updating and adding recipes for end-to-end speech recognition tasks. They modify and adapt scripts for different datasets like Librispeech and Tedlium3, focusing on data preparation, dictionary creation, and data formatting. Their commits demonstrate involvement in model training and testing, including the addition of RNN-T recipe and model average utilities. The user also adds features like Fp16 grad sync for ddp and updates docker configurations for triton.
pytorchend-to-endasrproduction-readyspeech-to-text
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial