Junyi Du is a Machine Learning Lead with 11 years of experience specializing in NLP, knowledge mining, and production-facing speech systems. He blends academic rigor from a 4.0 MS at USC with hands-on engineering across research labs and startups, shipping features from browser-side DL (WebGPU/WASM) to backend ASR pipelines. At Logseq he contributed full-stack search and RAG infrastructure and improved frontend editor/search components for a popular open-source knowledge management app, while at wenet he enhanced end-to-end speech recognition tooling for production use. Comfortable in languages from Clojure(Script) to Python and C++, he focuses on turning research ideas into deployable systems and developer-facing tooling. Notably, he combines expertise in vector/semantic search and lightweight client-side ML to push AI capabilities closer to users.
11 years of coding experience
6 years of employment as a software developer
Master of Science - MS, Computer Science, 4.0/4.0, Master of Science - MS, Computer Science, 4.0/4.0 at University of Southern California
Bachelor of Science - BS, Computer Science, 89.83/100, Bachelor of Science - BS, Computer Science, 89.83/100 at South China Agricultural University
A privacy-first, open-source platform for knowledge management and collaboration. Download link: http://github.com/logseq/logseq/releases. roadmap: https://logseq.io/p/NX4mc_ggEV
Role in this project:
Full-stack Developer
Contributions:443 reviews, 224 commits, 160 PRs in 1 year 2 months
Contributions summary:Junyi primarily contributed to the Logseq project by fixing and enhancing various frontend components. They addressed issues related to user interface elements, particularly for the editor and search components. These changes involved modifications to JavaScript, ClojureScript, and CSS code, demonstrating a focus on frontend development within the knowledge management application's codebase. Additionally, the user expanded the project's configuration, reflecting a broader understanding of the project's structure.
Production First and Production Ready End-to-End Speech Recognition Toolkit
Role in this project:
Back-end Developer
Contributions:10 commits, 3 PRs, 2 comments in 3 months
Contributions summary:Junyi primarily contributed to the `wenet` repository by modifying scripts and code related to the end-to-end speech recognition toolkit. Their work includes adapting training scripts for different datasets, adding options to the decoding process to output search candidates, and introducing an option for the CtcPrefixBeamSearch algorithm. The user also made improvements to the data preparation scripts and corrected formatting issues.
speech-recognitione2e-modelspytorchasrtransformer
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.