Yuhao Zhang is a software engineer with 12 years of experience building scalable systems and shipping user-focused solutions, currently at Google after prior roles at Amazon. He blends full-stack JavaScript expertise (React, Node, Express, MongoDB) with deeper systems and NLP work, contributing backend fixes and feature additions to prominent Stanford NLP projects like CoreNLP and Stanza. His background includes MS in Computer Science from USC and a BS from UC Santa Barbara, and he has applied his skills to real-world product problems such as reducing customer on-hold time and improving gift-selection flows. Notably, his open-source contributions include time-annotation and Chinese NER enhancements to CoreNLP and build/dependency improvements plus lemmatization support for Stanza, reflecting both research-adjacent and production readiness. He approaches engineering through rapid learning and implementation, favoring practical improvements that bridge research tools and customer-facing products.
12 years of coding experience
2 years of employment as a software developer
Master of Science - MS, Computer Science, Master of Science - MS, Computer Science at University of Southern California
Bachelor of Science - BS, Computer Science, 3.85, Bachelor of Science - BS, Computer Science, 3.85 at UC Santa Barbara
Stanford NLP Python library for tokenization, sentence segmentation, NER, and parsing of many human languages
Role in this project:
Back-end Developer & DevOps Engineer
Contributions:3 releases, 8 reviews, 439 commits in 3 years 1 month
Contributions summary:Yuhao's contributions primarily revolved around improving the build process and dependency management of the StanfordNLP Stanza library. They fixed dependencies on internal scripts, updated the README and improved core functionality by addressing protobuf calls and handling of environment variables. They implemented a shell script to create basic lemma scripts and added the addition of a dictionary-based lemmatizer. They also implemented support for multi-word tokens and incorporated better file management for the data.
CoreNLP: A Java suite of core NLP tools for tokenization, sentence segmentation, NER, parsing, coreference, sentiment analysis, etc.
Role in this project:
Back-end Developer
Contributions:30 commits, 1 push in 3 years 11 months
Contributions summary:Yuhao primarily focused on modifying code related to the `SUTimeITest.java` file, which appears to be a testing file for the SUTime tool. These changes suggest work within the time-based annotation aspect of the `corenlp` project. This includes updates to handle and test Timex expressions within the time annotation pipeline. Further contributions included refactoring and code changes relating to CoreNLPProtos, specifically adding `docDate` and `calendar` attributes. Additionally, there are commits relating to the Chinese NER, and the related components, demonstrating the implementation of a new feature.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.