Chuck Tang is an ML engineer based in Berkeley with a decade of hands-on experience building and shipping large-scale ML training and runtime systems, currently focused on reasoning and agentic language models at PyTorch. He has been a top contributor to open-source LLM training stacks like MosaicML Composer and LLM-Foundry, improving Docker/CI workflows, performance profiling, and integration tests that keep foundation-model training reliable in production. At Databricks he implemented cutting-edge features—FP8 for mixture-of-experts and robust regression testing—and previously bridged sim-to-real gaps for LiDAR perception at Applied Intuition. Comfortable across research and production, he combines C++/ROS safety tooling from Berkeley AI research with modern PyTorch-based MLOps expertise, and has a track record of turning flaky tests and configs into deterministic pipelines.
10 years of coding experience
3 years of employment as a software developer
Master's degree EECS, Master's degree EECS at University of California, Berkeley
LLM training code for Databricks foundation models
Role in this project:
ML Engineer
Contributions:1 release, 97 reviews, 109 PRs in 1 year 5 months
Contributions summary:Chuck primarily contributed to the integration testing and the overall training process for the LLM-foundry project. They fixed an integration test, which involved mocking a dataset and ensuring the codebase reproduced a golden loss score. They also added checks to prevent errors in YAML configuration files during training and evaluation, ensuring that improperly formatted configurations trigger runtime errors to streamline the pipeline. The user made further contributions to improve model testing and streamline the overall development process.
Contributions:3 releases, 196 reviews, 176 PRs in 1 year 11 months
Contributions summary:Chuck focused on enhancing the development and deployment environment for the `composer` project, which is focused on ML training. Their contributions involved adding and updating Docker images to support new PyTorch versions, CUDA versions, and nightly builds. They also modified build scripts, tested the docker images in CI/CD pipelines, and addressed OOM issues. The user demonstrated knowledge of system performance by adjusting the test environment.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.