Charlie Ruan is a PhD student and systems-focused ML engineer based in Berkeley with six years of experience building high-performance ML and inference infrastructure. He has contributed substantial code to flagship open-source projects such as MLC-LLM and web-llm—adding model support (WizardLM, WizardCoder, Mistral), Python APIs, sliding window attention, and .safetensors handling—to enable universal local and in-browser LLM deployment. His background spans cloud and accelerator runtimes (TensorFlow core contributions and TVM WebGPU/WebAssembly work) as well as distributed training and systems research from roles at Carnegie Mellon and Cornell. Currently advised by Ion Stoica at Berkeley, he blends academic research on post-training and SkyRL with hands-on engineering that repeatedly bridges compiler, runtime, and model integration. A non-obvious strength: he shifts fluently between low-level runtime optimizations and higher-level model/tooling integration, making him effective at productionizing cutting-edge ML research.
6 years of coding experience
4 years of employment as a software developer
Doctor of Philosophy - PhD, Computer Science, Doctor of Philosophy - PhD, Computer Science at University of California, Berkeley
Bachelor of Science - BS, Computer Science, Operations Research, Bachelor of Science - BS, Computer Science, Operations Research at Cornell University
Master of Science - MS, Computer Science, Master of Science - MS, Computer Science at Carnegie Mellon University
Contributions:55 reviews, 311 PRs, 214 pushes in 1 year 6 months
Contributions summary:Charlie primarily contributed to the development of models and features within the web-llm project. Their work involved adding support for new models, specifically the WizardCoder, WizardMath, and Mistral families, and incorporating features such as sliding window attention (SWA) for enhanced performance. They also made modifications to the conversation templates and added support for different models to the simple chat example. Further, they enabled the integration of tools with models, particularly focusing on supporting a JSON format for the responses.
Universal LLM Deployment Engine with ML Compilation
Role in this project:
ML Engineer
Contributions:71 reviews, 133 PRs, 45 pushes in 1 year 9 months
Contributions summary:Charlie primarily contributed to the implementation of support for different language models within the MLC-LLM framework. Specifically, they added support for the WizardLM model, modifying the conversation template, and integrating it into the existing infrastructure. Furthermore, the user added a Python API and build arguments for building model in a Python script. They also added support for `.safetensors` files, enabling users to use a more secure format.
language-modelllmmachine-learning-compilationtvm
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.