Cheng Phoo is a software engineer and researcher with 11 years of experience specializing in multimodal perception and foundation models for vision and language. He completed a PhD at Cornell focusing on sample-efficient visual recognition, domain adaptation for remote sensing and LiDAR-assisted 3D detection, and published work on unsupervised discovery of mobile objects and few-shot learning. Cheng transitioned from academic research and high-profile internships (MIT-IBM Watson AI Lab, Meta FAIR) to industry roles at Apple as a postdoc researcher on multimodal LLMs and now builds multimodal foundation models for Waymo’s onboard perception. He blends deep theoretical grounding with production-oriented engineering, routinely translating pre-trained foundation models into resource-efficient, domain-adapted systems. Based in Cupertino, Cheng pairs a strong math and CS background (University of Michigan) with hands-on experience in large-scale perception problems that bridge camera, LiDAR, and language modalities. An interesting thread through his work is leveraging unlabeled repeated-traversal LiDAR data to improve detection and adaptation in low-annotation domains.
11 years of coding experience
9 years of employment as a software developer
Bachelor’s Degree Pure Mathematics and Computer Science, Bachelor’s Degree Pure Mathematics and Computer Science at University of Michigan College of Literature, Science, and the Arts
Doctor of Philosophy - PhD Computer Science - Computer Vision and Machine Learning, Doctor of Philosophy - PhD Computer Science - Computer Vision and Machine Learning at Cornell University
American Degree Transfer Program Actuarial Science, American Degree Transfer Program Actuarial Science at INTI
Contributions:7 commits, 6 pushes, 1 branch in 11 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.