Summary
Muhammad Maaz is a research scientist and final-year PhD candidate specializing in multimodal large language models and long-form video understanding, based at MBZUAI with advising from Salman Khan and Fahad Khan. He has eight years of experience bridging applied computer vision and research, with production deployments on edge devices and industry internships at Meta where he developed open, competitive MLLMs like PerceptionLM/PLM. His work includes Video-ChatGPT and a 100K high-quality video-instruction dataset plus the first quantitative evaluation framework for video conversation models, highlighting a rare combination of dataset engineering, evaluation rigor, and model innovation. Comfortable moving models from lab to product, he has built CPU-optimized vision systems for real-world retail and traffic applications and mentors students as a teaching assistant. Based in Beijing, he brings deep academic training (MS with 4.0 GPA) and hands-on engineering that enables reproducible, open research that competes with closed-state-of-the-art systems.
8 years of coding experience