Mark Saroufim is a machine learning engineer and founder with 11 years of experience building high-performance ML systems, from training LLMs and kernel optimization to production model serving. Based in San Francisco, he co-founded GPU MODE and Core Automation, leading community-facing kernel competitions, events, and projects like Popcorn that explore GPU kernel generation and efficiency. His open-source contributions to PyTorch (including Inductor/Dynamo improvements, FX testing, vmap support, and GPU instrumentation) and work on pytorch/serve (DynamoDB snapshot serializer) show a rare combination of ML research, compiler-level tooling, and backend engineering. Prior roles span Meta, Graphcore, Microsoft, and startups, with an academic foundation in ML and theoretical CS from UC San Diego. A less obvious strength is his cross-domain fluency—writing Java backend services, contributing low-level GPU tooling, and steering community-driven applied research initiatives.
11 years of coding experience
13 years of employment as a software developer
University of California, San Diego
BE, Electrical and Computer Engineering, BE, Electrical and Computer Engineering at American University of Beirut
A set of examples around pytorch in Vision, Text, Reinforcement Learning, etc.
Role in this project:
ML Engineer
Contributions:138 reviews, 50 commits, 174 PRs in 6 months
Contributions summary:Mark contributed to the PyTorch examples repository by modifying and testing various machine learning examples. Their work includes reverting changes related to normalized loss in actor-critic and REINFORCE reinforcement learning examples. The user also added, disabled, and deleted tests related to FX (FX is a PyTorch library for program transformation). Further changes involved updating build scripts and reverting device configurations, indicating involvement in training scripts and infrastructure.
Serve, optimize and scale PyTorch models in production
Role in this project:
Backend Developer
Contributions:921 reviews, 590 commits, 568 PRs in 1 year 9 months
Contributions summary:Mark contributed to the development of a DynamoDB-based snapshot serializer and associated testing for the `pytorch/serve` repository. This primarily involved writing Java code, which interacts with AWS DynamoDB to manage and store snapshots. Further, the user updated the Java dependencies, upgraded the model archive, and addressed issues involving text generation, demonstrating proficiency in backend development tasks and a specific focus on enhancing the server's data persistence capabilities using AWS services.
pytorchmachine-learningmlopsservingdocker
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.