Senior Open Source ML Engineer at Amazon Web Services (AWS)
San Diego, California, United States
Join Prog.AI to see contacts
Join Prog.AI to see contacts
Summary
🤩
Rockstar
🎓
Top School
Omri S is a Senior Open Source ML Engineer with 13 years in software and nine years remote specializing in ML platforms, Kubernetes, DevOps, and cloud-native infrastructure across gaming, retail, banking, and hospitality. Previously a Principal ML Engineer at Roblox, he redesigned schedulers and built asynchronous serving, CI/CD controllers, dynamic storage, and monitoring that scaled inference and notebook workflows for production ML. Now at AWS, he continues to bridge open-source tooling and specialized hardware—contributing to vLLM with Trainium/Inferentia Neuron support and tensor parallelism—bringing uncommon expertise in optimizing LLM serving on accelerators. He combines hands-on systems programming (Python, Go, Java) with product-led adoption, mentorship, and incident leadership to move models from prototype to reliable production. Educated in information systems and operations research, he pairs analytical rigor with pragmatic platform design to accelerate ML velocity for teams.
14 years of coding experience
14 years of employment as a software developer
Master of Science (MS), Information Systems, Master of Science (MS), Information Systems at Weatherhead School of Management at Case Western Reserve University
A high-throughput and memory-efficient inference and serving engine for LLMs
Role in this project:
MLOps Engineer
Contributions:7 reviews, 7 PRs, 9 comments in 10 months
Contributions summary:Omri's contributions primarily focused on integrating and supporting AWS Neuron (Trainium/Inferentia) within the vLLM framework. They added documentation regarding Neuron installation and usage. Their work involved bug fixes and feature implementations, including enabling tensor parallelism and ensuring compatibility with different Neuron SDK versions. The user also updated the code to reflect the changes of the neuron device, specifically addressing the block size settings.
A high-throughput and memory-efficient inference and serving engine for LLMs
Contributions:1 PR, 55 pushes, 12 branches in 8 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.