Omri S

Senior Open Source ML Engineer at Amazon Web Services (AWS)

San Diego, California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Omri S is a Senior Open Source ML Engineer with 13 years in software and nine years remote specializing in ML platforms, Kubernetes, DevOps, and cloud-native infrastructure across gaming, retail, banking, and hospitality. Previously a Principal ML Engineer at Roblox, he redesigned schedulers and built asynchronous serving, CI/CD controllers, dynamic storage, and monitoring that scaled inference and notebook workflows for production ML. Now at AWS, he continues to bridge open-source tooling and specialized hardware—contributing to vLLM with Trainium/Inferentia Neuron support and tensor parallelism—bringing uncommon expertise in optimizing LLM serving on accelerators. He combines hands-on systems programming (Python, Go, Java) with product-led adoption, mentorship, and incident leadership to move models from prototype to reliable production. Educated in information systems and operations research, he pairs analytical rigor with pragmatic platform design to accelerate ML velocity for teams.
code14 years of coding experience
job14 years of employment as a software developer
bookMaster of Science (MS), Information Systems, Master of Science (MS), Information Systems at Weatherhead School of Management at Case Western Reserve University
languagesEnglish, Hebrew
stackoverflow-logo

Stackoverflow

Stats
19reputation
2kreached
0answers
3questions
github-logo-circle

Github Skills (17)

python10
inference10
llm10
pytorch9
mlops9
transformer8
documentation8
shell6
exec6
php6
linux6
bash6
awk6
deepspeed3
roc3

Programming languages (17)

JavaC++CSSCGoMustacheHTMLGroovy

Github contributions (5)

github-logo-circle
vllm-project/vllm

Jun 2024 - Apr 2025

A high-throughput and memory-efficient inference and serving engine for LLMs
Role in this project:
userMLOps Engineer
Contributions:7 reviews, 7 PRs, 9 comments in 10 months
Contributions summary:Omri's contributions primarily focused on integrating and supporting AWS Neuron (Trainium/Inferentia) within the vLLM framework. They added documentation regarding Neuron installation and usage. Their work involved bug fixes and feature implementations, including enabling tensor parallelism and ensuring compatibility with different Neuron SDK versions. The user also updated the code to reflect the changes of the neuron device, specifically addressing the block size settings.
inferencegptllmpytorchmodel-serving
omrishiv/vllm

Jul 2024 - Apr 2025

A high-throughput and memory-efficient inference and serving engine for LLMs
Contributions:1 PR, 55 pushes, 12 branches in 8 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial