Vishwa Karia is a Senior Software Engineer at Meta specializing in AI infrastructure, with seven years of experience building scalable distributed training systems for PyTorch-based personalization models. He combines deep learning and high-performance computing expertise with cloud-native engineering, having driven large-scale improvements at AWS SageMaker—optimizing distributed training runtimes and reliability for multi-thousand GPU clusters. Vishwa has hands-on MLOps experience (including contributions to the aws/sagemaker-training-toolkit) and a track record of designing automated systems that balance security, performance, and customer experience. He mentors women and students in tech, serves on AI-for-good initiatives, and thrives on teams that aim to make a positive societal impact. Comfortable both in low-level distributed systems work and cross-functional leadership, he blends academic grounding from UCLA with product-driven delivery at major cloud and consumer platforms. Colleagues describe him as a pragmatic problem-solver who seeks elegant, production-ready solutions to hard ML infrastructure challenges.
7 years of coding experience
5 years of employment as a software developer
Bachelor of Technology - BTech Computer Engineering, Bachelor of Technology - BTech Computer Engineering at Veermata Jijabai Technological Institute (VJTI)
University of California, Los Angeles
Higher Secondary School Maharashtra State Board Exam Science, Higher Secondary School Maharashtra State Board Exam Science at Ramnarain Ruia College
Train machine learning models within a 🐳 Docker container using 🧠 Amazon SageMaker.
Role in this project:
MLOps Engineer
Contributions:9 reviews, 5 commits, 8 PRs in 6 months
Contributions summary:Vishwa primarily focused on enhancing the integration and functionality of the SageMaker training toolkit, particularly around distributed training. They addressed UTF-8 decoding exceptions, updated the environment configurations, and modified the SMDataParallel runner. Furthermore, the user integrated SMDDP collectives, impacting the performance of the SMDataParallel runner by adding conditional logic based on the communication backend used and availability of SMDDP libraries. They also updated libraries to validate SMDDP collective functionality.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.