Peter Mcaughan

Senior System Software Engineer, AI at NVIDIA

San Francisco, California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Peter Mcaughan is a Senior System Software Engineer specializing in AI with eight years of experience building high-performance ML inference and cloud systems from research to production. Based in San Francisco, he contributed significant optimizations to the widely used ONNX Runtime—adding memory configuration options, Whisper model export/integration, and CUDA attention kernels—before joining NVIDIA to work on system-level AI software. His background blends academic research in ubiquitous computing and sign-language recognition with hands-on Azure Compute and cloud platform engineering, giving him a pragmatic research-to-product perspective. Known for digging into performance and memory bottlenecks, he brings a knack for making large models run faster and leaner in real-world environments.
code8 years of coding experience
job6 years of employment as a software developer
bookMaster's degree, Computer Science, Master's degree, Computer Science at Georgia Institute of Technology
bookBachelor’s Degree, Computer Engineering, Bachelor’s Degree, Computer Engineering at Texas A&M University
bookUniversity of São Paulo
languagesEnglish
github-logo-circle

Github Skills (11)

cuda10
machine-learning10
deep-learning10
onnx10
python9
pytorch9
c-language8
ai-framework8
neural-network8
hardware-acceleration8
cprogramming-language8

Programming languages (4)

C++CSSJavaScriptPython

Github contributions (5)

github-logo-circle
microsoft/onnxruntime

Sep 2022 - Oct 2022

ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Role in this project:
userML Engineer
Contributions:28 reviews, 2 commits, 29 PRs in 1 month
Contributions summary:Peter's contributions primarily focused on enhancing the ONNX Runtime, specifically related to integrating and optimizing models for machine learning inference and training. Their work involved adding support for various memory configuration options and implementing features for the Whisper model, including export scripts and integration with the BeamSearch operation. They also addressed performance bottlenecks and memory usage issues during the model conversion process. Further contributions included CUDA kernel implementations and optimization efforts for the attention layer within the CUDA environment.
runtimetrainingtensorflowai-frameworkaccelerator
petermcaughan/pmcaughanSite

Jul 2017 - Jun 2019

Source code for my personal website
Contributions:36 pushes, 1 branch in 1 year 11 months
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial