Senior Deep Learning Software Development Engineer at NVIDIA
Sunnyvale, California, United States
Join Prog.AI to see contacts
Join Prog.AI to see contacts
Summary
🤩
Rockstar
🎓
Top School
Dick Carter is a Senior Deep Learning Software Development Engineer with over nine years focused on GPU-accelerated ML frameworks and a multi-decade background in hardware and research from Hewlett Packard Labs to NVIDIA. He specializes in optimizing convolutional neural network performance—contributing key cuDNN enhancements, dilated convolution and TensorCore support, and kernel-level optimizations to the widely used Apache MXNet project. At NVIDIA he drives tighter, higher-performance integration of NVIDIA tech into third-party GPU frameworks, building on past work on a general-purpose GPU programming toolkit and DARPA SyNAPSE research. Trained at MIT and Stanford in EE and computer engineering, he blends low-level hardware and compiler insight with practical deep learning engineering. Based in Sunnyvale, he combines production-grade performance tuning with a rare background in CPU architecture and memristor-inspired neuromorphic efforts, making him especially effective at bridging research innovations to deployable GPU software.
9 years of coding experience
37 years of employment as a software developer
BSEE, Electrical engineering and computer science, BSEE, Electrical engineering and computer science at Massachusetts Institute of Technology
MSEE, Computer engineering, MSEE, Computer engineering at Stanford University
Lightweight, Portable, Flexible Distributed/Mobile Deep Learning with Dynamic, Mutation-aware Dataflow Dep Scheduler; for Python, R, Julia, Scala, Go, Javascript and more
Role in this project:
Back-end Developer & ML Engineer
Contributions:37 reviews, 69 commits, 111 PRs in 5 years 5 months
Contributions summary:Dick's contributions primarily focused on enhancing the cuDNN integration for the MXNet deep learning framework. This included implementing support for dilated convolutions and TensorCore, indicating a focus on accelerating convolutional neural networks. The user also worked on optimizing the framework by avoiding unneeded backpropagation kernels, and by converting dot operations to linalg gemm, as well as working on various bug fixes and performance improvements related to the cuDNN backend.
Lightweight, Portable, Flexible Distributed/Mobile Deep Learning with Dynamic, Mutation-aware Dataflow Dep Scheduler; for Python, R, Julia, Scala, Go, Javascript and more
Contributions:348 pushes, 91 branches in 6 years
pythonschedulerdataflowmutationdata-science
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.