Senior Developer Technology Engineer - Visual Computing at NVIDIA
Aachen, North Rhine-Westphalia, Germany
Join Prog.AI to see contacts
Join Prog.AI to see contacts
Summary
🤩
Rockstar
🎓
Top School
Maximilian Müller is a Senior Developer Technology Engineer specializing in visual computing and CUDA-driven optimization of end-to-end software pipelines, with seven years of experience at NVIDIA. He focuses on squeezing performance from GPU stacks—contributing to high-profile projects like ONNX Runtime (TensorRT provider) and NVIDIA's Video Processing Framework—improving inference build times, tensor-core utilization, and CI/build tooling. With an academic background from RWTH Aachen and Aalto in electrical engineering and machine learning, he bridges research and production, turning deep learning prototypes into scalable, hardware-accelerated deployments. Outside of core engineering he’s built practical developer tooling and CI improvements and even coached mountain biking, reflecting a hands-on, team-oriented approach and appetite for applied problem solving.
7 years of coding experience
4 years of employment as a software developer
Machine Learning, Data Science, Machine Learning, Data Science at Aalto University
Master of Science - MS, Elektrotechnik, Eletronik und Kommunikationstechnik - Vertiefung Computer Engineering, Master of Science - MS, Elektrotechnik, Eletronik und Kommunikationstechnik - Vertiefung Computer Engineering at RWTH Aachen University
Allgemeine Hochschulreife, 1,6, Allgemeine Hochschulreife, 1,6 at Freiherr-vom-Stein Gymnasium Betzdorf
Set of Python bindings to C++ libraries which provides full HW acceleration for video decoding, encoding and GPU-accelerated color space and pixel format conversions
Role in this project:
MLOps Engineer
Contributions:53 reviews, 20 commits, 31 PRs in 2 months
Contributions summary:Maximilian made contributions focused on improving the continuous integration (CI) process and adding functionality to the existing samples. The user added new samples to the CI testing framework and updated existing sample code by modifying the existing scripts. Additionally, the user refactored the build processes and integrated code formatting. The user also addressed warnings and updated coding practices.
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Role in this project:
ML Engineer
Contributions:99 reviews, 38 PRs, 263 comments in 3 years 7 months
Contributions summary:Maximilian primarily contributed to the TensorRT (TRT) execution provider within the ONNX Runtime project, focusing on features related to optimizing and accelerating model inference. Their work includes enabling TensorRT timing caches for faster build times, exposing new build options, and refitting embedded engines. They also made improvements to CUDA and TRT integration, including NHWC data layout support for enhanced tensor core utilization and addressing compilation issues.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.