Przemyslaw P is a Deep Learning Algorithms Manager at NVIDIA with 11 years of experience designing and delivering inference optimization, deployment tooling and end-to-end model pipelines such as PyTriton and Triton Model Navigator. He combines hands-on performance engineering—contributing to high-profile open-source projects like TVM, MXNet, NVIDIA DALI and TransformerEngine—with people and product leadership honed at Samsung and deepsense.ai. His core strengths are GPU-aware memory and data pipeline optimizations, deterministic pipeline construction, and fixing subtle concurrency and graph-compilation bugs that improve real-world inference stability and throughput. He brings an uncommon blend of embedded, backend and DevOps experience, which helps bridge low-level CUDA/kernel work with cloud deployment concerns. Przemyslaw also holds an MBA and formal training in systems analysis, enabling him to translate technical tradeoffs into pragmatic product decisions. Based in Warsaw, he’s an active contributor to foundational ML infra where his commits have touched performance-critical codepaths used by many downstream projects.
11 years of coding experience
13 years of employment as a software developer
Bachelor of Engineering (B.Eng.), Information Technology, Bachelor of Engineering (B.Eng.), Information Technology at Politechnika Świętokrzyska w Kielcach
Master of Business Administration (MBA), MBA, Master of Business Administration (MBA), MBA at Akademia Leona Koźmińskiego
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper and Ada GPUs, to provide better performance with lower memory utilization in both training and inference.
Role in this project:
ML Engineer
Contributions:15 releases, 701 reviews, 25 commits in 3 months
Contributions summary:The user, Przemek Tredak, primarily contributed to the development and improvement of the `transformerengine` library, focusing on performance and functionality. Their work includes implementing new features, such as linking to documentation archives and adding pylint to the lint action. Przemek's contributions also involved fixing critical issues in the codebase, like out-of-bounds memory access in the C+T+dbias kernel. Significant changes were made to enhance the library's functionality with FP8 precision and integration with NVTX for improved debugging and profiling.
Lightweight, Portable, Flexible Distributed/Mobile Deep Learning with Dynamic, Mutation-aware Dataflow Dep Scheduler; for Python, R, Julia, Scala, Go, Javascript and more
Role in this project:
Back-end Developer & Performance Engineer
Contributions:1 release, 100 reviews, 151 commits in 4 years 10 months
Contributions summary:Przemyslaw primarily focused on optimizing the performance of the MXNet deep learning framework. Their contributions involved modifying GPU memory allocation, particularly in the `pooled_storage_manager.h` and `pooled_storage_manager.cc` files. They also worked on improving the image I/O pipeline for better performance, as well as accelerating operations with RTC and using pinned memory to avoid unnecessary memory copies. Additionally, the user modified code in the `elemwise_binary_op.h`, and `elemwise_sum.cc` files, which suggests work related to core computational elements.
pythonschedulerdataflowmutationdata-science
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.