Przemyslaw P

Deep Learning Algorithms Manager at NVIDIA

Warsaw, Masovian Voivodeship, Poland
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Przemyslaw P is a Deep Learning Algorithms Manager at NVIDIA with 11 years of experience designing and delivering inference optimization, deployment tooling and end-to-end model pipelines such as PyTriton and Triton Model Navigator. He combines hands-on performance engineering—contributing to high-profile open-source projects like TVM, MXNet, NVIDIA DALI and TransformerEngine—with people and product leadership honed at Samsung and deepsense.ai. His core strengths are GPU-aware memory and data pipeline optimizations, deterministic pipeline construction, and fixing subtle concurrency and graph-compilation bugs that improve real-world inference stability and throughput. He brings an uncommon blend of embedded, backend and DevOps experience, which helps bridge low-level CUDA/kernel work with cloud deployment concerns. Przemyslaw also holds an MBA and formal training in systems analysis, enabling him to translate technical tradeoffs into pragmatic product decisions. Based in Warsaw, he’s an active contributor to foundational ML infra where his commits have touched performance-critical codepaths used by many downstream projects.
code11 years of coding experience
job13 years of employment as a software developer
bookBachelor of Engineering (B.Eng.), Information Technology, Bachelor of Engineering (B.Eng.), Information Technology at Politechnika Świętokrzyska w Kielcach
bookMaster of Business Administration (MBA), MBA, Master of Business Administration (MBA), MBA at Akademia Leona Koźmińskiego
bookWarsaw University of Technology
languagesEnglish
stackoverflow-logo

Stackoverflow

Stats
326reputation
5kreached
5answers
0questions
github-logo-circle

Github Skills (41)

graph-algorithms10
pytorch10
c-language10
cublas10
python10
multithreading10
memory-management10
machine-learning10
dm10
deeplearning-ai10
p810
compiler-compiler10
deep-learning10
gpu10
optmization10

Programming languages (5)

C++CJupyter NotebookVim ScriptPython

Github contributions (5)

github-logo-circle
NVIDIA/TransformerEngine

Sep 2022 - Jan 2023

A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper and Ada GPUs, to provide better performance with lower memory utilization in both training and inference.
Role in this project:
userML Engineer
Contributions:15 releases, 701 reviews, 25 commits in 3 months
Contributions summary:The user, Przemek Tredak, primarily contributed to the development and improvement of the `transformerengine` library, focusing on performance and functionality. Their work includes implementing new features, such as linking to documentation archives and adding pylint to the lint action. Przemek's contributions also involved fixing critical issues in the codebase, like out-of-bounds memory access in the C+T+dbias kernel. Significant changes were made to enhance the library's functionality with FP8 precision and integration with NVTX for improved debugging and profiling.
better-performancememorypythontrainingbit
apache/mxnet

Mar 2017 - Jan 2022

Lightweight, Portable, Flexible Distributed/Mobile Deep Learning with Dynamic, Mutation-aware Dataflow Dep Scheduler; for Python, R, Julia, Scala, Go, Javascript and more
Role in this project:
userBack-end Developer & Performance Engineer
Contributions:1 release, 100 reviews, 151 commits in 4 years 10 months
Contributions summary:Przemyslaw primarily focused on optimizing the performance of the MXNet deep learning framework. Their contributions involved modifying GPU memory allocation, particularly in the `pooled_storage_manager.h` and `pooled_storage_manager.cc` files. They also worked on improving the image I/O pipeline for better performance, as well as accelerating operations with RTC and using pinned memory to avoid unnecessary memory copies. Additionally, the user modified code in the `elemwise_binary_op.h`, and `elemwise_sum.cc` files, which suggests work related to core computational elements.
pythonschedulerdataflowmutationdata-science
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial