Ivan Yashchuk is a Software Engineering Manager at NVIDIA with a decade of experience building high-performance numerical and ML systems, especially GPU-accelerated libraries and PyTorch internals. He combines hands-on backend development—contributions to CuPy, PETSc, Stan Math, Firedrake and CUDA.jl—with leadership in productionizing nvFuser and core PyTorch ops. His work spans automatic differentiation, batched linear algebra, and rigorous test automation, reflecting a rare mix of numerical math, low-level performance tuning, and software quality engineering. Trained across computer science, mathematics and mechanical engineering (PhD/masters), he moves fluidly between research and production, having surfaced subtle edge-case fixes (NaN/Inf/empty-matrix handling) and performance wrappers used by GPU toolchains. Based in the Basel area, he blends deep academic background with practical open-source impact on widely used scientific and ML projects.
9 years of coding experience
8 years of employment as a software developer
Doctor of Philosophy - PhD Computer Science, Doctor of Philosophy - PhD Computer Science at Aalto University
Bachelor’s Degree Mechanical Engineering and Production Technology, Bachelor’s Degree Mechanical Engineering and Production Technology at Häme University of Applied Sciences, HAMK
Master's degree Mathematics, Master's degree Mathematics at University of Helsinki
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Role in this project:
Back-end Developer & ML Engineer
Contributions:1247 reviews, 476 commits, 296 PRs in 2 years 6 months
Contributions summary:Ivan contributed to the reference implementations of several PyTorch operations, including `rot90`, `roll`, and `atleast_1d/2d/3d`, demonstrating an understanding of tensor manipulation and core PyTorch functionality. They added implementations for the `softmax`, `log_softmax`, and `logsumexp` functions. They also worked on integrating the `nvFuser` executor, which suggests experience with performance optimization and potentially compiler-level details. This included implementing `nvFuser` support for functions like `var`, and ensuring compatibility with AOT Autograd and PyTorch's tracing capabilities.
A flexible framework of neural networks for deep learning
Role in this project:
Back-end Developer & ML Engineer
Contributions:495 commits, 13 PRs, 130 comments in 6 months
Contributions summary:Ivan primarily contributed to the Chainer deep learning framework by implementing new trigonometric functions (tan, arcsin, arccos, arctan) for both the CUDA and native devices. They integrated these functions with appropriate pybind bindings and implemented the necessary gradient calculations, thus enabling automatic differentiation. Furthermore, they declared and began the implementation of Cholesky and QR kernels for the CUDA device.
cudapythonmxnetcaffe2flexible-framework
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.