Sergey Kozub

Compiler Engineer at NVIDIA

Zurich, Zurich, Switzerland
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Sergey Kozub is a Compiler Engineer at NVIDIA with 3+ years of professional experience building high-performance systems across compilers, data pipelines and financial platforms. He has deep systems and numerical expertise in C++, Python and SQL, honed through roles at Google (scalable data processing, Bigtable, multithreading) and open-source contributions to flagship ML projects like TensorFlow and XLA where he optimized GPU convolutions, cuDNN vectorization and sparse operations. A practical full-stack engineer from backend systems to UI, he also founded and led data products at SuperbIntel, producing proprietary datasets and predictive models. Ranked Grandmaster on Kaggle, Sergey combines algorithmic rigor with production-grade engineering and a track record of fixing subtle numerical/metadata bugs that improve correctness and performance.
code3 years of coding experience
job21 years of employment as a software developer
bookMaster’s Degree, Computer Science, Master’s Degree, Computer Science at Technical University of Moldova
languagesEnglish, Russian, German
github-logo-circle

Github Skills (22)

c-language10
gpu-programming10
tensorflow10
cuda10
xla10
compiler10
cprogramming-language10
cudnn10
linear-algebra10
machine-learning9
hla8
tensorrt8
tensor8
operation8
sparse-array7

Programming languages (3)

C++StarlarkPython

Github contributions (5)

github-logo-circle
openxla/xla

Aug 2022 - Jan 2023

A machine learning compiler for GPUs, CPUs, and ML accelerators
Role in this project:
userBackend Developer
Contributions:37 reviews, 15 commits, 1 PR in 5 months
Contributions summary:Sergey contributed to the XLA compiler, focusing on GPU-related optimizations and bug fixes. Their work included preventing metadata loss during compilation steps, specifically within the cuDNN vectorization process. The commits also involved refactoring and improvement of the compiler's handling of 1x1 convolutions and scatter operations, improving efficiency and correctness. The user also made changes to the build system, improving its functionality.
compilercommunity-drivenmachine-learningmodular
tensorflow/tensorflow

Aug 2022 - Nov 2022

An Open Source Machine Learning Framework for Everyone
Role in this project:
userBack-end Developer & ML Engineer
Contributions:4 reviews, 15 commits, 6 comments in 3 months
Contributions summary:Sergey primarily worked on optimizing and improving the XLA compiler within the TensorFlow framework. Their contributions focused on enhancing the handling of integer types, specifically within the HloEvaluator. They implemented tests for cuDNN-based convolution operations, added reordering flags and custom calls for convolution inputs, and worked on the implementation of sparse dot operations, including addressing issues in the sparse dot metadata loading. The user's work involved low-level code modifications and optimizations related to numerical computation and the integration of hardware-specific (GPU) functionalities within the XLA framework.
pythondata-sciencedeep-learningmlmachine-learning
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial