Goran Flegar

Munich, Bavaria, United Kingdom
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
Goran Flegar is a backend compiler and ML systems engineer with 11 years of experience focused on making large language models and GPU workloads faster and more efficient. Based in Munich, he contributes to high-impact open-source projects—improving ML compilers like XLA and Triton and adding Triton lowering and kernel pipelines into TensorFlow—to squeeze more performance out of GPUs with fewer watts. He specializes in compiler passes, tiling, vectorization and GPU launch lowering, frequently fixing subtle correctness and optimization issues that enable production-grade kernel generation. Notably, his work spans both research-adjacent transformations and pragmatic backend fixes, bridging the gap between ML algorithms and fast, deployable GPU code.
code11 years of coding experience
languagesCroatian, English
stackoverflow-logo

Stackoverflow

Stats
1,583reputation
36kreached
34answers
2questions
github-logo-circle

Github Skills (28)

c-language10
back-end-development10
llvm10
gpu-programming10
machine-learning10
vectorization10
mlr10
tiling10
triton10
deep-learning10
tensorflow10
gpu10
compiler-optimization10
compiler10
backend10

Programming languages (8)

TypeScriptC++ShellCSSLLVMCMakeMLIRPython

Github contributions (5)

github-logo-circle
triton-lang/triton

Jan 2023 - Jan 2023

Development repository for the Triton language and compiler
Role in this project:
userBackend Engineer
Contributions:29 reviews, 3 commits, 60 PRs in 1 day
Contributions summary:Goran primarily contributed to the backend aspects of the Triton language and compiler. They fixed build issues, correctly propagated address spaces through GEP, and fixed invalid casts related to the optimizer, showcasing a focus on compiler optimization and correctness. Further contributions included vectorizing s8 to bf16 casts, and fixing crashes related to boolean reductions, demonstrating performance improvements and bug fixes within the compiler's backend. Several commits address updates to LLVM and dependency bumps.
compilerprogramming-languagecode-generationtriton
openxla/xla

Aug 2022 - Dec 2022

A machine learning compiler for GPUs, CPUs, and ML accelerators
Role in this project:
userBack-end Developer
Contributions:3 reviews, 34 commits, 6 comments in 3 months
Contributions summary:Goran primarily focused on improving the machine learning compiler by cleaning up and refactoring the collapse-parallel-loops pass. Their contributions involved removing function restrictions, eliminating canonicalization in tests, and adding a new pass to convert structured GmlSt tiling into gpu.launch. Furthermore, they addressed imperfect tiling scenarios and resolved issues related to the GmlSt-to-GPU pass, ultimately making the code more efficient and robust. The user also handled vectorization in multiple areas to support the gml_st for gpu.launch process.
compilercommunity-drivenmachine-learningmodular
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial