Anton Malakhov is a seasoned software engineer with over a decade of experience in high-performance C++ systems, specializing in multi-threaded and parallel programming and performance optimization. During a long tenure at Intel he moved from senior software roles to ML performance engineering, contributing patented ideas that made it into products and demonstrating a willingness to take calculated technical risks. He’s an active open-source contributor to prominent projects in the LLVM/Numba ecosystem, adding Intel SVML auto-vectorization support and TBB-based threading to boost JIT and compiler performance. Based in Nizhniy Novgorod, Anton blends deep low-level expertise with practical delivery, often tackling thorny concurrency and vectorization edge cases that improve real-world throughput.
11 years of coding experience
23 years of employment as a software developer
MS (Specialist) Microprocessor Design and Computer Science, MS (Specialist) Microprocessor Design and Computer Science at ОмГТУ
A lightweight LLVM python binding for writing JIT compilers
Role in this project:
Back-end Developer / Performance Engineer
Contributions:8 commits, 3 PRs, 11 comments in 8 months
Contributions summary:Anton primarily focused on optimizing the performance of the `llvmlite` project by enabling Intel SVML-enabled auto-vectorization for transcendental math functions. They introduced a patch to integrate SVML functionality, including changes to the LLVM analysis and IR components. Furthermore, the user addressed a specific issue related to loop vectorization, adding a workaround and a test case to ensure correct behavior. Finally, the user updated the code with the latest version, which included additional bug fixes.
Contributions:43 commits, 9 PRs, 61 comments in 1 year 11 months
Contributions summary:Anton focused on implementing multi-threading support within the Numba compiler, specifically integrating the Intel TBB (Threading Building Blocks) library. Their contributions involved adding new threading capabilities, modifying build scripts to include TBB, and addressing potential issues related to forking in multi-threaded environments. Additionally, the user made changes to enable Intel SVML optimizations for transcendental functions and addressed code quality concerns, demonstrating a focus on performance and optimization.
compilerllvmnumpypythoncuda
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.