Rasmus Larsen is a Distinguished Engineer based in San Jose with over two decades of experience spanning numerical linear algebra, inverse problems, machine learning, and high-performance software for hardware accelerators. After a long tenure at Google Research where he advanced ML compilers and buffer assignment in XLA, he now leads linear algebra library work at NVIDIA, bridging algorithmic theory with production-grade performance engineering. His open-source contributions to projects like LFortran, Eigen, and XLA reflect deep expertise in compiler intrinsics, tensor operations, and memory liveness analysis that improve both correctness and speed. Trained as a PhD computer scientist with postdoctoral work at Stanford on inverse problems in helioseismology, he combines academic rigor with practical systems delivery. Colleagues describe him as someone who surfaces subtle numerical and performance pitfalls early—often by refactoring or renaming core primitives to avoid cross-library collisions—making complex algorithms reliably fast on modern hardware.
10 years of coding experience
25 years of employment as a software developer
Doctor of Philosophy - PhD Computer Science, Doctor of Philosophy - PhD Computer Science at Aarhus University
post doc Computer Science / Solar physics, post doc Computer Science / Solar physics at Stanford University
THIS MIRROR IS DEPRECATED -- New url: https://gitlab.com/libeigen/eigen
Role in this project:
Back-end Developer
Contributions:231 commits in 1 year 6 months
Contributions summary:Rasmus's commits primarily involved implementing and refining features within the Eigen library, specifically related to tensor operations. They developed vectorized clip functors for Eigen Tensors, enabling efficient clipping operations. Furthermore, the user refactored code to use `numext::maxi` and `numext::mini`, preventing collisions and implementing a rename from `scalar_clip_op` to `scalar_clamp_op` to prevent conflicts with other libraries. Additional contributions were focused on correcting potential performance bottlenecks in the library.
A machine learning compiler for GPUs, CPUs, and ML accelerators
Role in this project:
Backend Developer
Contributions:29 commits in 1 year 2 months
Contributions summary:Rasmus primarily made changes to the buffer assignment, specifically in the XLA (Accelerated Linear Algebra) service. These changes involved modifications to buffer allocation, buffer assignment for computations, and related code related to the liveness analysis of buffers. Further commits reveal internal modifications to the test suite, particularly within the XLA tests. The user also contributed to modifying fast math flags and fixing errors in the shape utility within the XLA service.
compilermachine-learning
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.