Valentin Andrei

Senior Staff Performance Engineer at Meta

San Francisco Bay Area United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Valentin Andrei is a Senior Staff Performance Engineer based in the San Francisco Bay Area with a 15-year career focus on software performance, AI/HPC, and platform architecture spanning CPUs, GPUs and accelerators. He blends hands-on kernel-level optimization (notably CUDA improvements to core PyTorch ops) with systems thinking—driving full-stack code optimizations, benchmarks, simulators and workload modeling across large-scale distributed systems. A proven technical leader at Meta and Intel, he repeatedly bridges deep optimization work with team leadership and cross-product execution. His PhD work applying ML to multi-speaker speech analysis underscores a rare combination of research rigor and production-first engineering. Colleagues rely on him for hard performance digs that yield measurable speedups and cleaner, maintainable codepaths.
code9 years of coding experience
job12 years of employment as a software developer
bookPOLITEHNICA București National University for Science and Technology
languagesEnglish, French, Spanish, Romanian
stackoverflow-logo

Stackoverflow

Stats
1reputation
0reached
0answers
0questions
github-logo-circle

Github Skills (12)

cuda10
pytorch10
tensor10
deep-learning10
vectorization10
gpu10
performance-optimization10
neural-network9
machine-learning9
c-language8
python8
cprogramming-language8

Programming languages (4)

C++CHTMLPython

Github contributions (5)

github-logo-circle
pytorch/pytorch

Oct 2022 - Oct 2022

Tensors and Dynamic neural networks in Python with strong GPU acceleration
Role in this project:
userPerformance Engineer
Contributions:20 reviews, 2 commits, 27 PRs in 2 days
Contributions summary:Valentin primarily focused on optimizing the performance of CUDA kernels within the PyTorch library. Their contributions involved rewriting and improving existing kernels, such as those for layer normalization and depthwise convolutions, to leverage techniques like shared memory, warp shuffles, and vectorized memory loads. The user provided performance analyses and benchmarks demonstrating significant speedups, and also addressed code quality by removing dead code and cosmetic changes. These efforts aimed to improve the execution speed of core tensor operations.
pythongpu-accelerationdeep-learninggpunumpy
valentinandrei/pytorch

Oct 2022 - Aug 2024

Tensors and Dynamic neural networks in Python with strong GPU acceleration
Contributions:211 pushes, 5 branches in 1 year 10 months
pythongpu-accelerationdeep-learninggpuacceleration
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial