Akira Naruse

Principal Developer Technology Engineer at NVIDIA

Japan
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Akira Naruse is a Principal Developer Technology Engineer at NVIDIA with over 9 years of focused experience in high-performance ML frameworks and GPU-accelerated computing, following a long research tenure at Fujitsu Laboratories. He has driven backend and profiling enhancements in flagship open-source projects such as Chainer and cuPy, adding NVTX-based instrumentation, cuDNN support, and performance-oriented features that aid fine-grained analysis on NVIDIA hardware. His work on graph convolutional models for Chainer Chemistry shows hands-on ML engineering skills bridging research-grade models and production-ready frameworks. Known for shipping robust testing and low-level CUDA integrations, he blends deep systems knowledge with practical ML model implementation. Based in Japan, he brings a rare combination of research discipline and developer-technology leadership that improves both developer experience and runtime performance.
code10 years of coding experience
job27 years of employment as a software developer
book名古屋大学 / Nagoya University
github-logo-circle

Github Skills (23)

debugging10
python10
chainer10
machine-learning10
deep-learning10
cupy10
profiling10
neural-network10
cuda10
graph-convolutional-networks10
parallelization9
parallel9
parallel-computing9
performance-analysis9
parallel-processing9

Programming languages (3)

Jupyter NotebookPythonCuda

Github contributions (5)

github-logo-circle
cupy/cupy

Jun 2017 - Dec 2022

NumPy & SciPy for GPU
Role in this project:
userBackend Developer
Contributions:112 reviews, 715 commits, 138 PRs in 5 years 6 months
Contributions summary:Akira's commits focus on enhancing the NVIDIA Tools Extension Library (NVTX) within the cuPy project. Their primary contributions involved adding and modifying wrappers for the NVIDIA Tools Extension Library (NVTX), including adding support for cuDNN, and enabling more fine-grained performance analysis. They implemented functionality to measure the impact of different functions. The user also created a testing framework to validate these changes.
gpunumpyscipycudacudnn
chainer/chainer

Jul 2016 - Oct 2019

A flexible framework of neural networks for deep learning
Role in this project:
userBackend Developer
Contributions:236 commits, 56 PRs, 264 comments in 3 years 3 months
Contributions summary:Akira's commits primarily focused on enhancing the CUDA-based machine learning framework, Chainer. The contributions included implementing NVIDIA Tools Extension Library (NVTX) wrappers for profiling and debugging, expanding the framework's support for features like soft targets within the softmax_cross_entropy, along with support for layer normalization, and for more efficient sparse matrix computations. The changes spanned multiple files in the core library, and the testing, suggesting a focus on improving the framework's capabilities and performance.
deep-learningneural-networkpythonneural-networksmachine-learning
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial