Brian Shi is a results-driven mechanical supplier quality leader with 11 years of experience building and scaling SQE teams for advanced consumer and data center products across Apple and NVIDIA. He specializes in thermal, mechanical, optical, and packaging quality for high-volume NPI and next-generation AI infrastructure, combining hands-on failure analysis with program-level supplier management. Based in San Jose, he has moved from manufacturing process bring-up at Intuitive Surgical to leading cross-commodity quality strategies for iPhone, Apple Watch, and NVIDIA interconnects and optics. He pairs technical depth in brittle materials and automated inspection with a pragmatic, hiring-focused management style that grows high-performing teams. Less obvious: he contributes to NVIDIA’s DeepLearningExamples ecosystem by improving TensorRT integration and transformer quantization, signaling a rare blend of hardware-quality expertise and familiarity with ML inference tooling. He holds a B.S. in Biomedical Engineering from UC Irvine and thrives on solving complex, multidisciplinary production challenges.
State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.
Role in this project:
ML Engineer
Contributions:39 commits, 37 PRs, 29 pushes in 1 year 3 months
Contributions summary:Bhsueh_NV primarily contributes to the `FasterTransformer` project, focusing on integrating and fixing issues related to TensorRT plugins. They implement and debug features specifically related to the Transformer models' execution using TensorRT for accelerated inference. Additionally, they work on model quantization techniques for improved performance and efficiency, evidenced by modifications to quantization-related scripts and model components.
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
Contributions:6 PRs, 227 pushes, 91 branches in 2 years 5 months
cppinferencelarge-language-modelsllmnvidia
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.