Summary
Jaegeun Han is an AI Solutions Architect with over a decade of hands-on experience designing and optimizing large-scale GPU-accelerated AI systems for data centers and production LLM pipelines. Currently at AMD, he architects and validates ROCm-based deployments and leads cross-functional programs to boost AI workload performance across GPU, CPU, and DPU platforms. Previously at NAVER he built and stabilized HyperCLOVA X production pipelines, scaled 2,000+ GPU clusters, and delivered long-context and LoRA inference enhancements for FasterTransformer. His background spans deep CUDA optimization, TensorRT tuning, and on-prem GPU cluster deployments from roles at NVIDIA, Samsung, and research labs, with open-source contributions to CUDA educational repos showing practical SGEMM and convolution optimizations. Based in Seoul, he combines systems-level engineering with customer-facing enablement to turn cutting-edge AI research into scalable, production-ready solutions.
10 years of coding experience