Cong Xu is an Applied Scientist in San Francisco with nine years of experience building high-performance distributed ML and HPC systems, currently working at Amazon after a deep-learning software engineering tenure at Intel. He combines PhD-level research in parallel I/O and MPI with hands-on expertise in distributed LLM training (Megatron, FSDP, DeepSpeed), NCCL/RDMA optimizations, and cluster administration for GPU and Infiniband environments. Proficient in C/C++, Python and Java, he bridges systems and ML stacks—from Lustre/MPI-IO and HDF5 to PyTorch/TensorFlow—enabling scalable training and big-data analytics pipelines. His background includes published research on scalable MPI AlltoallV algorithms and practical experience shipping MapReduce/Hive projects and full-stack data collection tools. Colleagues rely on him to translate low-level performance engineering into robust, production-ready ML workflows.
9 years of coding experience
11 years of employment as a software developer
Bachelor's degree, Computer Science, Bachelor's degree, Computer Science at Beijing University of Post and Telecommunications
Master's degree, Computer Science, Master's degree, Computer Science at Auburn University
An open source AutoML toolkit for automate machine learning lifecycle, including feature engineering, neural architecture search, model compression and hyper-parameter tuning.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.