Summary
Can Karakus is a machine learning and distributed systems engineer with seven years of experience building scalable training infrastructure and applied ML products in the San Francisco Bay Area. He has held research and production roles from academia to industry, culminating in applied-science and Member of Technical Staff positions at AWS, Contextual AI, and Anthropic. At AWS he architected and shipped model-parallel training libraries for SageMaker and led large-scale distributed training efforts; his PhD work at UCLA produced practical redundancy and straggler-mitigation techniques that doubled training speed in cluster settings. He combines rigorous theoretical background in electrical engineering with hands-on systems implementation across MPI, EC2, and production ML services. Colleagues would note his knack for turning research ideas into robust, deployable tooling that improves throughput and fault tolerance in real workloads. Based in the Bay Area, he brings a blend of academic depth and production engineering focus to large-model and distributed ML challenges.
7 years of coding experience
14 years of employment as a software developer
High School, High School at İzmir Fen Lisesi
BS Electrical&Electronics Engineering, BS Electrical&Electronics Engineering at Bilkent University
University of California, Los Angeles