Summary
Bryan De Oliveira is a Staff Research Scientist with a decade of experience advancing deep reinforcement learning and agent-centered AI across industry and academia from Goiás, Brazil. He leads interdisciplinary teams at CEIA/AKCIT, turning PoCs into production-grade systems for finance, energy, tourism, and agribusiness while mentoring junior researchers. His work spans sim-to-real humanoid control, policy-gradient theory (GRPO/PPO), embodied LLM planning, and information-seeking dialogue agents, with multiple arXiv publications and industry collaborations. Bryan blends hands-on MLOps and scalable ML systems engineering—BigQuery, Ray RLlib, BentoML, Seldon, Kubernetes—with fundamental research on agency, curiosity, and offline RL. He’s driven by building generalist agents that learn from experience and has a track record of shipping applied RL solutions for debt collection, marketing personalization, and ad optimization. An oft-overlooked strength is his ability to bridge deep theoretical research and pragmatic deployment pipelines, enabling rapid translation from experiments to billable APIs.
10 years of coding experience
12 years of employment as a software developer
Master of Science - MS, Computer Science, Master of Science - MS, Computer Science at Universidade Federal de Goiás
Portuguese, English