Summary
Felipe Soares is a Senior Research Scientist at NVIDIA with nine years of experience building post-training data pipelines, alignment workflows, and evaluation frameworks for large language models such as Nemotron. His work spans instruction tuning, preference data collection, human evaluation design, and large-scale curation with an emphasis on making models reliable and release-ready. Before focusing on LLMs he spent years in multilingual and biomedical NLP across academia and industry, which shaped his rigor around data quality, careful evaluation, and domain-sensitive modeling. He has published 35+ papers in top venues including NeurIPS, ACL, ICLR and The Lancet, and holds a PhD plus two research master’s degrees. Known for operationalizing human-in-the-loop evaluation and synthetic data practices, he bridges research, product, legal, and vendor stakeholders to turn experimental datasets into production-grade assets. Based in New York, he combines deep academic credentials with pragmatic engineering to improve model behavior in high-stakes settings.
9 years of coding experience
8 years of employment as a software developer
Doctor of Philosophy - PhD, Engineering, Doctor of Philosophy - PhD, Engineering at Federal University of Rio Grande do Sul
Doctor of Philosophy - PhD, Doctor of Philosophy - PhD at The University of Sheffield
Bachelor of Engineering (BEng), Engineering/Industrial Management, Bachelor of Engineering (BEng), Engineering/Industrial Management at University of Exeter
Portuguese, English, Spanish, Catalan