Summary
Andreea-maria Oncescu is a Senior Research Scientist based in Oxford with nine years of experience building multimodal AI systems, now focusing at Huawei on unifying audio (speech, sounds, music) with vision-language models using open-source architectures like Llama, ViT, and CLIP. She completed a DPhil at Oxford’s Visual Geometry Group where she advanced video and audio understanding, retrieval, and captioning, and has hands-on experience taking foundation models into production and end-to-end data pipelines. Her background includes industry collaborations and internships at Meta on self-supervised multimodal video understanding and text-to-audio retrieval, reflecting a strong blend of academic rigor and product-focused research. Known for bridging sensor- and web-integrated ML prototypes to large-scale multimodal LLMs, she brings both systems-level engineering (wearable prototypes, real-time monitoring) and cutting-edge model development to practical AI deployments.
9 years of coding experience
6 years of employment as a software developer
DPhil - Computer Vision and Natural Language, Engineering Science, DPhil - Computer Vision and Natural Language, Engineering Science at University of Oxford
Attend highly intensive weekly preparation sessions for International Olympiads in Physics and Astronomy, Attend highly intensive weekly preparation sessions for International Olympiads in Physics and Astronomy at International Computer High School of Bucharest
English, French, Romanian