Code for the ICML 2021 (long talk) paper: "ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision"
Role in this project:
ML Engineer Contributions:1 release, 18 commits, 1 PR in 9 months
Contributions summary:Wonjun primarily focused on implementing and improving demo applications for the ViLT model, specializing in vision-and-language tasks. They added interactive demos for image captioning and visual question answering (VQA), including the integration of Gradio for user interfaces. Furthermore, they updated the VQA demo by incorporating example prompts and refactoring URLs for data retrieval, demonstrating a focus on model usability and deployment. They also contributed to the model's utility code, modifying the training tasks.
convolutionvision-and-language
Contributions:156 pushes, 5 branches in 6 years 10 months
kim