Code and data to evaluate LLMs on the ENEM, the main standardized Brazilian university admission exams.
Contributions:14 PRs, 33 pushes, 15 branches in 1 year 9 months
artificial-intelligencellm-inferencellmsmultimodal-datasets
A framework for few-shot evaluation of autoregressive language models.
Contributions:17 PRs, 22 pushes, 11 branches in 9 months
language-model