EuroEval
EuroEval Portuguese
How well a model handles Portuguese, as the unweighted mean of EuroEval's primary metric over its 8 Portuguese datasets (sst2_pt, harem, scala_pt, multi_wiki_qa_pt, publico, alba_mcq_pt, cultura_viva_pt, winogrande_pt). The primary metric differs by task — MCC for classification, knowledge and common-sense reasoning, micro-F1 without MISC for entity recognition, F1 for reading comprehension, chrF++ for summarization, METEOR for simplification — and is always the first of the two the board prints. Only models measured on every dataset are included, so each mean covers the same ground. Most rows are the validation split; some older open-weight rows are the test split, which EuroEval ranks in the same table.
ModelsModels with a result.167
Top-3 spreadPoints from first to third.0.6 pts
YearYear of release.N/A
Result and price
Results
175 results#ModelResultThe score from the source.