EuroEval Portuguese

How well a model handles Portuguese, as the unweighted mean of EuroEval's primary metric over its 8 Portuguese datasets (sst2_pt, harem, scala_pt, multi_wiki_qa_pt, publico, alba_mcq_pt, cultura_viva_pt, winogrande_pt). The primary metric differs by task — MCC for classification, knowledge and common-sense reasoning, micro-F1 without MISC for entity recognition, F1 for reading comprehension, chrF++ for summarization, METEOR for simplification — and is always the first of the two the board prints. Only models measured on every dataset are included, so each mean covers the same ground. Most rows are the validation split; some older open-weight rows are the test split, which EuroEval ranks in the same table.

ModelsModels with a result.167
Top resultBest result on this test.74.7%GPT-6 Astra
Top-3 spreadPoints from first to third.0.6 pts
YearYear of release.N/A

Result and price

020406080100$0.03$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

175 results
#ModelResultThe score from the source.
01GPT-6 Astra74.7% ±0.4
02Gemini 3.7 Flash74.4% ±0.5
03GLM-5.3-Flash74.1% ±0.6
04Gemini 3.6 Flash72.6% ±0.3
05GPT-5.6 Sol70.5% ±0.4
07Qwen3.6 27B FP868.7% ±0.5
08Qwen3.8-27B68.5% ±0.5
09Gemma 4 31B67.5% ±0.6
10Claude Sonnet 4.667.0% ±0.5
11Qwen3.5 35B A3B FP865.5% ±0.6
12GPT-5.6 Terra65.4% ±0.5
19GPT-5.262.5% ±0.4
20Gemma 4 26B A4B62.0% ±0.5
23GPT-5.6 Luna60.7% ±0.4
24Gemma 4 12B IT60.7% ±0.6
27Llama 3.1-70B59.3% ±0.4
28GPT-OSS 120B59.3% ±0.5
33Gemma 3 27B Pt57.7% ±0.4
36Qwen3.5 9B Base57.1% ±0.8
40GPT-5 mini55.9% ±0.4
43Qwen3 4B Thinking55.6% ±0.3
46Gemma 4 31B55.3% ±0.7
47Qwen3.5 9B55.0% ±0.6
48GLM-4.7-Flash54.9% ±0.8
49GPT-4.154.4% ±0.5
51Ministral 3 8B Base53.7% ±0.8
52Gemma 3 27B53.5% ±0.4
53Claude Haiku 4.553.4% ±0.6
56Gemma 3 12B Pt52.6% ±0.4
57Exaone 4.0 32B52.4% ±0.6
59Apertus V1.5 70B51.9% ±0.5
60Gemma 4 E4B IT51.5% ±0.7
63Nesso 4B49.6% ±0.4
67Gemma 3 12B49.3% ±0.3
68GPT-5 nano49.0% ±0.6
70Qwen3.5 4B48.7% ±0.6
73Qwen3.5 4B Base47.9% ±0.9
75Domyn Small V1.046.5% ±0.7
77Gemma 4 12B45.5% ±0.8
80EuroLLM 22B44.9% ±0.7
88Gemma 4 26B A4B43.5% ±0.9
89Gemma 4 E2B IT43.5% ±0.6
91Aya Expanse 8B43.4% ±0.6
95Llama 3.1-8B41.1% ±0.5
96Qwen3.5 2B41.0% ±0.6
97DFM Mimir40.7% ±0.5
99AMALIA 9B SFT40.3% ±0.6
100AMALIA 9B DPO39.3% ±0.6
102SmolLM3 3B Base38.3% ±0.5
103EuroLLM 9B38.1% ±0.4
105Gemma 3 4B37.7% ±0.5
108Gemma 3 4B Pt37.5% ±0.4
109Luciole 23B Base37.1% ±0.7
111LFM2.5 8B A1B Base36.7% ±0.6
113Apertus V1.1 4B36.0% ±0.7
115Qwen3.5 0.8B34.6% ±0.6
116Olmo 3 32B33.6% ±0.4
117Llama 3.2 3B Instruct33.6% ±0.4
119Occiglot 7B Eu532.4% ±0.5
123Qwen3.5 2B Base31.3% ±0.9
126Olmo 3 7B30.9% ±0.4
127Olmo 3 7B Instruct30.6% ±0.3
130Luciole 8B Base28.9% ±0.8
131Aya 23 8B28.4% ±0.9
132YuLan Mini Instruct27.9% ±0.4
133Gemma 4 E4B27.6% ±0.9
136Apertus V1.1 1.5B26.2% ±0.8
138LFM2 1.2B24.4% ±0.4
140Gemma 3 1B IT23.4% ±0.4
141Salamandra 7B23.2% ±0.5
142Qwen3.5 0.8B Base23.2% ±0.8
143Gemma 3 1B Pt22.7% ±0.4
145LFM2.5-8B-A1B22.0% ±1.0
147LFM2 700M20.4% ±0.4
148Apertus V1.1 0.5B20.1% ±0.6
149Llama 3.2 1B19.6% ±0.5
150Gemma 4 E2B19.5% ±0.8
151Pleias 1.2B Preview17.7% ±0.7
152Luciole 1B Base17.3% ±0.7
153LFM2 350M16.9% ±0.5
154Pleias 3B Preview15.4% ±0.8
157SmolLM2 360M14.1% ±0.7
158Llama 3.2 1B Instruct14.0% ±0.6
159Nesso 0.4B Agentic13.7% ±0.6
160Ling Mini 2.012.0% ±0.7
161Gemma 3 270M11.9% ±1.0
164Gemma 3 270M IT9.1% ±0.8
167Zagreus 0.4B Ita6.9% ±0.6
168SmolLM2 135M6.8% ±0.4
173Monad1.5% ±0.3
174Muse 3B1.0% ±0.3