EuroEval Spanish

How well a model handles Spanish, as the unweighted mean of EuroEval's primary metric over its 8 Spanish datasets (sentiment_headlines_es, conll_es, scala_es, mlqa_es, mlsum_es, include_es, multiloko_es, winogrande_es). The primary metric differs by task — MCC for classification, knowledge and common-sense reasoning, micro-F1 without MISC for entity recognition, F1 for reading comprehension, chrF++ for summarization, METEOR for simplification — and is always the first of the two the board prints. Only models measured on every dataset are included, so each mean covers the same ground. Most rows are the validation split; some older open-weight rows are the test split, which EuroEval ranks in the same table.

ModelsModels with a result.161
Top resultBest result on this test.70.0%GPT-6 Astra
Top-3 spreadPoints from first to third.2.6 pts
YearYear of release.N/A

Result and price

020406080100$0.03$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

169 results
#ModelResultThe score from the source.
01GPT-6 Astra70.0% ±0.4
02Gemini 3.7 Flash68.3% ±0.4
03GLM-5.3-Flash67.4% ±0.6
04Gemini 3.6 Flash66.8% ±0.4
06GPT-5.6 Sol66.3% ±0.4
08GPT-5.6 Terra62.1% ±0.4
09Qwen3.6 27B FP861.4% ±0.6
13Qwen3.8-27B59.0% ±0.6
14Claude Sonnet 4.658.5% ±0.5
16GPT-5.6 Luna57.9% ±0.4
17Qwen3.5 35B A3B FP857.6% ±0.7
22Gemma 4 31B55.8% ±0.7
24GPT-5 mini54.2% ±0.6
27GPT-5.253.5% ±0.4
28GPT-OSS 120B53.5% ±0.4
32GPT-4.152.8% ±0.5
34Qwen3 4B Thinking51.5% ±0.4
39Claude Haiku 4.549.7% ±0.4
40Gemma 4 26B A4B49.4% ±0.5
41GLM-4.7 Flash NVFP449.4% ±0.6
42Gemma 3 27B49.2% ±0.4
43GPT-5 nano49.0% ±0.5
47Exaone 4.0 32B48.3% ±0.7
49Apertus V1.5 70B47.3% ±0.6
52Qwen3.5 9B46.1% ±0.7
55Gemma 3 12B45.1% ±0.4
56Gemma 4 12B IT45.0% ±0.6
60Gemma 3 27B Pt43.8% ±0.4
61Nesso 4B43.7% ±0.3
66Gemma 4 31B40.8% ±0.8
69Ministral 3 8B Base40.2% ±0.6
70Qwen3.5 9B Base40.1% ±0.5
73Aya Expanse 8B39.2% ±0.5
74Gemma 3 12B Pt39.1% ±0.4
78Gemma 4 E4B IT37.5% ±0.6
79Gemma 4 E2B IT37.4% ±0.7
81Qwen3.5 4B36.8% ±0.7
82Domyn Small V1.036.4% ±0.7
85Qwen3.5 4B Base35.4% ±0.6
86Gemma 4 26B A4B35.4% ±0.6
87Olmo 3 32B35.1% ±0.3
88EuroLLM 9B34.8% ±0.4
89AMALIA 9B DPO34.7% ±0.5
92AMALIA 9B SFT34.2% ±0.6
96EuroLLM 22B32.7% ±0.4
98Gemma 4 12B32.3% ±0.8
99Gemma 3 4B31.9% ±0.4
101Qwen3.5 2B31.5% ±0.5
103Llama 3.1-8B31.3% ±0.5
104Olmo 3 7B30.9% ±0.4
105SmolLM3 3B Base30.8% ±0.4
107Gemma 3 4B Pt30.1% ±0.4
108Llama 3.2 3B Instruct29.6% ±0.4
109Olmo 3 7B Instruct29.1% ±0.3
110Occiglot 7B Eu528.8% ±0.4
112DFM Mimir27.9% ±0.6
115Apertus V1.1 4B27.5% ±0.6
117Qwen3.5 2B Base27.1% ±0.6
118Luciole 23B Base26.8% ±0.7
119Qwen3.5 0.8B26.5% ±0.7
123LFM2.5-8B-A1B24.1% ±1.0
124Gemma 4 E4B23.6% ±0.7
125Luciole 8B Base22.4% ±0.9
127LFM2 1.2B22.3% ±0.4
128TildeOpen 30B22.0% ±0.8
129Qwen3.5 0.8B Base21.8% ±0.6
131YuLan Mini Instruct20.5% ±0.5
134Apertus V1.1 1.5B19.0% ±0.7
136Aya 23 8B18.1% ±0.6
137LFM2 700M17.8% ±0.3
139Salamandra 7B17.2% ±0.3
140Gemma 3 1B Pt17.1% ±0.4
142Gemma 3 1B IT16.2% ±0.3
143Apertus V1.1 0.5B15.9% ±0.6
144Luciole 1B Base15.3% ±0.8
145Llama 3.2 1B15.1% ±0.5
146Pleias 3B Preview14.9% ±0.5
147Gemma 4 E2B14.6% ±1.1
148Pleias 1.2B Preview14.6% ±0.6
149LFM2 350M14.0% ±0.3
151Nesso 0.4B Agentic12.4% ±0.6
152Llama 3.2 1B Instruct12.3% ±0.4
153SmolLM2 360M12.2% ±0.4
154Ling Mini 2.011.3% ±0.6
155Gemma 3 270M10.8% ±0.5
157Pleias 350M Preview10.3% ±0.4
158MiniCPM5-1B9.7% ±0.7
159SmolLM2 135M9.1% ±0.4
162Gemma 3 270M IT7.6% ±0.5
164Zagreus 0.4B Ita4.9% ±0.3
167Monad1.5% ±0.3
168Muse 3B0.8% ±0.4