EuroEval Dutch

How well a model handles Dutch, as the unweighted mean of EuroEval's primary metric over its 9 Dutch datasets (dbrd, conll_nl, dutch_cola, squad_nl, wiki_lingua_nl, include_nl, multiloko_nl, winogrande_nl, duidelijke_taal). The primary metric differs by task — MCC for classification, knowledge and common-sense reasoning, micro-F1 without MISC for entity recognition, F1 for reading comprehension, chrF++ for summarization, METEOR for simplification — and is always the first of the two the board prints. Only models measured on every dataset are included, so each mean covers the same ground. Most rows are the validation split; some older open-weight rows are the test split, which EuroEval ranks in the same table.

ModelsModels with a result.154
Top resultBest result on this test.74.7%GPT-6 Astra
Top-3 spreadPoints from first to third.3.0 pts
YearYear of release.N/A

Result and price

020406080100$0.03$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

161 results
#ModelResultThe score from the source.
01GPT-6 Astra74.7% ±0.3
02Gemini 3.7 Flash73.1% ±0.4
03Gemini 3.6 Flash71.8% ±0.3
04GLM-5.3-Flash69.9% ±0.5
06GPT-5.6 Sol68.7% ±0.4
08GPT-5.6 Terra65.0% ±0.5
09Qwen3.6 27B FP864.0% ±0.4
11Qwen3.5 35B A3B FP863.3% ±0.6
14Gemma 4 31B62.7% ±0.4
15GPT-4.162.6% ±0.5
17Qwen3.8-27B62.1% ±0.5
21GPT-5 mini60.1% ±0.4
22Claude Sonnet 4.659.7% ±0.4
24GPT-5.6 Luna59.3% ±0.6
25GPT-OSS 120B57.8% ±0.4
29Gemma 4 12B IT57.3% ±0.5
32Gemma 3 27B Pt56.5% ±0.4
33Gemma 4 26B A4B56.4% ±0.5
35GPT-5.256.2% ±0.5
36Qwen3.5 9B Base55.6% ±0.7
42Qwen3 4B Thinking53.8% ±0.4
44GPT-5 nano53.7% ±0.5
47Claude Haiku 4.553.4% ±0.5
48Gemma 3 12B Pt53.4% ±0.4
49GLM-4.7-Flash53.3% ±0.6
53Gemma 3 12B52.4% ±0.3
55Qwen3.5 9B52.2% ±0.5
56Apertus V1.5 70B52.2% ±0.6
57Gemma 4 31B51.8% ±0.7
63Ministral 3 8B Base50.1% ±0.7
66Gemma 4 E4B IT49.8% ±0.5
69Qwen3.5 4B Base48.3% ±0.9
71EuroLLM 22B46.7% ±0.6
74Nesso 4B45.8% ±0.3
77Aya Expanse 8B45.5% ±0.7
78Gemma 4 26B A4B45.3% ±0.7
79Gemma 4 12B44.9% ±0.7
80Qwen3.5 4B44.6% ±0.7
82Domyn Small V1.044.3% ±0.6
83Llama 3.1-8B43.8% ±0.4
84Gemma 4 E2B IT43.1% ±0.5
85EuroLLM 9B43.0% ±0.5
86Gemma 3 4B Pt41.6% ±0.4
89Exaone 4.0 32B41.2% ±0.5
93Qwen3.5 2B40.2% ±0.5
94Luciole 23B Base40.0% ±0.6
95Gemma 3 4B39.8% ±0.6
96Apertus V1.1 4B38.4% ±0.6
99SmolLM3 3B Base37.4% ±0.4
102DFM Mimir36.5% ±0.5
103LFM2.5 8B A1B Base36.3% ±0.7
105Qwen3.5 2B Base35.7% ±0.7
106Occiglot 7B Eu535.2% ±0.4
110Luciole 8B Base33.6% ±0.5
112Gemma 4 E4B33.2% ±0.8
113Aya 23 8B33.1% ±0.6
116Qwen3.5 0.8B31.9% ±0.5
121Lizzy 7B28.8% ±0.5
124Apertus V1.1 1.5B27.1% ±0.5
125Qwen3.5 0.8B Base27.0% ±0.7
127Salamandra 7B26.2% ±0.3
128Llama 3.2 1B25.3% ±0.4
129Qwen3 0.6B24.8% ±0.4
131Gemma 3 1B Pt24.4% ±0.6
132YuLan Mini Instruct24.3% ±0.4
134Gemma 4 E2B24.1% ±0.6
135Apertus V1.1 0.5B23.7% ±0.8
136Luciole 1B Base23.3% ±0.7
139Gemma 3 1B IT21.5% ±0.3
140LFM2.5-8B-A1B21.1% ±0.6
141Pleias 1.2B Preview20.8% ±0.6
142LFM2 1.2B18.2% ±0.4
144Llama 3.2 1B Instruct15.0% ±0.5
145SmolLM2 360M14.2% ±0.5
146Pleias 350M Preview14.0% ±0.5
147Gemma 3 270M13.6% ±0.6
148LFM2 700M13.3% ±0.3
151LFM2 350M10.4% ±0.4
152Ling Mini 2.09.9% ±0.6
153SmolLM2 135M9.9% ±0.4
153Gemma 3 270M IT9.9% ±0.5
159Monad2.1% ±0.2
160Muse 3B0.4% ±0.4