EuroEval Polish

How well a model handles Polish, as the unweighted mean of EuroEval's primary metric over its 7 Polish datasets (polemo2, kpwr_ner, scala_pl, poquad, psc, llmzszl, winogrande_pl). The primary metric differs by task — MCC for classification, knowledge and common-sense reasoning, micro-F1 without MISC for entity recognition, F1 for reading comprehension, chrF++ for summarization, METEOR for simplification — and is always the first of the two the board prints. Only models measured on every dataset are included, so each mean covers the same ground. Most rows are the validation split; some older open-weight rows are the test split, which EuroEval ranks in the same table.

ModelsModels with a result.189
Top resultBest result on this test.68.1%GLM-5.3-Flash
Top-3 spreadPoints from first to third.0.7 pts
YearYear of release.N/A

Result and price

020406080100$0.03$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

198 results
#ModelResultThe score from the source.
01GLM-5.3-Flash68.1% ±0.4
02Gemini 3.7 Flash67.9% ±0.4
03GPT-6 Astra67.4% ±0.4
04Gemini 3.6 Flash65.9% ±0.4
05Qwen3.8-27B65.4% ±0.5
06Qwen3.6-27B65.3% ±0.5
07Qwen3.6 27B FP865.0% ±0.4
08Gemini 3 Pro64.8% ±0.5
10GPT-5.6 Sol62.8% ±0.4
12Qwen3.5 35B A3B FP862.3% ±0.6
14Gemma 4 31B61.6% ±0.4
18Claude Sonnet 4.659.9% ±0.5
20GPT-5.6 Terra59.0% ±0.5
24GPT-OSS 120B57.6% ±0.5
25GPT-5 mini57.2% ±0.3
26GPT-5.257.1% ±0.4
28Llama 3.1-70B56.5% ±0.5
33Gemma 3 27B Pt55.7% ±0.4
34Gemma 4 26B A4B55.5% ±0.7
37Gemma 4 12B IT55.2% ±0.6
38Voxtral Small 24B55.1% ±0.2
45GPT-5.6 Luna54.2% ±0.4
46GPT-4.154.1% ±0.5
48GLM-4.7-Flash53.5% ±0.6
51GPT-5 nano52.8% ±0.5
52Qwen3.5 9B52.7% ±0.5
56Gemma 3 27B51.8% ±0.3
57Qwen3 4B Thinking51.4% ±0.3
60Ministral 3 8B Base50.3% ±0.6
64Gemma 3 12B Pt49.7% ±0.5
65Claude Haiku 4.549.5% ±0.6
71Gemma 3 12B47.7% ±0.3
73Gemma 4 31B46.7% ±0.9
75GaMS3 12B46.3% ±0.8
76Gemma 4 E4B IT46.2% ±0.7
86PLLuM 12B Chat44.4% ±0.5
87Gemma 3N E4B IT44.3% ±0.3
88Qwen3.5 4B44.2% ±0.8
89PLLuM 12B Nc Chat44.2% ±0.6
90Qwen3.5 9B Base44.0% ±0.9
91Nesso 4B43.5% ±0.3
92Qwen3.5 4B Base43.5% ±0.7
94Gemma 4 12B43.1% ±0.7
96Domyn Small V1.041.7% ±0.5
97Llama PLLuM 8B Chat41.6% ±0.5
98Llama 3.1-8B41.4% ±0.4
99Exaone 4.0 32B41.3% ±0.6
103Aya Expanse 8B40.8% ±0.5
104Apertus V1.5 70B40.8% ±0.6
105Llama 3.1 8B Instruct40.8% ±0.3
106BgGPT Gemma 3 4B IT40.4% ±0.6
108EuroLLM 9B Instruct39.9% ±0.5
111Gemma 4 E2B IT38.9% ±0.7
113Gemma 4 26B A4B38.5% ±0.8
115Llama PLLuM 8B Base38.0% ±0.8
117Granite 4.0 Micro37.3% ±0.3
118Gemma 3 4B Pt37.0% ±0.3
119EuroLLM 9B36.6% ±0.6
121Apertus V1.1 4B35.8% ±0.8
122EuroLLM 22B35.0% ±0.5
123Gemma 3 4B34.8% ±0.3
125Qwen3.5 2B34.6% ±0.7
127Olmo 3 32B34.1% ±0.4
128Olmo 3 7B33.4% ±0.4
130LFM2.5 8B A1B Base32.9% ±0.8
132Granite 4.0 Micro32.0% ±0.3
133Llama 3.2 3B Instruct31.6% ±0.3
134SmolLM3 3B Base31.4% ±0.5
136YugoGPT31.1% ±0.7
140Occiglot 7B Eu529.4% ±0.4
142Aya 23 8B28.4% ±0.7
143Qwen3.5 0.8B28.2% ±0.5
146Olmo 3 7B Instruct27.6% ±0.4
149Luciole 23B Base26.5% ±0.6
151DFM Mimir25.4% ±0.6
152Gemma 4 E4B25.2% ±1.0
156Luciole 8B Base22.5% ±1.2
157Qwen3.5 0.8B Base21.7% ±0.7
158Gemma 3 1B IT21.4% ±0.4
159Gemma 3 1B Pt20.9% ±0.5
160YuLan Mini Instruct20.3% ±0.4
161Apertus V1.1 1.5B20.3% ±0.8
162Gemma 4 E2B20.2% ±0.9
164Llama 3.2 1B19.9% ±0.5
168Salamandra 7B15.4% ±0.5
170Apertus V1.1 0.5B14.0% ±0.9
172Pleias 1.2B Preview12.1% ±0.6
173Luciole 1B Base11.9% ±0.7
176LFM2.5-8B-A1B10.8% ±0.7
178EuroLLM 1.7B10.5% ±0.3
179Pleias 3B Preview10.4% ±0.5
180Gemma 3 270M10.3% ±0.8
181LFM2 1.2B10.2% ±0.6
183LFM2 700M8.5% ±0.4
184Gemma 3 270M IT8.1% ±0.5
188LFM2 350M6.3% ±0.3
189SmolLM2 360M6.2% ±0.5
193SmolLM2 135M4.4% ±0.5
195Monad1.5% ±0.2
196Muse 3B0.9% ±0.9
197Ling Mini 2.00.7% ±0.7