EuroEval Italian

How well a model handles Italian, as the unweighted mean of EuroEval's primary metric over its 8 Italian datasets (sentipolc16, multinerd_it, scala_it, squad_it, ilpost_sum, include_it, multiloko_it, winogrande_it). The primary metric differs by task — MCC for classification, knowledge and common-sense reasoning, micro-F1 without MISC for entity recognition, F1 for reading comprehension, chrF++ for summarization, METEOR for simplification — and is always the first of the two the board prints. Only models measured on every dataset are included, so each mean covers the same ground. Most rows are the validation split; some older open-weight rows are the test split, which EuroEval ranks in the same table.

ModelsModels with a result.160
Top resultBest result on this test.77.8%GPT-6 Astra
Top-3 spreadPoints from first to third.2.7 pts
YearYear of release.N/A

Result and price

020406080100$0.03$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

168 results
#ModelResultThe score from the source.
01GPT-6 Astra77.8% ±0.3
02Gemini 3.7 Flash76.9% ±0.3
03Gemini 3.6 Flash75.1% ±0.3
04GLM-5.3-Flash73.3% ±0.4
06GPT-5.6 Sol70.1% ±0.4
07GPT-5.6 Terra67.8% ±0.4
08Qwen3.6 27B FP867.7% ±0.5
09Gemma 4 31B66.8% ±0.4
12Qwen3.8-27B66.2% ±0.5
15Qwen3.5 35B A3B FP865.7% ±0.5
16Claude Sonnet 4.665.3% ±0.5
22Gemma 4 26B A4B61.1% ±0.5
23GPT-5 mini60.1% ±0.3
24GPT-5.260.1% ±0.4
25GPT-5.6 Luna60.0% ±0.5
26Gemma 4 12B IT59.5% ±0.6
27GPT-OSS 120B58.4% ±0.6
32GLM-4.7-Flash57.1% ±0.5
33GPT-4.157.1% ±0.4
35Qwen3.5 9B56.9% ±0.6
41Qwen3 4B Thinking55.4% ±0.4
44GPT-5 nano53.8% ±0.5
45Gemma 3 27B Pt53.8% ±0.6
47Qwen3.5 4B53.3% ±0.7
48Claude Haiku 4.553.1% ±0.5
52Qwen3.5 9B Base51.3% ±0.7
54Gemma 4 E4B IT50.9% ±0.6
55Ministral 3 8B Base50.5% ±0.8
57Gemma 4 31B50.3% ±0.9
59Exaone 4.0 32B49.7% ±0.7
60Gemma 3 12B49.2% ±0.3
61Nesso 4B49.1% ±0.3
62Apertus V1.5 70B48.7% ±0.7
64Domyn Small V1.048.1% ±0.6
72Gemma 3 12B Pt46.8% ±0.5
73Qwen3.5 4B Base46.4% ±0.7
78EuroLLM 22B44.0% ±0.6
80Aya Expanse 8B43.6% ±0.7
84Gemma 4 E2B IT42.3% ±0.8
86Olmo 3 32B41.5% ±0.5
87Gemma 4 12B41.5% ±0.9
89Gemma 4 26B A4B41.0% ±0.9
94Llama 3.1-8B38.5% ±0.6
95Qwen3.5 2B38.1% ±0.7
96AMALIA 9B DPO37.7% ±0.7
98AMALIA 9B SFT36.9% ±0.6
99Gemma 3 4B36.7% ±0.4
100SmolLM3 3B Base36.4% ±0.5
101EuroLLM 9B36.2% ±0.4
103Gemma 3 4B Pt35.4% ±0.4
104Olmo 3 7B35.2% ±0.4
107DFM Mimir33.6% ±0.6
108LFM2.5 8B A1B Base33.3% ±0.6
109Apertus V1.1 4B33.3% ±0.8
111Llama 3.2 3B Instruct32.5% ±0.5
112Occiglot 7B Eu531.8% ±0.4
113Qwen3.5 0.8B31.4% ±0.6
115Olmo 3 7B Instruct31.1% ±0.4
117Luciole 23B Base29.0% ±0.9
123Gemma 4 E4B26.4% ±0.9
126Qwen3.5 0.8B Base24.4% ±0.8
127YuLan Mini Instruct24.4% ±0.4
128Luciole 8B Base23.6% ±0.9
130LFM2.5-8B-A1B20.8% ±1.0
131Gemma 3 1B IT20.7% ±0.4
134Aya 23 8B20.1% ±0.8
136LFM2 1.2B19.5% ±0.3
137Gemma 3 1B Pt19.3% ±0.4
138Salamandra 7B18.6% ±0.5
139Apertus V1.1 1.5B18.2% ±0.8
141Gemma 4 E2B16.6% ±0.9
142Pleias 1.2B Preview16.4% ±0.5
143LFM2 700M15.4% ±0.3
144Apertus V1.1 0.5B15.3% ±0.7
145LFM2 350M15.2% ±0.4
147Pleias 3B Preview15.0% ±0.6
148Llama 3.2 1B14.7% ±0.5
149Luciole 1B Base14.1% ±0.7
150Nesso 0.4B Agentic13.1% ±0.7
151Ling Mini 2.012.9% ±0.6
153Pleias 350M Preview12.4% ±0.5
154SmolLM2 360M12.2% ±0.5
156Llama 3.2 1B Instruct11.7% ±0.4
157Gemma 3 270M9.9% ±0.6
160SmolLM2 135M8.5% ±0.3
161Gemma 3 270M IT7.5% ±0.5
163Zagreus 0.4B Ita6.2% ±0.8
166Muse 3B2.3% ±1.4
167Monad2.0% ±0.3