EuroEval French

How well a model handles French, as the unweighted mean of EuroEval's primary metric over its 8 French datasets (allocine, eltec, scala_fr, fquad, orange_sum, include_fr, multiloko_fr, hellaswag_fr). The primary metric differs by task — MCC for classification, knowledge and common-sense reasoning, micro-F1 without MISC for entity recognition, F1 for reading comprehension, chrF++ for summarization, METEOR for simplification — and is always the first of the two the board prints. Only models measured on every dataset are included, so each mean covers the same ground. Most rows are the validation split; some older open-weight rows are the test split, which EuroEval ranks in the same table.

ModelsModels with a result.163
Top resultBest result on this test.77.1%Gemini 3.7 Flash
Top-3 spreadPoints from first to third.1.5 pts
YearYear of release.N/A

Result and price

020406080100$0.03$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

171 results
#ModelResultThe score from the source.
01Gemini 3.7 Flash77.1% ±0.3
02GPT-6 Astra76.3% ±0.4
03Gemini 3.6 Flash75.6% ±0.3
05GLM-5.3-Flash73.5% ±0.4
06GPT-5.6 Sol72.0% ±0.4
07GPT-5.6 Luna68.9% ±0.4
08GPT-5.6 Terra68.9% ±0.4
09Claude Sonnet 4.668.5% ±0.4
10Gemma 4 31B67.9% ±0.4
16Qwen3.6 27B FP867.0% ±0.5
18Qwen3.8-27B66.8% ±0.4
21Gemma 4 26B A4B65.2% ±0.5
22Qwen3.5 35B A3B FP864.2% ±0.5
25GPT-4.162.3% ±0.3
27GPT-5.262.0% ±0.4
28Gemma 4 12B IT61.9% ±0.5
33GPT-5 mini60.1% ±0.4
40Llama 3.1-70B58.4% ±0.5
43GPT-OSS 120B58.4% ±0.4
45Gemma 3 27B Pt58.0% ±0.5
47Claude Haiku 4.557.6% ±0.4
48Gemma 3 27B57.5% ±0.3
49Ministral 3 8B Base57.3% ±0.6
52GLM-4.7-Flash57.1% ±0.6
55Qwen3.5 9B55.5% ±0.6
56Gemma 3 12B55.4% ±0.3
57Qwen3 4B Thinking55.3% ±0.3
62Apertus V1.5 70B54.7% ±0.5
63Gemma 3 12B Pt54.5% ±0.4
66Exaone 4.0 32B53.9% ±0.6
67GPT-5 nano53.9% ±0.5
68Gemma 4 E4B IT53.7% ±0.5
75Nesso 4B52.1% ±0.4
77Qwen3.5 4B51.5% ±0.5
78Gemma 4 31B50.0% ±0.8
83Gemma 4 E2B IT47.5% ±0.6
84Aya Expanse 8B47.4% ±0.6
87Qwen3.5 9B Base46.7% ±1.0
89Qwen3.5 2B46.1% ±0.6
90Llama 3.1-8B45.9% ±0.5
94Gemma 4 26B A4B44.4% ±1.0
95Gemma 3 4B44.4% ±0.6
96EuroLLM 22B43.9% ±0.4
97EuroLLM 9B43.7% ±0.4
98SmolLM3 3B Base43.2% ±0.5
99Gemma 4 12B42.8% ±0.9
102AMALIA 9B DPO42.5% ±0.5
103LFM2.5 8B A1B Base42.1% ±0.5
104Gemma 3 4B Pt41.9% ±0.5
105AMALIA 9B SFT40.7% ±0.5
106DFM Mimir40.2% ±0.4
109Llama 3.2 3B Instruct39.2% ±0.4
110Apertus V1.1 4B38.7% ±0.7
113Occiglot 7B Eu537.4% ±0.4
115Qwen3.5 4B Base37.0% ±0.8
116Luciole 23B Base37.0% ±0.7
118Domyn Small V1.035.9% ±1.2
122Aya 23 8B33.0% ±0.7
123Luciole 8B Base32.5% ±0.6
124Qwen3.5 0.8B31.7% ±0.5
125YuLan Mini Instruct31.6% ±0.3
127Apertus V1.1 1.5B30.4% ±0.6
128Gemma 4 E4B30.1% ±0.9
129LFM2.5-8B-A1B29.6% ±0.8
130Qwen3.5 2B Base29.3% ±0.8
134Salamandra 7B25.2% ±0.4
135LFM2 1.2B24.5% ±0.3
136TildeOpen 30B24.4% ±0.5
137Gemma 4 E2B24.3% ±0.8
138Gemma 3 1B Pt24.3% ±0.4
139Llama 3.2 1B24.2% ±0.3
140Gemma 3 1B IT24.0% ±0.4
141Qwen3.5 0.8B Base23.8% ±0.8
144LFM2 700M23.2% ±0.3
145Pleias 3B Preview22.9% ±0.5
146Apertus V1.1 0.5B22.0% ±0.6
147Pleias 1.2B Preview21.6% ±0.6
149Luciole 1B Base20.8% ±0.8
150SmolLM2 360M17.5% ±0.6
152Gemma 3 270M16.4% ±0.8
153Llama 3.2 1B Instruct15.7% ±0.6
154Nesso 0.4B Agentic15.6% ±0.9
156LFM2 350M14.1% ±0.6
158Pleias 350M Preview13.7% ±0.4
160MiniCPM5-1B11.7% ±0.9
161Gemma 3 270M IT10.9% ±0.7
162SmolLM2 135M9.8% ±0.5
164Zagreus 0.4B Ita8.2% ±0.7
165Ling Mini 2.07.6% ±0.8
169Monad1.8% ±0.2
170Muse 3B0.6% ±0.6