EuroEval
EuroEval Italian
How well a model handles Italian, as the unweighted mean of EuroEval's primary metric over its 8 Italian datasets (sentipolc16, multinerd_it, scala_it, squad_it, ilpost_sum, include_it, multiloko_it, winogrande_it). The primary metric differs by task — MCC for classification, knowledge and common-sense reasoning, micro-F1 without MISC for entity recognition, F1 for reading comprehension, chrF++ for summarization, METEOR for simplification — and is always the first of the two the board prints. Only models measured on every dataset are included, so each mean covers the same ground. Most rows are the validation split; some older open-weight rows are the test split, which EuroEval ranks in the same table.
ModelsModels with a result.160
Top-3 spreadPoints from first to third.2.7 pts
YearYear of release.N/A
Result and price
Results
168 results#ModelResultThe score from the source.