How we make the score
We compare each result with the results of other models on the same benchmark.
We add the results in eight categories. Each category has a weight.
- Agentic
- 20%
- Coding
- 20%
- Reasoning
- 15%
- Multimodal
- 10%
- Knowledge
- 15%
- Multilingual
- 5%
- Instruction
- 5%
- Math
- 10%
A model gets a score when it has results from 8 benchmarks in 4 categories.
The ± range shows how much a score can change. Lists put models in order by the low end of the range.