NOVA-63

A broad multilingual benchmark row from Qwen's launch comparisons intended to measure cross-lingual capability beyond a single language family.

  • Not in index
  • Multilingual
ModelsModels with a result.7
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.N/A

Result and price

020406080100$1$3$10$30RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

7 results
#ModelResultThe score from the source.
Qwen3.5 397B59.1%
Qwen3.7 Max59.0%
Qwen3.7 Plus58.8%
Qwen3.6 Plus57.9%
Kimi K2.556.0%
GLM-555.1%