MMAnswerBench

A multimodal mathematical reasoning benchmark that tests whether models can answer visually grounded math questions correctly.

  • Not in index
  • Math
ModelsModels with a result.10
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.N/A

Result and price

020406080100$0.3$1$3$10$30RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

10 results
#ModelResultThe score from the source.
GLM-5.291.0%
Kimi K2.686.0%
GLM-5.183.8%
Qwen3.6 Plus83.8%
GLM-582.5%
Kimi K2.581.8%
Qwen3.5 397B80.9%
Qwen3.6-27B80.8%