CharXiv Reasoning

A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.

  • In index
  • Multimodal
ModelsModels with a result.39
Top resultBest result on this test.93.5%Claude Mythos 5
Top-3 spreadPoints from first to third.2.1 pts
YearYear of release.N/A

Result and price

020406080100$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

39 results
#ModelResultThe score from the source.
04Kimi K391.3%
14Muse Spark86.4%
19GPT-5.482.8%
21GPT-5.282.1%
22Inkling82.0%
26MiMo-V2.581.0%
38Grok 4.2060.9%