Massive Multi-discipline Multimodal Understanding Pro

A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.

ModelsModels with a result.41
Top resultBest result on this test.83.9%Gemini 3.1 Pro
Top-3 spreadPoints from first to third.0.9 pts
YearYear of release.2024

Result and price

020406080100$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

41 results
#ModelResultThe score from the source.
05Kimi K381.6%
07GPT-5.581.2%
07GPT-5.481.2%
11Muse Spark80.4%
13GPT-5.279.5%
22MiniMax M378.1%
22Grok 4.378.1%
24MiMo-V2.577.9%
31Grok 4.2075.2%
35Inkling73.5%