Gert Labs Composite Game Benchmark

A game-environment benchmark that evaluates AI models in novel games covering strategic planning, resource management, spatial reasoning, cooperation, and theory of mind.

ModelsModels with a result.52
Top resultBest result on this test.73.0%Claude Opus 4.8
Top-3 spreadPoints from first to third.7.4 pts
YearYear of release.2026

Result and price

020406080100$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

52 results
#ModelResultThe score from the source.
02GPT-5.572.9%
04GPT-5.464.9%
12GLM-5.160.1%
20GLM-551.0%
26MiMo-V2.546.9%
28GPT-5.246.5%
30Grok 4.343.9%
31Qwen3 Max43.7%
33Grok 442.3%
35GPT-5.141.2%
37GLM-4.740.0%
42Grok 4.2038.4%
52GPT-4.125.6%