ARC-AGI-1 (public eval)

Public set of abstract reasoning puzzles.

ModelsModels with a result.98
Top resultBest result on this test.99.0%Claude Fable 5.1
Top-3 spreadPoints from first to third.0.0 pts
YearYear of release.N/A

Result and cost

020406080100$0.001$0.003$0.01$0.03$0.1$0.3$1$3$10RESULTUS DOLLARS PER TASK ยท LOG SCALE

Results

98 results
#ModelResultThe score from the source.CostUS dollars to run one task.
27Inkling90.6%$0.216
33GLM-5.280.4%$0.175
35GPT-5 Pro77.0%$4.03
38Grok 4 (Thinking)73.7%$0.818
39Kimi K2.573.1%$0.14
44GPT-5high effort65.9%$0.392
45o3high effort64.3%$0.404
47o3-prohigh effort63.3%$3.92
49DeepSeek V3.261.6%$0.07
51MiniMax M2.559.1%$0.09
52GLM-558.6%$0.14
63o3-minihigh effort46.6%$0.305
74Claude Opus 435.5%$0.349
75Claude Sonnet 4.535.4%$0.069
77Codex Mini (Latest)33.4%$0.127
78Claude Sonnet 433.0%$0.07
82Deepseek R1 (05/28)27.0%$0.047
83Claude Haiku 4.526.6%$0.022
88O1 Mini13.2%$0.124
90GPT-4.111.8%$0.032
91Magistral Medium8.9%$0.107
92Magistral Small8.6%$0.029
93Grok 3 [Beta]8.4%$0.074
95GPT-4.1 mini7.2%<$0.01
96Llama 4 Maverick7.1%<$0.01
97Llama 4 Scout2.4%<$0.01
98GPT-4.1 nano1.8%<$0.01