ARC-AGI-1 (semi-private)

Abstract reasoning puzzles held out from training data.

ModelsModels with a result.109
Top resultBest result on this test.98.5%Claude Fable 5
Top-3 spreadPoints from first to third.1.0 pts
YearYear of release.N/A

Result and cost

020406080100$0.001$0.003$0.01$0.03$0.1$0.3$1$3$10$30$100RESULTUS DOLLARS PER TASK · LOG SCALE

Results

109 results
#ModelResultThe score from the source.CostUS dollars to run one task.
12GPT-5.2 (Refine.)94.5%$11.4
35Grok 4 (Refine.)79.6%$8.42
36Inkling79.5%$0.3
37GLM-5.277.0%$0.192
39Gemini 3 Pro75.0%$0.493
42GPT-5 Pro70.2%$4.78
43Grok 4 (Thinking)66.7%$1.01
44GPT-5high effort65.7%$0.509
45Kimi K2.565.3%$0.15
46MiniMax M2.563.7%$0.07
49o3high effort60.8%$0.5
50o3-prohigh effort59.3%$4.16
54DeepSeek V3.257.0%$0.08
62GLM-544.7%$0.17
70o3-minihigh effort34.5%$0.399
83Claude Sonnet 4.525.5%$0.081
85Claude Sonnet 423.8%$0.081
86Claude Opus 422.5%$0.404
87Deepseek R1 (05/28)21.2%$0.046
94DeepSeek-R115.8%$0.06
95Claude Haiku 4.514.3%$0.026
96O1 Mini14.0%$0.135
98GPT-4.510.3%$0.29
100Magistral Medium5.9%$0.102
102Grok 3 [Beta]5.5%$0.093
102GPT-4.15.5%$0.039
104Magistral Small5.0%$0.04
105GPT-4o4.5%$0.05
106Llama 4 Maverick4.4%<$0.01
107GPT-4.1 mini3.5%<$0.01
108Llama 4 Scout0.5%<$0.01
109GPT-4.1 nano0.0%<$0.01