ARC-AGI-2 (public eval)

Public set of abstract reasoning puzzles.

ModelsModels with a result.104
Top resultBest result on this test.98.3%Claude Fable 5.1
Top-3 spreadPoints from first to third.1.3 pts
YearYear of release.N/A

Result and cost

020406080100$0.003$0.01$0.03$0.1$0.3$1$3$10$30RESULTUS DOLLARS PER TASK ยท LOG SCALE

Results

104 results
#ModelResultThe score from the source.CostUS dollars to run one task.
29Inkling37.8%$0.713
33Grok 4 (Thinking)21.1%$2.08
34GLM-5.220.8%$0.229
38GPT-5 Pro13.3%$8.01
39Kimi K2.512.1%$0.3
41GPT-5high effort9.6%$0.774
48GLM-55.4%$0.29
48MiniMax M2.55.4%$0.19
59DeepSeek V3.23.9%$0.13
61Claude Sonnet 4.53.8%$0.131
63o3high effort2.9%$0.9
69Claude Sonnet 42.1%$0.131
73Claude Opus 41.3%$0.663
81DeepSeek-R10.3%$0.08
84Claude Haiku 4.50.0%$0.043
84Grok 3 [Beta]0.0%$0.14
84Magistral Medium0.0%$0.106
84GPT-4.10.0%$0.069
84Magistral Small0.0%$0.05
84GPT-4.1 mini0.0%$0.014
84Llama 4 Maverick0.0%$0.013
84GPT-4o0.0%$0.08
84Llama 4 Scout0.0%<$0.01
84GPT-4.1 nano0.0%<$0.01
84GPT-4o mini0.0%$0.01
84Gemini 1.5 Pro0.0%$0.04
84GPT-4.50.0%$2.07
84O1 Mini0.0%$0.184