ProofBench v1.1

Automated theorem proving benchmark

ModelsModels with a result.33
Top resultBest result on this test.100.0%Claude Fable 5.1
Top-3 spreadPoints from first to third.1.0 pts
UpdatedDate the source changed the results.21 Sept 2026

Result and cost

020406080100$0.03$0.1$0.3$1$3$10$30RESULTUS DOLLARS PER TASK · LOG SCALE

Results

33 results
#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.
01Claude Fable 5.1100.0%$2.78 min
01Alephprover100.0%$9.3518 min
03GPT-6 Astra99.0%$1.715 min
03Claude Opus 599.0%$1.7910 min
05Claude Fable 595.0%$5.2414 min
06Aristotle86.0%
07GPT-5.6 Sol83.0%$1.429 min
08Claude Sonnet 577.0%$2.2816 min
09Hy4 preview75.0%$0.33623 min
10GPT-5.6 Terra74.0%$1.0113 min
11GPT-5.6 Luna60.0%$0.0979 min
12Gemini 3.7 Flash58.0%$0.5597 min
12Qwen3.8 Max58.0%$1.4431 min
14DeepSeek V4 Flash 042356.0%$0.03715 min
15DeepSeek V4.1 Flash54.0%$0.1269 min
16Grok 4.651.0%$0.75914 min
17GLM-5.349.0%$2.0842 min
18Gemini 3.8 Flash48.0%$0.6036 min
19Muse Spark 1.243.0%$0.4248 min
20DeepSeek V4 Pro 042333.0%$0.06212 min
21Gemini 3.5 Flash31.0%$0.5417 min
21Grok 4.531.0%$0.57411 min
23Grok 4.726.0%$1.4813 min
23Gemini 3.1 Pro Preview26.0%$0.78415 min
25MiMo-V2.5-Pro22.0%$0.0920 min
26GLM-5.3-Flash21.0%$0.22471 min
27MiniMax M318.0%$0.41912 min
28Qwen3.8-27B16.0%$2.324 min
28MiMo-V2.516.0%$0.05621 min
30Mistral Medium 3.59.0%$1.138 min
31Inkling-Small6.0%$0.0828 min
32Mercury 2.53.0%$0.0443 min
33Inkling0.0%$0.1997 min