CaseLaw v2

Private question-answer benchmark over Canadian court-cases.

ModelsModels with a result.52
Top resultBest result on this test.79.3%Grok 4.3
Top-3 spreadPoints from first to third.9.4 pts
UpdatedDate the source changed the results.4 May 2026

Result and cost

020406080100$0.0003$0.001$0.003$0.01$0.03$0.1$0.3RESULTUS DOLLARS PER TASK · LOG SCALE

Results

52 results
#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.
01Grok 4.379.3%$0.0443 s
02GPT-5.173.4%$0.04324 s
03GPT-4.169.9%$0.0476.7 s
04GPT-5 mini68.5%$0.01143 s
05Claude Opus 4.768.4%$0.16218 s
06GPT-566.5%$0.0672 min
07GPT-5.566.2%$0.15525 s
08GPT-5.266.0%$0.10486 s
09Grok 465.8%$0.15932 s
10Kimi K2 Thinking65.7%$0.0152 min
10Grok 4 Fast (Reasoning)65.7%<$0.0124 s
12Gemini 3.1 Pro Preview64.8%$0.13657 s
13Command A64.5%$0.09842 s
14Claude Sonnet 4.664.0%$0.14536 s
15Gemini 2.5 Pro63.9%$0.25744 s
16GPT-5.463.8%$0.10157 s
17Muse Spark63.1%<$0.0162 s
18Claude Opus 4.5 Thinking62.6%$0.22917 s
21Mistral Large61.4%$0.01941 s
22MiniMax M2.760.9%<$0.0122 s
23Grok 4.1 Fast (Reasoning)60.5%$0.01119 s
24Qwen3.5 Plus Thinking59.7%$0.032 min
24GPT-4o59.7%$0.06741 s
26DeepSeek V4 Pro 042359.4%$0.01348 s
27Kimi K2.5 Thinking58.7%$0.0254 min
28Trinity-Large-Thinking57.9%<$0.0115 s
30Qwen3.5 Flash55.9%<$0.0144 s
31Gemini 3 Flash Preview55.8%$0.0254 min
31MiniMax M2.155.8%<$0.0137 s
33DeepSeek V3.2 (Thinking)55.4%$0.01866 s
35GLM-4.754.9%$0.02160 s
36Grok 4.20 (Reasoning)54.4%$0.06811 s
37DeepSeek V3.153.9%$0.01741 s
38MiniMax M2.553.5%$0.01231 s
39Qwen3.6-27B53.2%$0.02442 s
40Gemini 3 Pro53.1%$0.13832 s
41Gemma 4 31B52.6%$029 s
41GPT-5 nano52.6%<$0.0156 s
43GLM-5 Thinking52.5%$0.0342 min
44GPT-5.4 nano51.9%<$0.017.6 s
45GPT-5.4 mini51.7%$0.04158 s
46GLM-5.151.6%$0.03449 s
47Qwen3.6 Plus51.4%$0.04744 s
48Qwen3 Max51.2%$0.0852 min
49GPT-OSS 120B48.8%<$0.0168 s
50Qwen 3.6 Max (preview)47.9%$0.05243 s
51Mistral Medium 3.544.2%$0.08565 s
52GPT-OSS 20B43.8%<$0.0186 s