Public Benefits Bench v1.1

Can AI help people navigate SNAP benefits?

ModelsModels with a result.35
Top resultBest result on this test.76.9%Claude Opus 5
Top-3 spreadPoints from first to third.6.5 pts
UpdatedDate the source changed the results.21 Sept 2026

Result and cost

020406080100$0.01$0.03$0.1$0.3$1$3$10$30RESULTUS DOLLARS PER TASK · LOG SCALE

Results

35 results
#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.
01Claude Opus 576.9%$3.9539 min
02Claude Fable 5.174.9%$7.4861 min
03Claude Fable 570.4%$4.5923 min
04Hy4 preview68.6%$0.25364 min
05GLM-5.368.5%$0.94742 min
06Muse Spark 1.268.5%$0.58811 min
07Claude Opus 4.868.1%$1.8915 min
08Qwen3.8 Max67.1%$0.99162 min
09Grok 4.666.8%$0.91629 min
10GPT-5.6 Sol66.5%$8.8563 min
11Claude Sonnet 566.0%$1.9424 min
12Grok 4.765.6%$2.2619 min
13Gemini 3.8 Flash65.3%$1.019 min
14DeepSeek V4.1 Flash64.3%$0.07313 min
15MiniMax M364.1%$0.25414 min
16DeepSeek V4 Pro 042362.9%$0.4618 min
17Claude Sonnet 4.662.4%$1.1117 min
18GPT-5.6 Terra62.4%$1.222 min
19GLM-5.161.8%$0.34112 min
20Muse Spark 1.161.3%$0.2789 min
21GPT-5.6 Luna61.2%$0.2742 min
22GPT-5.560.9%$3.9742 min
23Gemini 3.5 Flash59.5%$1.118 min
24Inkling-Small59.4%$0.09611 min
25Inkling58.5%$0.36816 min
26Gemini 3.6 Flash56.6%$0.3859 min
27Claude Haiku 4.554.3%$0.2256 min
28Gemini 3.1 Pro Preview53.8%$0.3776 min
29Gemini 3.5 Flash-Lite51.8%$0.0964 min
30Grok 4.351.7%$0.2646 min
31Ling 3.0 Flash50.7%$0.058 min
32Mercury 2.545.3%$0.0685 min
33Grok 4.1 Fast (Reasoning)44.8%$0.0314 min
34Laguna M.143.8%—0.0 s
35Laguna XS.241.4%—0.0 s