JobBench

An occupational agent benchmark for professional workflows that workers say they most want delegated to AI.

ModelsModels with a result.30
Top resultBest result on this test.64.9%Muse Spark 1.3
Top-3 spreadPoints from first to third.3.2 pts
YearYear of release.2026

Result and price

020406080100$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

30 results
#ModelResultThe score from the source.
09Kimi K352.9%
12GPT-5.542.7%
13GPT-5.438.9%
16GPT-5.234.3%