τ²-bench Airline

Share of tasks an agent completes on one attempt, against a simulated user who holds information it must ask for.

ModelsModels with a result.22
Top resultBest result on this test.84.0%Claude Opus 4.5
Top-3 spreadPoints from first to third.1.5 pts
UpdatedDate the source changed the results.2 Mar 2026

Result and price

020406080100$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

22 results
#ModelResultThe score from the source.CostUS dollars to run one task.
12GPT-5Sierra62.5%$0.134
17GPT-4.1Sierra56.0%
21o3OpenAI52.0%