τ²-bench Retail

Share of tasks an agent completes on one attempt, against a simulated user who holds information it must ask for.

ModelsModels with a result.23
Top resultBest result on this test.85.3%Gemini 3.0 Pro
Top-3 spreadPoints from first to third.1.1 pts
UpdatedDate the source changed the results.30 Apr 2026

Result and price

020406080100$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

23 results
#ModelResultThe score from the source.CostUS dollars to run one task.
07GPT-5Sierra81.6%$0.106
16GPT-4.1Sierra74.0%
17o3OpenAI73.9%