BioMysteryBench

Anthropic's benchmark of whether AI agents can solve real-world bioinformatics mysteries from raw data

ModelsModels with a result.11
Top resultBest result on this test.79.3%GPT-6 Astra
Top-3 spreadPoints from first to third.7.0 pts
UpdatedDate the source changed the results.21 Sept 2026

Result and cost

020406080100$0.03$0.1$0.3$1$3$10RESULTUS DOLLARS PER TASK · LOG SCALE

Results

11 results
#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.
01GPT-6 AstraTerminus 279.3%$1.725 min
01Claude Opus 5Terminus 279.3%$3.2937 min
03Grok 4.6Terminus 272.2%$1.2314 min
04GPT-5.6 SolTerminus 271.1%$1.9318 min
05Hy4 previewTerminus 269.3%$0.36224 min
07Muse Spark 1.2Terminus 264.8%$0.99524 min
10GPT-5.6 LunaTerminus 261.5%$0.08919 min
11Gemini 3.6 FlashTerminus 258.5%$1.2513 min