Vals AI
BioMysteryBench
Anthropic's benchmark of whether AI agents can solve real-world bioinformatics mysteries from raw data
ModelsModels with a result.11
Top-3 spreadPoints from first to third.7.0 pts
UpdatedDate the source changed the results.21 Sept 2026
Result and cost
Results
11 results#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.