Vals AI
MedQA
Evaluating language model bias in medical questions.
ModelsModels with a result.91
Top-3 spreadPoints from first to third.0.1 pts
UpdatedDate the source changed the results.16 Apr 2026
Result and cost
Results
92 results#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.