HealthBench Hard

A harder subset of OpenAI's HealthBench for evaluating open-ended medical and health reasoning with rubric-based grading.

  • In index
  • Knowledge
ModelsModels with a result.11
Top resultBest result on this test.42.8%Muse Spark
Top-3 spreadPoints from first to third.6.2 pts
YearYear of release.N/A

Result and price

020406080100$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

11 results
#ModelResultThe score from the source.
01Muse Spark42.8%
02GPT-5.440.1%
07GPT-6 Luna31.4%
08GPT-6 Sol30.1%
10Grok 4.2020.3%