HealthBench length-adjusted score

HealthBench score after applying a verbosity penalty to model responses.

  • Not in index
  • Knowledge
ModelsModels with a result.5
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.N/A

Result and price

020406080100$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

5 results
#ModelResultThe score from the source.
GPT-6 Astra58.3%
GPT-6 Luna54.5%
GPT-6 Sol53.2%