PostTrainBench v1.1

Post-training four base language models across seven weighted benchmarks, with ten hours and one H100 per run.

ModelsModels with a result.14
Top resultBest result on this test.49.3%Claude Opus 5.5
Top-3 spreadPoints from first to third.5.0 pts
YearYear of release.2026

Result and price

020406080100$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

14 results
#ModelResultThe score from the source.
08Kimi K332.0%
09GLM-5.231.7%
11GPT-5.527.2%
12Grok 4.523.4%
14GPT-5.419.0%