PostTrain Bench

A software-engineering benchmark for post-training infrastructure and implementation tasks, evaluated through the official Harbor implementation.

  • Not in index
  • Coding
ModelsModels with a result.3
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.N/A

Results

3 results
#ModelResultThe score from the source.
GLM-5.339.8%
Kimi K336.6%