FrontierCode 1.1 Main
Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics.
ModelsModels with a result.14
Top-3 spreadPoints from first to third.1.0 pts
YearYear of release.N/A
Result and price
Results
14 results#ModelResultThe score from the source.