FrontierCode 1.1 Main

Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics.

  • In index
  • Coding
ModelsModels with a result.14
Top resultBest result on this test.54.4%Claude Opus 5.5
Top-3 spreadPoints from first to third.1.0 pts
YearYear of release.N/A

Result and price

020406080100$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

14 results
#ModelResultThe score from the source.
05SWE-250.0%
08GPT-5.543.0%
10SWE-1.742.3%