Inference
Explore
Leaderboard
Benchmarks
Methodology
24 Sept 2026
Benchmarks
/
cursorBench32
cursorBench32
In index
Coding
Models
?
Models with a result.
18
Top result
?
Best result on this test.
73.4
%
Claude Fable 5.1
Top-3 spread
?
Points from first to third.
2.9
pts
Year
?
Year of release.
N/A
Result and price
0
20
40
60
80
100
$1
$3
$10
$30
$100
RESULT
PARETO
OUTPUT PRICE PER 1M TOKENS · LOG SCALE
Results
18 results
#
Model
Result
?
The score from the source.
01
Claude Fable 5.1
73.4
%
02
Grok 4.6
70.8
%
03
Claude Fable 5
70.5
%
04
Claude Opus 5
70.0
%
05
Gemini 3.8 Flash
69.2
%
06
GPT-5.6 Sol
67.2
%
07
Grok 4.5
66.7
%
08
GPT-5.6 Terra
64.9
%
09
Claude Opus 4.8
62.3
%
10
Claude Sonnet 5
61.5
%
11
GPT-5.6 Luna
61.1
%
12
MA
Kimi K3
60.8
%
13
GPT-5.5
58.4
%
14
CU
Composer 2.5
56.1
%
15
GLM-5.2
55.0
%
16
Gemini 3.6 Flash
53.5
%
17
MA
Kimi K2.7 Code
49.7
%
18
Gemini 3.5 Flash
48.8
%
cursorBench32 results — Inference360