Inference
Explore
Leaderboard
Benchmarks
Methodology
24 Sept 2026
Benchmarks
/
Terminal-Bench 2.1 · BenchLM
Terminal-Bench 2.1 · BenchLM
In index
Agentic
Models
?
Models with a result.
62
Top result
?
Best result on this test.
92.8
%
SWE-2
Top-3 spread
?
Points from first to third.
2.2
pts
Year
?
Year of release.
N/A
Result and price
0
20
40
60
80
100
$0.03
$0.1
$0.3
$1
$3
$10
$30
$100
RESULT
PARETO
OUTPUT PRICE PER 1M TOKENS · LOG SCALE
Results
62 results
#
Model
Result
?
The score from the source.
01
CO
SWE-2
92.8
%
02
GPT-5.6 Sol
91.9
%
03
DeepSeek V4.1 Flash
90.6
%
04
MiMo-V2.6-Pro
89.9
%
05
Gemini 3.8 Flash
89.4
%
06
Muse Spark 1.3
88.8
%
07
MA
Kimi K3
88.3
%
08
GLM-5.3
88.2
%
09
Claude Mythos 5
88.0
%
10
DeepSeek V4 Pro 0813
87.9
%
11
MiMo-V2.6-Flash
87.6
%
12
GPT-5.6 Terra
87.4
%
13
Qwen3.8 Max
86.6
%
14
OA
Ornith-1.5-397B
86.1
%
15
Gemini 3.7 Flash
85.8
%
16
TE
Hy4 preview
85.4
%
17
ST
Step 5 Preview
85.0
%
18
GPT-5.6 Luna
84.7
%
19
Claude Fable 5
84.3
%
19
GLM-5.3-Flash
84.3
%
21
Grok 4.5
83.3
%
22
Muse Spark 1.2
82.9
%
23
DeepSeek V4 Flash 0731
82.7
%
24
SA
Sakana Fugu-Ultra
82.1
%
25
CO
SWE-1.7
81.5
%
26
GLM-5.2
81.0
%
27
Claude Sonnet 5
80.4
%
28
SA
Sakana Fugu
80.2
%
29
Muse Spark 1.1
80.0
%
30
SA
Atria Dawn Preview
78.3
%
31
DA
Ornith-1.0-397B
77.5
%
32
Gemini 3.5 Flash
76.2
%
33
DS
dots3-note Preview
75.1
%
34
Claude Opus 4.8
74.6
%
35
Qwen3.8-27B
73.0
%
36
Seed 2.1 Pro
71.0
%
37
AP
Apodex 1.1
70.8
%
38
PO
Laguna S 2.1
70.2
%
39
MC
Quasar 438B
69.3
%
40
OA
Ornith-1.5-35B-A3B
67.8
%
41
Seed 2.1 Turbo
67.6
%
42
MiniMax M3
66.0
%
43
PA
Pokee-Isaac 28B
65.1
%
44
TM
Inkling-Small
64.7
%
45
DA
Ornith-1.0-35B
64.2
%
46
TM
Inkling
63.8
%
47
MAI-Code-1.1-Flash
62.9
%
48
ST
Step 3.7 Flash
59.5
%
49
UP
Solar Pro 4
57.0
%
49
IN
Ling 3.0 Flash
57.0
%
51
Nemotron 3 Ultra
56.4
%
52
Gemini 3.5 Flash-Lite
54.0
%
53
PM
Ternary Bonsai 2 27B
52.8
%
54
Muse Glimmer 30B
51.7
%
55
OA
Ornith-1.5-9B
46.2
%
56
LA
K-EXAONE 2.0
43.8
%
57
DA
Ornith-1.0-9B
43.1
%
58
ST
A.X K2
36.0
%
59
IB
Granite 4.2 30B
29.2
%
60
Nemotron 3.5 Lightning 30B A3B NVFP4
23.5
%
61
IB
Granite 4.2 8B
20.6
%
62
OP
MiniCPM5-2B
8.6
%
Terminal-Bench 2.1 · BenchLM results — Inference360