Inference
Explore
Leaderboard
Benchmarks
Methodology
24 Sept 2026
Benchmarks
/
Humanity's Last Exam
Center for AI Safety et al.
Humanity's Last Exam
In index
Knowledge
Source ↗
Models
?
Models with a result.
59
Top result
?
Best result on this test.
65.0
%
Claude Fable 5.1
Top-3 spread
?
Points from first to third.
0.5
pts
Year
?
Year of release.
2025
Result and price
0
20
40
60
80
100
$0.03
$0.1
$0.3
$1
$3
$10
$30
$100
$300
RESULT
PARETO
OUTPUT PRICE PER 1M TOKENS · LOG SCALE
Results
59 results
#
Model
Result
?
The score from the source.
01
Claude Fable 5.1
65.0
%
02
Claude Opus 5
64.7
%
03
Claude Mythos 5
64.5
%
04
Muse Spark 1.1
62.1
%
05
GPT-5.4 Pro
58.7
%
06
Claude Opus 4.8
57.9
%
07
Claude Sonnet 5
57.4
%
08
GPT-5.5 Pro
57.2
%
09
AP
Apodex 1.1
56.1
%
10
MA
Kimi K3
56.0
%
11
TE
Hy4 preview
55.4
%
12
Claude Opus 4.7 (Adaptive)
54.7
%
12
GLM-5.2
54.7
%
14
Claude Opus 4.6
53.0
%
15
DS
dots3-note Preview
52.6
%
16
GLM-5.1
52.3
%
17
GPT-5.5
52.2
%
18
GPT-5.4
52.1
%
19
Muse Spark
50.4
%
19
GLM-5
50.4
%
21
Claude Sonnet 4.6
49.0
%
22
MiMo-V2.5-Pro
48.0
%
23
TM
Inkling-Small
47.8
%
24
IN
Agents-A1
47.6
%
25
ST
Step 5 Preview
46.5
%
26
TM
Inkling
46.0
%
27
OA
Ornith-1.5-397B
44.6
%
28
Qwen3.8 Max
43.6
%
29
DeepSeek V4 Pro 0813
42.7
%
30
GPT-5.4 mini
41.5
%
31
Qwen3.7 Max
41.4
%
32
Gemini 3.5 Flash
40.2
%
33
GPT-5.4 nano
37.7
%
34
DeepSeek V4.1 Flash
36.8
%
35
Qwen3.8-Omni-Flash
36.5
%
36
Qwen3.8-Flash-Next
35.9
%
37
Grok 4.3
35.0
%
38
DeepSeek V4 Flash 0731
34.8
%
39
MA
Kimi K2.6
34.7
%
39
Qwen3.7 Plus
34.7
%
41
Qwen3.8-27B
30.8
%
41
Claude Opus 4.5
30.8
%
43
MA
Kimi K2.5
30.1
%
44
Qwen3.6 Plus
28.8
%
45
Qwen3.5 397B
28.7
%
46
ST
A.X K2
27.8
%
47
Nemotron 3 Ultra
26.7
%
48
Gemma 4 31B
26.5
%
49
OA
Ornith-1.5-35B-A3B
25.6
%
50
TE
Hy3 Preview
25.5
%
51
GLM-4.7
24.8
%
52
Qwen3.6-27B
24.0
%
53
IN
Ling 3.0 Flash
22.7
%
54
Qwen3.6-35B-A3B
21.4
%
55
OA
Ornith-1.5-9B
20.2
%
56
Gemini 2.5 Pro
18.8
%
57
LA
K-EXAONE 2.0
18.3
%
58
Gemma 4 26B A4B
17.2
%
59
OP
MiniCPM5-2B
8.9
%
Humanity's Last Exam results — Inference360