Claude Opus 4.7 (Adaptive)

Claude Opus 4.7 (Adaptive) is a reasoning model from Anthropic in the Claude Opus 4.7 family. 24 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.66.2 ±4.8
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.
Input / 1MUS dollars per 1M input tokens.$5
Output / 1MUS dollars per 1M output tokens.$25
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
66.1
CodingCode writing and repair.
68.3
ReasoningLogic problems and puzzles.
63.1
MultimodalTasks with images and text.
63.0
KnowledgeFacts and expert knowledge.
69.6
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
48.6
MathMath problems.
N/A

Results

24 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
Graduate-Level Google-Proof Q&AKnowledge94.2%65.2David Rein et al.
GPQA DiamondKnowledge94.2%65.2David Rein et al.
Artificial Analysis GPQA DiamondKnowledge91.4%62.2Artificial Analysis
CharXiv ReasoningMultimodal91.0%69.7CharXiv authors
τ²-Bench Tool-Agent-User EvaluationAgentic88.6%62.4Victor Barres et al.
Software Engineering Benchmark VerifiedCoding87.6%67.8Carlos E. Jimenez et al.
CharXiv Reasoning without toolsMultimodal82.1%CharXiv authors
BrowseCompAgentic79.3%66.4OpenAI
Artificial Analysis MMMU-ProMultimodal78.8%63.4Artificial Analysis
Artificial Analysis Long Context ReasoningReasoning78.7%62.7Artificial Analysis
OSWorld-VerifiedAgentic78.0%66.4Tianbao Xie et al.
MCP AtlasAgentic77.3%66.4OpenAI
Artificial Analysis Coding IndexCoding73.6%70.8Artificial Analysis
CyberGymAgentic73.1%66.2Zhun Wang et al.
SWE-bench ProCoding64.3%66.2Xiang Deng et al.
OpenAI MRCR v2 8-needle 128K-256KReasoning59.2%OpenAI
Artificial Analysis IFBenchInstruction58.6%48.6Artificial Analysis
Humanity's Last ExamKnowledge54.7%75.1Center for AI Safety et al.
Artificial Analysis Omniscience AccuracyKnowledge48.9%74.4Artificial Analysis
Humanity's Last Exam without toolsKnowledge46.9%68.5OpenAI
Artificial Analysis ITBench-AAAgentic46.7%Artificial Analysis
JobBenchAgentic45.9%65.7Yuetai Li et al.
OfficeQA ProMultimodal43.6%55.9OfficeQA Pro authors
Artificial Analysis Humanity's Last ExamKnowledge42.3%70.4Artificial Analysis
GDPval-AA normalizedAgentic41.9%67.7Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge40.7%73.3Artificial Analysis
Artificial Analysis Agentic IndexAgentic39.5%69.4Artificial Analysis
OSWorld 2.0Agentic18.2%64.6Mengqi Yuan et al.
Critical Physics TasksReasoning12.0%63.5Artificial Analysis

24 benchmarks count, from 26 of 29 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeDeepSeek V4 Flash 073164.6 · $0.28

More from Anthropic

Claude Opus 5.582.8Claude Fable 5.181.4Claude Opus 577.9Claude Fable 577.5Claude Opus 4.871.1