GPT-6 Astra

GPT-6 Astra is a reasoning model from OpenAI in the GPT-6 family. 48 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.83.5 ±2.8
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.28/s
Input / 1MUS dollars per 1M input tokens.$10 batch $5
Output / 1MUS dollars per 1M output tokens.$50 batch $25 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1444 (#59)

The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.

2,693 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
82.0
CodingCode writing and repair.
81.9
ReasoningLogic problems and puzzles.
85.4
MultimodalTasks with images and text.
74.2
KnowledgeFacts and expert knowledge.
81.0
MultilingualTasks in many languages.
95.0
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
84.0

Results

48 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
OpenAI MRCR v2 8-needle 256K-512KReasoning100.0%Meta Superintelligence Labs
IOICoding100.0%89.221 Sept 2026Vals AI
OTIS Mock AIME 2024-2025Math100.0%67.4max effortEpoch AI
ProofBench v1.1Math99.0%89.521 Sept 2026Vals AI
FrontierMath-Tier-4-v2-PrivateMath97.6%94.8high effortEpoch AI
ARC-AGI-1 (semi-private)Reasoning97.5%73.2max effortARC Prize Foundation
OpenAI MRCR v2 8-needle 512K-1MReasoning96.3%Meta Superintelligence Labs
Artificial Analysis GPQA DiamondKnowledge96.1%67.0Artificial Analysis
Graduate-Level Google-Proof Q&AKnowledge96.0%66.9David Rein et al.
GPQA DiamondKnowledge96.0%66.9David Rein et al.
BenchCAD Vision2Code voxel IoU with toolsMultimodal95.9%Zhang et al. and Anthropic
GPQA diamondKnowledge95.8%66.7max effortEpoch AI
ARC-AGI-2 (semi-private)Reasoning95.0%87.6max effortARC Prize Foundation
FrontierMath-Tiers-1-3-v2-PrivateMath93.7%84.2max effortEpoch AI
ScreenSpot ProMultimodal92.7%75.1Kaixin Li et al.
BrowseCompAgentic91.5%76.6OpenAI
Vibe Code Bench v1.1Coding89.6%79.4OpenHands21 Sept 2026Vals AI
Terminal-Bench 2.1Agentic87.3%75.621 Sept 2026Vals AI
Artificial Analysis MMMU-ProMultimodal86.9%73.3Artificial Analysis
Mystery Game PuzzlesReasoning84.0%95.0max effortEpoch AI
Artificial Analysis Long Context ReasoningReasoning80.7%64.1Artificial Analysis
Furniture AssemblyReasoning80.0%95.0max effortEpoch AI
EuroEval ItalianMultilingual77.8%95.0EuroEval
EuroEval SwedishMultilingual77.6%95.0EuroEval
Artificial Analysis Coding IndexCoding76.9%73.2Artificial Analysis
EuroEval FrenchMultilingual76.3%95.0EuroEval
EBR-benchReasoning76.2%95.0max effortEpoch AI
SimpleQA VerifiedKnowledge75.6%91.6max effortEpoch AI
EuroEval PortugueseMultilingual74.7%95.0EuroEval
EuroEval DutchMultilingual74.7%95.0EuroEval
DeepSWEAgentic74.1%77.6Datacurve AI
OSWorld 2.0Agentic72.6%90.5Mengqi Yuan et al.
Chess PuzzlesReasoning72.0%95.0max effortEpoch AI
EuroEval SpanishMultilingual70.0%89.7EuroEval
Artificial Analysis AutomationBenchAgentic68.5%83.1Artificial Analysis
HealthBench Professional raw scoreKnowledge68.2%Anthropic
EuroEval GermanMultilingual68.0%87.2EuroEval
ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobAgentic68.0%95.0NeoCognition
Code MigrationCoding67.7%88.321 Sept 2026Vals AI
EuroEval PolishMultilingual67.4%86.4EuroEval
FrontierSWE v2Coding65.5%90.6Proximal
HealthBench ProfessionalKnowledge64.7%Rebecca Soskin Hicks et al.
Terminal-Bench-Science 0.1Agentic64.6%Terminal-Bench-Science Team
FrontierCode 1.1 ExtendedCoding64.5%Cognition
ARC-AGI-3 (semi-private)Reasoning62.7%max effortARC Prize Foundation
Artificial Analysis Omniscience AccuracyKnowledge62.6%91.3Artificial Analysis
Agents' Last ExamAgentic59.3%95.0DeepSeek-AI
HealthBench length-adjusted scoreKnowledge58.3%Anthropic
Terminal-Bench 4.0.0Agentic58.2%95.0max effort · Codex21 Sept 2026Terminal-Bench
Humanity's Last Exam with toolsAgentic57.2%70.0DeepSeek-AI
Terminal-Bench 4.0Agentic57.1%95.021 Sept 2026Vals AI
HealthBench raw scoreKnowledge56.9%Anthropic
Artificial Analysis SciCodeCoding56.5%71.3Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge54.7%83.9Artificial Analysis
FrontierCode 1.1 MainCoding53.3%81.0Cognition
Artificial Analysis Intelligence IndexKnowledge52.7%88.3Artificial Analysis
GDPval-AA normalizedAgentic52.1%75.6Artificial Analysis
Artificial Analysis Agentic IndexAgentic51.5%79.1Artificial Analysis
Artificial Analysis AnalystAgentAgentic51.2%77.4Artificial Analysis
MirrorCodeCoding46.7%high effortEpoch AI
ExploitGymAgentic42.4%91.5Zhun Wang et al.
AutomationBenchAgentic41.4%84.8Moonshot AI
Artificial Analysis Tau3-BankingAgentic41.4%68.9Artificial Analysis
GeneBench-ProReasoning37.8%OpenAI
HealthBench HardKnowledge36.6%78.3Meta AI
Medical Long Context Reasoning (MLCR-AA)Reasoning35.0%68.3Wisedocs and Artificial Analysis
Critical Physics TasksReasoning31.7%95.0Artificial Analysis
Artificial Analysis GDP.pdfAgentic31.0%82.8Artificial Analysis
Vibe Code Bench 1-100Coding27.6%82.2OpenHands16 Sept 2026Vals AI
Agent Arena task outcomeAgentic17.787.3max effort15 Sept 2026LMArena
Agent Arena command recoveryAgentic7.375.6max effort15 Sept 2026LMArena
ProgramBenchCoding5.5%21 Sept 2026Vals AI
FrontierMath-ErdosMath2.9%max effortEpoch AI
Agent Arena steerabilityAgentic-0.566.8max effort15 Sept 2026LMArena

48 benchmarks count, from 60 of 74 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEpoch AI, collected directlyCC BY — free to use and redistribute with attributionARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repositoryTerminal-Bench, collected directlyNo licence stated for the leaderboard. The harness repo is Apache-2.0LMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Claude Opus 5.582.8 · $20

More from OpenAI

GPT-5.6 Sol76.7GPT-5.6 Terra73.4GPT-6 Sol76.9GPT-5.572.3GPT-5.5 Pro77.9