Benchmarks
345 tests · 143 in index
Agentic
86 testsBenchmarkModelsModels with a result.Top resultBest result on this test.
τ²-Bench Tool-Agent-User EvaluationVictor Barres et al.132GLM-5.299.1%GDPval-AA normalized107Claude Opus 5.567.3%Terminal-Bench 2.0Vals AI65GPT-5.573.2%Terminal-Bench 2.1Vals AI63GPT-6 Astra87.3%Terminal-Bench 2.1 · BenchLM62SWE-292.8%LiveBench Agentic CodingLiveBench55DeepSeek V4.1 Flash Max77.3%Gert Labs Composite Game BenchmarkGert Labs52Claude Opus 4.873.0%Terminal-Bench 1.0Vals AI45GPT-5.263.8%BrowseComp44Atria Dawn Preview92.5%Agent Arena command recoveryLMArena43Claude Opus 513.1Agent Arena steerabilityLMArena43Claude Opus 4.811.1Agent Arena task outcomeLMArena43Claude Fable 5.119.8Terminal-Bench 2.0 · BenchLM41GPT-5.582.0%Claw-EvalBowen Ye et al.39Ornith-1.5-397B81.4%MCP Atlas38Muse Spark 1.188.1%DeepSWEDatacurve AI36Muse Spark 1.375.4%OSWorld-VerifiedTianbao Xie et al.34Qwen3.8 Max86.1%JobBenchYuetai Li et al.30Muse Spark 1.364.9%Terminal-Bench 4.0Vals AI29GPT-6 Astra57.1%CyberGymZhun Wang et al.27MiMo-V2.6-Flash95.1%APEX-Agents-AAArtificial Analysis / Mercor26Gemini 3.5 Flash47.1%τ²-bench BankingSierra Research26Qwen3.8 Max55.2%τ²-bench RetailSierra Research23Gemini 3.0 Pro85.3%Berkeley Function Calling Leaderboard v422BTL-388.5%OSWorld 2.0Mengqi Yuan et al.22GPT-6 Astra72.6%Toolathlon22Muse Spark 1.175.6%τ²-bench AirlineSierra Research22Claude Opus 4.584.0%Humanity's Last Exam with tools20Claude Opus 5.567.7%Toolathlon-Verified20Claude Opus 580.6%AutomationBench19DeepSeek V4.1 Flash54.8%ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobNeoCognition18Claude Fable 5.172.0%DeepSearchQA18Atria Dawn Preview96.0%τ²-bench TelecomSierra Research18Qwen3 Max Thinking98.2%τ³-Bench Tool-Agent-User EvaluationSierra Research18Mercury 2.596.0%Artificial Analysis EnterpriseOps-GymArtificial Analysis17Claude Fable 551.1%Agents' Last Exam15GPT-6 Astra59.3%Artificial Analysis AnalystAgentArtificial Analysis15Gemini 3.7 Flash60.0%Artificial Analysis AutomationBenchArtificial Analysis15Claude Opus 5.569.5%Terminal-Bench 4.0.0Terminal-Bench15GPT-6 Astra58.2%WideResearch15Hy4 preview83.9%Artificial Analysis GDP.pdfArtificial Analysis14GPT-6 Astra31.0%Artificial Analysis Tau3-BankingArtificial Analysis14Grok 4.650.7%ExploitGymZhun Wang et al.14GPT-6 Astra42.4%Terminal-Bench 3.0Ryan Marten et al.13Claude Opus 542.7%Artificial Analysis Harvey LAB-AAArtificial Analysis12Kimi K394.6%VITA-BenchMeituan LongCat Team12Qwen3.7 Max47.9%APEX-Agents10Grok 4.657.5%QwenClawBench10Qwen3.7 Max64.3%
Not in index
ResearchClawBench19—AndroidWorld9—Artificial Analysis ITBench-AAArtificial Analysis8—DeepPlanningDeepPlanning authors7—PinchBenchKilo Code6—Terminal-Bench Hard6—MCP-Tasks5—CoWorkBench4—Data Research and Analysis with Complex Operations4—MLE-Bench Lite3—Terminal-Bench-Science 0.1Terminal-Bench-Science Team3—Toolathlon Verified average assistant turns3—Toolathlon Verified Pass cubed3—Toolathlon Verified Pass@33—WebArena-Verified Browser Agent BenchmarkAmine El Hattami et al.3—BankerToolBench2—Legal Agent Benchmark all-pass rate — Harvey held-out set2—Legal Agent Benchmark mean criterion-pass rate — Harvey held-out set2—MCP-Atlas mean claim coverage2—MM-ClawBench2—OSWorld2—SpreadsheetBench 22—terminalBench32—WebVoyager2—Berkeley Function Calling Leaderboard v3Shishir G. Patil et al.1—CTI-REALM1—CWE-BenchCollinear1—CybenchAndy K. Zhang et al.1—DECK-Bench (Internal)1—GDPval rubrics1—General AI Assistants1—Kimi Claw 24/7 Bench1—Legal Agent Benchmark all-pass rate — Anthropic harness1—Legal Agent Benchmark mean criterion-pass rate — Anthropic harness1—MCPMark-VerifiedMCPMark1—MobileWorld1—Multi-Agent BrowseComp — 10-agent team prerelease configuration1—τ²-Bench Airline DomainVictor Barres et al.1—Coding
62 testsBenchmarkModelsModels with a result.Top resultBest result on this test.
LiveCodeBenchVals AI136Claude Fable 5.190.5%Artificial Analysis SciCodeArtificial Analysis94Claude Opus 5.566.9%Vibe Code Bench v1.1Vals AI92Claude Fable 590.4%SWE-benchVals AI84Claude Opus 597.0%Software Engineering Benchmark VerifiedCarlos E. Jimenez et al.75Claude Opus 596.0%SWE-bench ProXiang Deng et al.71Claude Opus 5.589.9%SWE-bench Verified · SWE-bench/experimentsSWE-bench team62Claude Opus 4.576.8%IOI v1Vals AI60Claude Opus 591.7%Code MigrationVals AI58GPT-6 Astra67.7%LiveBench CodingLiveBench55Claude Opus 5.589.3%SWE-bench Verified · swebench.comSWE-bench team46Claude 4.5 Opus74.4%SkillsBenchVals AI34DeepSeek V4.1 Flash69.8%LiveCodeBench v6LiveCodeBench maintainers31Sakana Fugu-Ultra93.2%SWE-Bench verifiedEpoch AI31Claude Opus 4.783.5%NL2Repo29DeepSeek V4.1 Flash65.4%IOIVals AI28GPT-6 Astra100.0%Scientific Code Benchmark27Sakana Fugu60.1%cursorBench3218Claude Fable 5.173.4%Vibe Code Bench 1-100Vals AI18Claude Opus 528.5%React Native EvalsCallstack16Composer 296.1%FrontierSWE v2Proximal15GPT-6 Astra65.5%SWE-bench Lite · swebench.comSWE-bench team15Claude 4 Sonnet56.7%FrontierCode 1.1 Main14Claude Opus 5.554.4%VulcanBench v3VulcanBench contributors14Grok 4.589.9%SWE-bench Lite · SWE-bench/experimentsSWE-bench team13Claude 4 Sonnet57.5%OpenHarmony Bench v1.0OpenHarmony Bench authors12GLM-5.360.8%ProgramBench: Can Language Models Rebuild Programs From Scratch?John Yang et al.12Claude Opus 593.0%LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeNaman Jain et al.7Qwen3.7 Max91.6%SWE-bench Verified (mini-swe-agent-v2)5Claude Opus 4.675.6%
Not in index
ProgramBenchVals AI42—SWE-bench Multilingual · SWE-bench/experimentsSWE-bench team14—SWE-RebenchNebius13—LiveCodeBench ProLiveCodeBench Pro authors8—MirrorCodeEpoch AI8—cursorBench317—FrontierCode 1.1 Extended7—SWE-bench (full test split)SWE-bench team7—cursorBench406—Bug Hunt BenchPawel Huryn4—MLS-Bench LiteMLS-Bench4—SWE-MarathonAbundant AI and BenchFlow4—FrontierSWEEvan Chu et al.3—PostTrain Bench3—VulcanBench Coding Intelligence Index v1VulcanBench contributors3—Artificial Analysis LiveCodeBenchArtificial Analysis2—DeepSeek DSBench FullStack2—DeepSeek DSBench Hard2—Evaluating Large Language Models Trained on CodeMark Chen et al.2—Kimi Code Bench v22—LiveCodeBench Pass@1 with Chain-of-Thought2—LiveCodeBench v5LiveCodeBench maintainers2—BigCodeBench1—EEBench V1 core corpusatopile1—KernelBench Hard1—Multi-SWE Bench1—PaperBench1—ProgramBench hidden-test pass rate after episode 1Yang et al.1—Spider 2.0-LiteSpider 2.0 authors1—SVG-Bench1—SWE-Atlas Codebase QnA1—VIBE V21—VIBE-Pro1—Reasoning
25 testsBenchmarkModelsModels with a result.Top resultBest result on this test.
Critical Physics TasksArtificial Analysis180GPT-5.6 Sol32.3%Artificial Analysis Long Context Reasoning177Kimi K388.7%Chess PuzzlesEpoch AI122GPT-6 Astra72.0%ARC-AGI-2 (semi-private)ARC Prize Foundation113GPT-6 Astra95.0%ARC-AGI-1 (semi-private)ARC Prize Foundation109Claude Fable 598.5%Mystery Game PuzzlesEpoch AI62GPT-6 Astra84.0%LiveBench Data AnalysisLiveBench55GPT-6 Astra Max83.0%LiveBench ReasoningLiveBench55GPT-6 Astra Max92.7%EBR-benchEpoch AI20GPT-6 Astra76.2%Furniture AssemblyEpoch AI19GPT-6 Astra80.0%Medical Long Context Reasoning (MLCR-AA)Wisedocs and Artificial Analysis16Claude Fable 5.171.1%
Not in index
ARC-AGI-3 (semi-private)ARC Prize Foundation15—LongBench v2LongBench v2 authors14—MRCRv29—AI-Needle4—MRCR 1M4—BIG-Bench HardMirac Suzgun et al.3—CorpusQA 1M2—Discrete Reasoning Over Paragraphs2—GeneBench-Pro2—OpenAI MRCR v2 8-needle 128K-256K2—OpenAI MRCR v2 8-needle 256K-512K2—OpenAI MRCR v2 8-needle 512K-1M2—OpenAI MRCR v2 8-needle 64K-128K2—Graphwalks BFS 0K-128K1—Multimodal
60 testsBenchmarkModelsModels with a result.Top resultBest result on this test.
Artificial Analysis MMMU-ProArtificial Analysis101Claude Opus 5.587.7%MMMU ProVals AI90Claude Fable 5.190.6%Massive Multi-discipline Multimodal Understanding ProMMMU-Pro authors41Gemini 3.1 Pro83.9%CharXiv Reasoning39Claude Mythos 593.5%ScreenSpot ProKaixin Li et al.18GPT-6 Astra92.7%ERQA13Qwen3.8 Max77.8%OfficeQA ProOfficeQA Pro authors12Claude Opus 5.567.7%RealWorldQA12Qwen3.8-Flash-Next88.5%Massive Multi-discipline Multimodal UnderstandingMMMU authors11Qwen3.6 Plus86.0%SimpleVQA11Qwen3.7 Plus81.7%V*11Kimi K2.696.9%
Not in index
MathVision19—CharXiv Reasoning without tools14—VideoMMMU11—MMMU-Pro with Python9—MedXpertQA Multimodal8—SWE-bench Multimodal · swebench.comSWE-bench team8—ZeroBench8—BabyVisionMeta AI7—Multimodal Multi-disciplinary Video UnderstandingMMVU benchmark maintainers7—SWE-bench Multimodal · SWE-bench/experimentsSWE-bench team7—OmniDocBench 1.56—RefCOCO averageRefCOCO dataset authors6—Video-MME with subtitle6—LVBench5—MathVision with Python5—OCRBench V2OCRBench authors5—BabyVision with Python4—Chartography with image and code toolsSurge AI and Anthropic4—CountBench4—MLVU mean average4—Vision2Web4—AI2D test split3—BenchCAD Vision2Code voxel IoU with toolsZhang et al. and Anthropic3—CC-OCR3—PerceptionBench (Internal)3—Video-MMEVideo-MME benchmark team3—ZeroBench_main with Python3—BenchCAD Vision2Code voxel IoU without toolsZhang et al. and Anthropic2—Blueprint-Bench 2Google DeepMind2—Chartography without toolsSurge AI and Anthropic2—GDP.pdf mean criteria pass rate without toolsSurge AI and Anthropic2—Liquid image-to-JSON extraction JSON validity2—Liquid image-to-JSON extraction schema consistency F12—Liquid image-to-JSON extraction VLM judge score2—ODINW132—OfficeQA2—Video-MME without subtitle2—A Benchmark for Visual Question Answering using World KnowledgeDustin Schwenk et al.1—Anthropic biomedical-image-analysis evaluation1—CharXiv Descriptive and Reasoning Combined1—DynaMath1—GDP.pdf mean criteria pass rate with toolsSurge AI and Anthropic1—MMLongBench-Doc1—MMSearch-Plus1—MStar1—olmOCR-BenchAllen Institute for AI1—OmniDocBench1—OmniDocBench v1.6Linke Ouyang et al.1—WorldVQA ForceAnswer1—Knowledge
48 testsBenchmarkModelsModels with a result.Top resultBest result on this test.
GPQA diamondEpoch AI186GPT-6 Astra95.8%Artificial Analysis Humanity's Last ExamArtificial Analysis185Claude Opus 5.561.4%Artificial Analysis GPQA DiamondArtificial Analysis180GPT-6 Astra96.1%Artificial Analysis Omniscience Accuracy176Claude Fable 5.167.2%GPQA DiamondVals AI128Gemini 3.1 Pro Preview95.5%MMLU ProVals AI128Claude Fable 5.192.4%Graduate-Level Google-Proof Q&ADavid Rein et al.84GPT-6 Astra96.0%SimpleQA VerifiedEpoch AI67GPT-6 Astra75.6%GPQA Diamond · BenchLMDavid Rein et al.65GPT-6 Astra96.0%Humanity's Last ExamCenter for AI Safety et al.59Claude Fable 5.165.0%LiveBench LanguageLiveBench55Claude Fable 590.7%Massive Multitask Language Understanding ProfessionalYubo Wang et al.46Qwen3.7 Max89.6%Humanity's Last Exam without tools38Claude Opus 5.564.4%SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesXiaoxuan Du et al.20Claude Opus 4.695.0%HealthBench Hard11Muse Spark42.8%MMLU-Redux11Claude Opus 4.596.6%MMLU-Pro first-party comparison snapshot6Claude Opus 4.689.1%
Not in index
HealthBench ProfessionalRebecca Soskin Hicks et al.11—Massive Multitask Language UnderstandingDan Hendrycks et al.7—C-EvalC-Eval authors6—HLE-VerifiedWeiqi Zhai et al.6—LABBench2: An Improved Benchmark for AI Systems Performing Biology ResearchJon M. Laurent et al.6—BioMysteryBench Human Difficult5—HealthBench length-adjusted score5—HealthBench Professional raw score5—HealthBench raw score5—MedXpertQA Text5—MMMLU5—Artificial Analysis Openness IndexArtificial Analysis4—BioMysteryBench Human Solvable4—FrontierScience Research4—Artificial Analysis MMLU-ProArtificial Analysis3—Measuring Short-Form Factuality in Large Language ModelsJason Wei et al.3—Anthropic Protein Design evaluation2—Benchling Molecular Biology Protocols Understanding2—Chinese-SimpleQA2—LatchBio SingleCellBench2—LatchBio SpatialBench Verified2—Molecular Biology Protocols Troubleshooting2—AGIEval1—Anthropic de novo protein-binder design evaluation1—Anthropic medicinal-chemistry evaluation1—Anthropic Organic Chemistry V2 evaluation1—Anthropic protein-design library-ranking task1—Axiom Bio morphology-to-molecule matching1—Chinese Massive Multitask Language Understanding1—FrontierScience1—ProteinGym Hard1—Multilingual
22 testsBenchmarkModelsModels with a result.Top resultBest result on this test.
EuroEval PolishEuroEval189GLM-5.3-Flash68.1%EuroEval SwedishEuroEval173GPT-6 Astra77.6%EuroEval GermanEuroEval169GPT-6 Astra68.0%EuroEval PortugueseEuroEval167GPT-6 Astra74.7%EuroEval FrenchEuroEval163Gemini 3.7 Flash77.1%EuroEval SpanishEuroEval161GPT-6 Astra70.0%EuroEval ItalianEuroEval160GPT-6 Astra77.8%EuroEval DutchEuroEval154GPT-6 Astra74.7%
Not in index
MGSMVals AI70—MMLU-ProXMMLU-ProX authors12—SWE-bench Multilingual · swebench.comSWE-bench team12—NOVA-637—Artificial Analysis Global-MMLU-LiteArtificial Analysis4—INCLUDE4—PolyMath3—Global MMLUSingh et al.2—MAXIFE2—Multi-task Indic Language Understanding BenchmarkVerma et al.2—Long-context translation evaluation (Cohere)1—MKQA-11 multilingual retrievalLiquid AI1—NanoBEIR Multilingual ExtendedLiquid AI1—WMT26 General Machine Translation (all languages)WMT organizers1—Instruction
5 testsBenchmarkModelsModels with a result.Top resultBest result on this test.
Math
37 testsBenchmarkModelsModels with a result.Top resultBest result on this test.
OTIS Mock AIME 2024-2025Epoch AI168GPT-6 Astra100.0%AIMEVals AI91Gemini 3.1 Pro Preview98.1%MATH level 5Epoch AI88GPT-598.1%FrontierMath-Tiers-1-3-v2-PrivateEpoch AI71GPT-6 Astra93.7%FrontierMath-2025-02-28-PrivateEpoch AI68GPT-5.551.7%MATH 500Vals AI57Gemini 3 Pro96.4%LiveBench MathematicsLiveBench55Claude Opus 5.597.1%FrontierMath-Tier-4-2025-07-01-PrivateEpoch AI54AI Co-Mathematician47.9%FrontierMath-Tier-4-v2-PrivateEpoch AI53GPT-6 Astra97.6%ProofBench v1.1Vals AI33Claude Fable 5.1100.0%AIME 202628GLM-5.299.2%Harvard-MIT Mathematics Tournament February 202623Qwen3.7 Max97.1%American Invitational Mathematics Examination 202516MAI-Thinking-197.0%Harvard-MIT Mathematics Tournament February 202510GLM-597.5%MATH-500 Problem SetDan Hendrycks et al.7Ternary Bonsai 2 27B98.8%
Not in index
MMAnswerBench10—Harvard-MIT Mathematics Tournament November 20259—IMOAnswerBench9—Apex8—AIME25 first-party comparison snapshot6—FrontierMath-ErdosEpoch AI5—OEIS Open LiteEpoch AI5—Apex Shortlist3—Artificial Analysis MATH-500Artificial Analysis3—Grade School Math 8K3—OEIS OpenEpoch AI3—United States of America Mathematical Olympiad 2026Mathematical Association of America3—International Mathematical Olympiad 20262—American Invitational Mathematics Examination 20241—Artificial Analysis AIME 2025Artificial Analysis1—ArXivMath August 2026 with tools1—ArXivMath August 2026 without tools1—ArXivMath June 2026 with tools1—ArXivMath June 2026 without tools1—International Physics Olympiad 2025 (Theory)1—RiemannBench with tools1—RiemannBench without tools1—