Benchmarks

345 tests · 143 in index

Agentic

86 tests
BenchmarkModelsModels with a result.Top resultBest result on this test.
τ²-Bench Tool-Agent-User EvaluationVictor Barres et al.132GLM-5.299.1%GDPval-AA normalized107Claude Opus 5.567.3%Terminal-Bench 2.0Vals AI65GPT-5.573.2%Terminal-Bench 2.1Vals AI63GPT-6 Astra87.3%Terminal-Bench 2.1 · BenchLM62SWE-292.8%LiveBench Agentic CodingLiveBench55DeepSeek V4.1 Flash Max77.3%Gert Labs Composite Game BenchmarkGert Labs52Claude Opus 4.873.0%Terminal-Bench 1.0Vals AI45GPT-5.263.8%BrowseComp44Atria Dawn Preview92.5%Agent Arena command recoveryLMArena43Claude Opus 513.1Agent Arena steerabilityLMArena43Claude Opus 4.811.1Agent Arena task outcomeLMArena43Claude Fable 5.119.8Terminal-Bench 2.0 · BenchLM41GPT-5.582.0%Claw-EvalBowen Ye et al.39Ornith-1.5-397B81.4%MCP Atlas38Muse Spark 1.188.1%DeepSWEDatacurve AI36Muse Spark 1.375.4%OSWorld-VerifiedTianbao Xie et al.34Qwen3.8 Max86.1%JobBenchYuetai Li et al.30Muse Spark 1.364.9%Terminal-Bench 4.0Vals AI29GPT-6 Astra57.1%CyberGymZhun Wang et al.27MiMo-V2.6-Flash95.1%APEX-Agents-AAArtificial Analysis / Mercor26Gemini 3.5 Flash47.1%τ²-bench BankingSierra Research26Qwen3.8 Max55.2%τ²-bench RetailSierra Research23Gemini 3.0 Pro85.3%Berkeley Function Calling Leaderboard v422BTL-388.5%OSWorld 2.0Mengqi Yuan et al.22GPT-6 Astra72.6%Toolathlon22Muse Spark 1.175.6%τ²-bench AirlineSierra Research22Claude Opus 4.584.0%Humanity's Last Exam with tools20Claude Opus 5.567.7%Toolathlon-Verified20Claude Opus 580.6%AutomationBench19DeepSeek V4.1 Flash54.8%ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobNeoCognition18Claude Fable 5.172.0%DeepSearchQA18Atria Dawn Preview96.0%τ²-bench TelecomSierra Research18Qwen3 Max Thinking98.2%τ³-Bench Tool-Agent-User EvaluationSierra Research18Mercury 2.596.0%Artificial Analysis EnterpriseOps-GymArtificial Analysis17Claude Fable 551.1%Agents' Last Exam15GPT-6 Astra59.3%Artificial Analysis AnalystAgentArtificial Analysis15Gemini 3.7 Flash60.0%Artificial Analysis AutomationBenchArtificial Analysis15Claude Opus 5.569.5%Terminal-Bench 4.0.0Terminal-Bench15GPT-6 Astra58.2%WideResearch15Hy4 preview83.9%Artificial Analysis GDP.pdfArtificial Analysis14GPT-6 Astra31.0%Artificial Analysis Tau3-BankingArtificial Analysis14Grok 4.650.7%ExploitGymZhun Wang et al.14GPT-6 Astra42.4%Terminal-Bench 3.0Ryan Marten et al.13Claude Opus 542.7%Artificial Analysis Harvey LAB-AAArtificial Analysis12Kimi K394.6%VITA-BenchMeituan LongCat Team12Qwen3.7 Max47.9%APEX-Agents10Grok 4.657.5%QwenClawBench10Qwen3.7 Max64.3%
Not in index
ResearchClawBench19AndroidWorld9Artificial Analysis ITBench-AAArtificial Analysis8DeepPlanningDeepPlanning authors7PinchBenchKilo Code6Terminal-Bench Hard6MCP-Tasks5CoWorkBench4Data Research and Analysis with Complex Operations4MLE-Bench Lite3Terminal-Bench-Science 0.1Terminal-Bench-Science Team3Toolathlon Verified average assistant turns3Toolathlon Verified Pass cubed3Toolathlon Verified Pass@33WebArena-Verified Browser Agent BenchmarkAmine El Hattami et al.3BankerToolBench2Legal Agent Benchmark all-pass rate — Harvey held-out set2Legal Agent Benchmark mean criterion-pass rate — Harvey held-out set2MCP-Atlas mean claim coverage2MM-ClawBench2OSWorld2SpreadsheetBench 22terminalBench32WebVoyager2Berkeley Function Calling Leaderboard v3Shishir G. Patil et al.1CTI-REALM1CWE-BenchCollinear1CybenchAndy K. Zhang et al.1DECK-Bench (Internal)1GDPval rubrics1General AI Assistants1Kimi Claw 24/7 Bench1Legal Agent Benchmark all-pass rate — Anthropic harness1Legal Agent Benchmark mean criterion-pass rate — Anthropic harness1MCPMark-VerifiedMCPMark1MobileWorld1Multi-Agent BrowseComp — 10-agent team prerelease configuration1τ²-Bench Airline DomainVictor Barres et al.1

Coding

62 tests
BenchmarkModelsModels with a result.Top resultBest result on this test.
LiveCodeBenchVals AI136Claude Fable 5.190.5%Artificial Analysis SciCodeArtificial Analysis94Claude Opus 5.566.9%Vibe Code Bench v1.1Vals AI92Claude Fable 590.4%SWE-benchVals AI84Claude Opus 597.0%Software Engineering Benchmark VerifiedCarlos E. Jimenez et al.75Claude Opus 596.0%SWE-bench ProXiang Deng et al.71Claude Opus 5.589.9%SWE-bench Verified · SWE-bench/experimentsSWE-bench team62Claude Opus 4.576.8%IOI v1Vals AI60Claude Opus 591.7%Code MigrationVals AI58GPT-6 Astra67.7%LiveBench CodingLiveBench55Claude Opus 5.589.3%SWE-bench Verified · swebench.comSWE-bench team46Claude 4.5 Opus74.4%SkillsBenchVals AI34DeepSeek V4.1 Flash69.8%LiveCodeBench v6LiveCodeBench maintainers31Sakana Fugu-Ultra93.2%SWE-Bench verifiedEpoch AI31Claude Opus 4.783.5%NL2Repo29DeepSeek V4.1 Flash65.4%IOIVals AI28GPT-6 Astra100.0%Scientific Code Benchmark27Sakana Fugu60.1%cursorBench3218Claude Fable 5.173.4%Vibe Code Bench 1-100Vals AI18Claude Opus 528.5%React Native EvalsCallstack16Composer 296.1%FrontierSWE v2Proximal15GPT-6 Astra65.5%SWE-bench Lite · swebench.comSWE-bench team15Claude 4 Sonnet56.7%FrontierCode 1.1 Main14Claude Opus 5.554.4%VulcanBench v3VulcanBench contributors14Grok 4.589.9%SWE-bench Lite · SWE-bench/experimentsSWE-bench team13Claude 4 Sonnet57.5%OpenHarmony Bench v1.0OpenHarmony Bench authors12GLM-5.360.8%ProgramBench: Can Language Models Rebuild Programs From Scratch?John Yang et al.12Claude Opus 593.0%LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeNaman Jain et al.7Qwen3.7 Max91.6%SWE-bench Verified (mini-swe-agent-v2)5Claude Opus 4.675.6%
Not in index
ProgramBenchVals AI42SWE-bench Multilingual · SWE-bench/experimentsSWE-bench team14SWE-RebenchNebius13LiveCodeBench ProLiveCodeBench Pro authors8MirrorCodeEpoch AI8cursorBench317FrontierCode 1.1 Extended7SWE-bench (full test split)SWE-bench team7cursorBench406Bug Hunt BenchPawel Huryn4MLS-Bench LiteMLS-Bench4SWE-MarathonAbundant AI and BenchFlow4FrontierSWEEvan Chu et al.3PostTrain Bench3VulcanBench Coding Intelligence Index v1VulcanBench contributors3Artificial Analysis LiveCodeBenchArtificial Analysis2DeepSeek DSBench FullStack2DeepSeek DSBench Hard2Evaluating Large Language Models Trained on CodeMark Chen et al.2Kimi Code Bench v22LiveCodeBench Pass@1 with Chain-of-Thought2LiveCodeBench v5LiveCodeBench maintainers2BigCodeBench1EEBench V1 core corpusatopile1KernelBench Hard1Multi-SWE Bench1PaperBench1ProgramBench hidden-test pass rate after episode 1Yang et al.1Spider 2.0-LiteSpider 2.0 authors1SVG-Bench1SWE-Atlas Codebase QnA1VIBE V21VIBE-Pro1

Reasoning

25 tests
BenchmarkModelsModels with a result.Top resultBest result on this test.

Multimodal

60 tests
BenchmarkModelsModels with a result.Top resultBest result on this test.
Artificial Analysis MMMU-ProArtificial Analysis101Claude Opus 5.587.7%MMMU ProVals AI90Claude Fable 5.190.6%Massive Multi-discipline Multimodal Understanding ProMMMU-Pro authors41Gemini 3.1 Pro83.9%CharXiv Reasoning39Claude Mythos 593.5%ScreenSpot ProKaixin Li et al.18GPT-6 Astra92.7%ERQA13Qwen3.8 Max77.8%OfficeQA ProOfficeQA Pro authors12Claude Opus 5.567.7%RealWorldQA12Qwen3.8-Flash-Next88.5%Massive Multi-discipline Multimodal UnderstandingMMMU authors11Qwen3.6 Plus86.0%SimpleVQA11Qwen3.7 Plus81.7%V*11Kimi K2.696.9%
Not in index
MathVision19CharXiv Reasoning without tools14VideoMMMU11MMMU-Pro with Python9MedXpertQA Multimodal8SWE-bench Multimodal · swebench.comSWE-bench team8ZeroBench8BabyVisionMeta AI7Multimodal Multi-disciplinary Video UnderstandingMMVU benchmark maintainers7SWE-bench Multimodal · SWE-bench/experimentsSWE-bench team7OmniDocBench 1.56RefCOCO averageRefCOCO dataset authors6Video-MME with subtitle6LVBench5MathVision with Python5OCRBench V2OCRBench authors5BabyVision with Python4Chartography with image and code toolsSurge AI and Anthropic4CountBench4MLVU mean average4Vision2Web4AI2D test split3BenchCAD Vision2Code voxel IoU with toolsZhang et al. and Anthropic3CC-OCR3PerceptionBench (Internal)3Video-MMEVideo-MME benchmark team3ZeroBench_main with Python3BenchCAD Vision2Code voxel IoU without toolsZhang et al. and Anthropic2Blueprint-Bench 2Google DeepMind2Chartography without toolsSurge AI and Anthropic2GDP.pdf mean criteria pass rate without toolsSurge AI and Anthropic2Liquid image-to-JSON extraction JSON validity2Liquid image-to-JSON extraction schema consistency F12Liquid image-to-JSON extraction VLM judge score2ODINW132OfficeQA2Video-MME without subtitle2A Benchmark for Visual Question Answering using World KnowledgeDustin Schwenk et al.1Anthropic biomedical-image-analysis evaluation1CharXiv Descriptive and Reasoning Combined1DynaMath1GDP.pdf mean criteria pass rate with toolsSurge AI and Anthropic1MMLongBench-Doc1MMSearch-Plus1MStar1olmOCR-BenchAllen Institute for AI1OmniDocBench1OmniDocBench v1.6Linke Ouyang et al.1WorldVQA ForceAnswer1

Knowledge

48 tests
BenchmarkModelsModels with a result.Top resultBest result on this test.
GPQA diamondEpoch AI186GPT-6 Astra95.8%Artificial Analysis Humanity's Last ExamArtificial Analysis185Claude Opus 5.561.4%Artificial Analysis GPQA DiamondArtificial Analysis180GPT-6 Astra96.1%Artificial Analysis Omniscience Accuracy176Claude Fable 5.167.2%GPQA DiamondVals AI128Gemini 3.1 Pro Preview95.5%MMLU ProVals AI128Claude Fable 5.192.4%Graduate-Level Google-Proof Q&ADavid Rein et al.84GPT-6 Astra96.0%SimpleQA VerifiedEpoch AI67GPT-6 Astra75.6%GPQA Diamond · BenchLMDavid Rein et al.65GPT-6 Astra96.0%Humanity's Last ExamCenter for AI Safety et al.59Claude Fable 5.165.0%LiveBench LanguageLiveBench55Claude Fable 590.7%Massive Multitask Language Understanding ProfessionalYubo Wang et al.46Qwen3.7 Max89.6%Humanity's Last Exam without tools38Claude Opus 5.564.4%SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesXiaoxuan Du et al.20Claude Opus 4.695.0%HealthBench Hard11Muse Spark42.8%MMLU-Redux11Claude Opus 4.596.6%MMLU-Pro first-party comparison snapshot6Claude Opus 4.689.1%
Not in index
HealthBench ProfessionalRebecca Soskin Hicks et al.11Massive Multitask Language UnderstandingDan Hendrycks et al.7C-EvalC-Eval authors6HLE-VerifiedWeiqi Zhai et al.6LABBench2: An Improved Benchmark for AI Systems Performing Biology ResearchJon M. Laurent et al.6BioMysteryBench Human Difficult5HealthBench length-adjusted score5HealthBench Professional raw score5HealthBench raw score5MedXpertQA Text5MMMLU5Artificial Analysis Openness IndexArtificial Analysis4BioMysteryBench Human Solvable4FrontierScience Research4Artificial Analysis MMLU-ProArtificial Analysis3Measuring Short-Form Factuality in Large Language ModelsJason Wei et al.3Anthropic Protein Design evaluation2Benchling Molecular Biology Protocols Understanding2Chinese-SimpleQA2LatchBio SingleCellBench2LatchBio SpatialBench Verified2Molecular Biology Protocols Troubleshooting2AGIEval1Anthropic de novo protein-binder design evaluation1Anthropic medicinal-chemistry evaluation1Anthropic Organic Chemistry V2 evaluation1Anthropic protein-design library-ranking task1Axiom Bio morphology-to-molecule matching1Chinese Massive Multitask Language Understanding1FrontierScience1ProteinGym Hard1

Multilingual

22 tests
BenchmarkModelsModels with a result.Top resultBest result on this test.

Instruction

5 tests
BenchmarkModelsModels with a result.Top resultBest result on this test.

Math

37 tests
BenchmarkModelsModels with a result.Top resultBest result on this test.
OTIS Mock AIME 2024-2025Epoch AI168GPT-6 Astra100.0%AIMEVals AI91Gemini 3.1 Pro Preview98.1%MATH level 5Epoch AI88GPT-598.1%FrontierMath-Tiers-1-3-v2-PrivateEpoch AI71GPT-6 Astra93.7%FrontierMath-2025-02-28-PrivateEpoch AI68GPT-5.551.7%MATH 500Vals AI57Gemini 3 Pro96.4%LiveBench MathematicsLiveBench55Claude Opus 5.597.1%FrontierMath-Tier-4-2025-07-01-PrivateEpoch AI54AI Co-Mathematician47.9%FrontierMath-Tier-4-v2-PrivateEpoch AI53GPT-6 Astra97.6%ProofBench v1.1Vals AI33Claude Fable 5.1100.0%AIME 202628GLM-5.299.2%Harvard-MIT Mathematics Tournament February 202623Qwen3.7 Max97.1%American Invitational Mathematics Examination 202516MAI-Thinking-197.0%Harvard-MIT Mathematics Tournament February 202510GLM-597.5%MATH-500 Problem SetDan Hendrycks et al.7Ternary Bonsai 2 27B98.8%
Not in index
MMAnswerBench10Harvard-MIT Mathematics Tournament November 20259IMOAnswerBench9Apex8AIME25 first-party comparison snapshot6FrontierMath-ErdosEpoch AI5OEIS Open LiteEpoch AI5Apex Shortlist3Artificial Analysis MATH-500Artificial Analysis3Grade School Math 8K3OEIS OpenEpoch AI3United States of America Mathematical Olympiad 2026Mathematical Association of America3International Mathematical Olympiad 20262American Invitational Mathematics Examination 20241Artificial Analysis AIME 2025Artificial Analysis1ArXivMath August 2026 with tools1ArXivMath August 2026 without tools1ArXivMath June 2026 with tools1ArXivMath June 2026 without tools1International Physics Olympiad 2025 (Theory)1RiemannBench with tools1RiemannBench without tools1