Ornith-1.5-35B-A3B

Ornith-1.5-35B-A3B is a reasoning model from Ornith AI in the Ornith 1.5 family. 14 benchmarks count toward its score, in 3 categories.

partialShows if the model has enough results for an index.
IndexOverall score out of 100.Unranked
CoverageShare of the index weight with results.55%
SpeedOutput tokens per second.
Input / 1MUS dollars per 1M input tokens.Free
Output / 1MUS dollars per 1M output tokens.Free
ContextMaximum tokens in one request.262K
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
55.9
CodingCode writing and repair.
61.3
ReasoningLogic problems and puzzles.
N/A
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
55.5
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
N/A

Results

14 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
Graduate-Level Google-Proof Q&AKnowledge89.2%60.6David Rein et al.
GPQA DiamondKnowledge89.2%60.6David Rein et al.
Software Engineering Benchmark VerifiedCoding79.0%60.9Carlos E. Jimenez et al.
Claw-EvalAgentic72.5%72.3Bowen Ye et al.
MCP AtlasAgentic70.2%61.0OpenAI
Terminal-Bench 2.1 (provider run)Agentic67.8%64.1DeepSeek-AI
WideResearchAgentic67.8%52.6Qwen
Terminal-Bench 2.1 (provider run)Agentic67.8%64.1DeepSeek-AI
BrowseCompAgentic67.6%56.6OpenAI
SWE-bench ProCoding59.6%61.6Xiang Deng et al.
Toolathlon-VerifiedAgentic48.7%50.8Moonshot AI
NL2RepoCoding46.2%61.2MiniMax
Humanity's Last Exam with toolsAgentic33.4%47.0DeepSeek-AI
Humanity's Last ExamKnowledge25.6%50.5Center for AI Safety et al.
Humanity's Last Exam without toolsKnowledge25.6%50.5OpenAI
DeepSWEAgentic22.0%39.8Datacurve AI
Terminal-Bench 3.0Agentic5.1%58.8Ryan Marten et al.

14 benchmarks count, from 17 of 17 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

More from Ornith AI

Ornith-1.5-397BUnrankedOrnith-1.5-9BUnranked