Muse Spark 1.2

Muse Spark 1.2 is a reasoning model from Meta in the Muse Spark family. 32 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.69.2 ±3.4
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.126/s
Input / 1MUS dollars per 1M input tokens.$1.25
Output / 1MUS dollars per 1M output tokens.$4.25
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1489 (#11)

The index is a score out of 100. The ± range shows how much it can change.

3,227 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
69.4
CodingCode writing and repair.
67.3
ReasoningLogic problems and puzzles.
70.0
MultimodalTasks with images and text.
67.5
KnowledgeFacts and expert knowledge.
68.5
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
77.3
MathMath problems.
68.3

Results

32 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
LiveBench MathematicsMath91.2%69.8xhigh effort25 Jun 2026LiveBench
Artificial Analysis GPQA DiamondKnowledge90.4%61.2Artificial Analysis
LiveBench ReasoningReasoning90.0%81.0xhigh effort25 Jun 2026LiveBench
MMLU ProKnowledge88.3%59.61 Sept 2026Vals AI
VulcanBench v3Coding87.0%72.7VulcanBench contributors
SWE-benchCoding86.6%67.01 Sept 2026Vals AI
MMMU ProMultimodal86.1%67.51 Sept 2026Vals AI
Terminal-Bench 2.1 (provider run)Agentic82.9%73.0DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic82.9%73.0DeepSeek-AI
Vibe Code Bench v1.1Coding79.1%75.1OpenHands21 Sept 2026Vals AI
Artificial Analysis Long Context ReasoningReasoning79.0%62.9Artificial Analysis
LiveBench LanguageKnowledge78.6%65.4xhigh effort25 Jun 2026LiveBench
LiveBench CodingCoding77.5%66.4xhigh effort25 Jun 2026LiveBench
LiveBench Data AnalysisReasoning76.5%62.2xhigh effort25 Jun 2026LiveBench
LiveBench Instruction FollowingInstruction74.3%77.3xhigh effort25 Jun 2026LiveBench
Artificial Analysis Coding IndexCoding72.2%69.8Artificial Analysis
Terminal-Bench 2.1Agentic69.7%65.221 Sept 2026Vals AI
SimpleQA VerifiedKnowledge60.3%77.4xhigh effortEpoch AI
DeepSWEAgentic59.3%66.9Datacurve AI
LiveBench Agentic CodingAgentic57.6%70.9xhigh effort25 Jun 2026LiveBench
Artificial Analysis SciCodeCoding57.4%72.5Artificial Analysis
SkillsBenchCoding53.0%68.0OpenHands11 Sept 2026Vals AI
IOI v1Coding49.5%68.19 Aug 2026Vals AI
GDPval-AA normalizedAgentic49.1%73.3Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge45.5%73.9Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge45.4%70.0Artificial Analysis
Artificial Analysis Agentic IndexAgentic44.0%73.0Artificial Analysis
ProofBench v1.1Math43.0%66.821 Sept 2026Vals AI
Artificial Analysis Intelligence IndexKnowledge39.6%71.9Artificial Analysis
Code MigrationCoding30.0%64.121 Sept 2026Vals AI
IOICoding21.8%55.121 Sept 2026Vals AI
Critical Physics TasksReasoning17.7%75.5Artificial Analysis
FrontierSWE v2Coding12.0%61.2Proximal
Agent Arena command recoveryAgentic7.375.6xhigh effort15 Sept 2026LMArena
Terminal-Bench 4.0Agentic5.6%63.421 Sept 2026Vals AI
Agent Arena task outcomeAgentic-2.065.2xhigh effort15 Sept 2026LMArena
Agent Arena steerabilityAgentic-4.662.2xhigh effort15 Sept 2026LMArena

32 benchmarks count, from 37 of 37 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Qwen3.8-Flash-Next66.4 · FreeGPT-6 Luna68.4 · $0.5DeepSeek V4.1 Flash70.6 · $0.6

More from Meta

Muse Spark 1.375.0Muse Spark 1.167.0Muse Spark61.6Muse Glimmer 30B53.5Llama 4 Maverick33.6