Beam

Beam is a reasoning model from Reflection AI. 12 benchmarks count toward its index, in 5 categories.

  • Reasoning
  • Input: text
  • Released 2026-10
availableShows if the model has enough results for an index.
IndexOverall score out of 100.61.3±7.7Median 50.7
CoverageShare of the index weight with results.70%
SpeedOutput tokens per second.—Median 62/s
First tokenSeconds to the first output token.N/AMedian 1.6 s
Input / 1MUS dollars per 1M input tokens.N/AMedian $0.6
Output / 1MUS dollars per 1M output tokens.N/AMedian $2.5
ContextMaximum tokens in one request.N/AMedian 262K
EloLMArena rating and rank.N/AMedian 1427

Capabilities

Score per category, out of 100.
Out of 100
AgenticMulti-step tasks with tools.5 counted62.945 / 184
CodingCode writing and repair.3 counted62.442 / 165
ReasoningLogic problems and puzzles.None countedN/A—
MultimodalTasks with images and text.None countedN/A—
KnowledgeFacts and expert knowledge.2 counted61.049 / 228
MultilingualTasks in many languages.None countedN/A—
InstructionTasks with strict rules in the prompt.1 counted65.3—
MathMath problems.1 counted55.2—

Results

12 counted
BenchmarkThe test name.ResultThe score from the source.PlacePlace among results on this benchmark.
AIME 2026Math97.8%2 / 29
SWE-bench ProCoding65.5%11 / 73
Instruction Following BenchmarkInstruction79.7%11 / 43
MCP AtlasAgentic78.7%10 / 39
Scientific Code BenchmarkCoding49.7%8 / 28
GPQA Diamond · BenchLMKnowledge90.5%23 / 66
BrowseCompAgentic77.4%23 / 45
DeepSearchQAAgentic80.1%11 / 19
DeepSWEAgentic44.4%38 / 41
LongBench v2Reasoning65.5%—

Sources

BenchLM benchmark aggregationCC BY-NC 4.0 · Data from BenchLM.ai