Hy4 preview

Hy4 preview is a reasoning model from Tencent in the Hy4 family. 24 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.70.2 ±4.3
CoverageShare of the index weight with results.90%
SpeedOutput tokens per second.40/s
Input / 1MUS dollars per 1M input tokens.$0.834
Output / 1MUS dollars per 1M output tokens.$2.5
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
70.3
CodingCode writing and repair.
65.6
ReasoningLogic problems and puzzles.
73.8
MultimodalTasks with images and text.
78.5
KnowledgeFacts and expert knowledge.
67.0
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
79.8

Results

24 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
Graduate-Level Google-Proof Q&AKnowledge92.3%63.5David Rein et al.
GPQA DiamondKnowledge92.3%63.5David Rein et al.
Terminal-Bench 2.1 (provider run)Agentic85.4%74.5DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic85.4%74.5DeepSeek-AI
WideResearchAgentic83.9%70.1Qwen
MCP AtlasAgentic83.7%71.2OpenAI
BankerToolBenchAgentic78.6%MiniMax
CyberGymAgentic78.4%69.8Zhun Wang et al.
Vibe Code Bench v1.1Coding77.5%74.4OpenHands21 Sept 2026Vals AI
Data Research and Analysis with Complex OperationsAgentic77.2%Anthropic
ProofBench v1.1Math75.0%79.821 Sept 2026Vals AI
ApexMath74.2%DeepSeek-AI
Toolathlon-VerifiedAgentic74.1%72.3Moonshot AI
OfficeQA ProMultimodal66.2%78.5OfficeQA Pro authors
SWE-bench ProCoding65.7%67.5Xiang Deng et al.
DeepSWEAgentic64.3%70.5Datacurve AI
JobBenchAgentic61.7%76.5Yuetai Li et al.
IOICoding59.3%71.521 Sept 2026Vals AI
NL2RepoCoding58.9%71.3MiniMax
Humanity's Last Exam with toolsAgentic55.4%68.2DeepSeek-AI
Humanity's Last ExamKnowledge55.4%75.7Center for AI Safety et al.
Terminal-Bench 2.1Agentic55.1%56.521 Sept 2026Vals AI
Code MigrationCoding47.4%75.321 Sept 2026Vals AI
Humanity's Last Exam without toolsKnowledge43.4%65.5OpenAI
APEX-AgentsAgentic37.1%70.3Moonshot AI / APEX-Agents benchmark authors
PostTrain BenchCoding35.6%Moonshot AI
AutomationBenchAgentic32.1%68.9Moonshot AI
SWE-MarathonCoding31.9%Abundant AI and BenchFlow
Agents' Last ExamAgentic22.8%62.7DeepSeek-AI
ProgramBench: Can Language Models Rebuild Programs From Scratch?Coding17.5%33.5John Yang et al.
Critical Physics TasksReasoning16.9%73.8Artificial Analysis
Agent Arena task outcomeAgentic9.878.515 Sept 2026LMArena
Agent Arena command recoveryAgentic7.475.815 Sept 2026LMArena
Terminal-Bench 4.0Agentic5.1%63.021 Sept 2026Vals AI
ProgramBenchCoding0.0%21 Sept 2026Vals AI
Agent Arena steerabilityAgentic-1.266.015 Sept 2026LMArena

24 benchmarks count, from 30 of 36 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

GPT-6 Luna68.4 · $0.5DeepSeek V4.1 Flash70.6 · $0.6MiMo-V2.6-Pro73.5 · $0.87

More from Tencent

Hy357.5Hy3 Preview55.6AuKUnrankedHunyuan A13B InstructUnrankedHy ASR 3.0 previewUnranked