Tencent
availableShows if the model has enough results for an index.Hy4 preview
Hy4 preview is a reasoning model from Tencent in the Hy4 family. 24 benchmarks count toward its score, in 6 categories.
IndexOverall score out of 100.70.2 ±4.3
CoverageShare of the index weight with results.90%
SpeedOutput tokens per second.40/s
Input / 1MUS dollars per 1M input tokens.$0.834
Output / 1MUS dollars per 1M output tokens.$2.5
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.N/A
The index is a score out of 100. The ± range shows how much it can change.
CapabilitiesScore per category, out of 100.
Out of 100Results
24 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| Graduate-Level Google-Proof Q&A | Knowledge | 92.3% | 63.5 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 92.3% | 63.5 | — | — | David Rein et al. |
| Terminal-Bench 2.1 (provider run) | Agentic | 85.4% | 74.5 | — | — | DeepSeek-AI |
| Terminal-Bench 2.1 (provider run) | Agentic | 85.4% | 74.5 | — | — | DeepSeek-AI |
| WideResearch | Agentic | 83.9% | 70.1 | — | — | Qwen |
| MCP Atlas | Agentic | 83.7% | 71.2 | — | — | OpenAI |
| BankerToolBench | Agentic | 78.6% | — | — | — | MiniMax |
| CyberGym | Agentic | 78.4% | 69.8 | — | — | Zhun Wang et al. |
| Vibe Code Bench v1.1 | Coding | 77.5% | 74.4 | OpenHands | 21 Sept 2026 | Vals AI |
| Data Research and Analysis with Complex Operations | Agentic | 77.2% | — | — | — | Anthropic |
| ProofBench v1.1 | Math | 75.0% | 79.8 | — | 21 Sept 2026 | Vals AI |
| Apex | Math | 74.2% | — | — | — | DeepSeek-AI |
| Toolathlon-Verified | Agentic | 74.1% | 72.3 | — | — | Moonshot AI |
| OfficeQA Pro | Multimodal | 66.2% | 78.5 | — | — | OfficeQA Pro authors |
| SWE-bench Pro | Coding | 65.7% | 67.5 | — | — | Xiang Deng et al. |
| DeepSWE | Agentic | 64.3% | 70.5 | — | — | Datacurve AI |
| JobBench | Agentic | 61.7% | 76.5 | — | — | Yuetai Li et al. |
| IOI | Coding | 59.3% | 71.5 | — | 21 Sept 2026 | Vals AI |
| NL2Repo | Coding | 58.9% | 71.3 | — | — | MiniMax |
| Humanity's Last Exam with tools | Agentic | 55.4% | 68.2 | — | — | DeepSeek-AI |
| Humanity's Last Exam | Knowledge | 55.4% | 75.7 | — | — | Center for AI Safety et al. |
| Terminal-Bench 2.1 | Agentic | 55.1% | 56.5 | — | 21 Sept 2026 | Vals AI |
| Code Migration | Coding | 47.4% | 75.3 | — | 21 Sept 2026 | Vals AI |
| Humanity's Last Exam without tools | Knowledge | 43.4% | 65.5 | — | — | OpenAI |
| APEX-Agents | Agentic | 37.1% | 70.3 | — | — | Moonshot AI / APEX-Agents benchmark authors |
| PostTrain Bench | Coding | 35.6% | — | — | — | Moonshot AI |
| AutomationBench | Agentic | 32.1% | 68.9 | — | — | Moonshot AI |
| SWE-Marathon | Coding | 31.9% | — | — | — | Abundant AI and BenchFlow |
| Agents' Last Exam | Agentic | 22.8% | 62.7 | — | — | DeepSeek-AI |
| ProgramBench: Can Language Models Rebuild Programs From Scratch? | Coding | 17.5% | 33.5 | — | — | John Yang et al. |
| Critical Physics Tasks | Reasoning | 16.9% | 73.8 | — | — | Artificial Analysis |
| Agent Arena task outcome | Agentic | 9.8 | 78.5 | — | 15 Sept 2026 | LMArena |
| Agent Arena command recovery | Agentic | 7.4 | 75.8 | — | 15 Sept 2026 | LMArena |
| Terminal-Bench 4.0 | Agentic | 5.1% | 63.0 | — | 21 Sept 2026 | Vals AI |
| ProgramBench | Coding | 0.0% | — | — | 21 Sept 2026 | Vals AI |
| Agent Arena steerability | Agentic | -1.2 | 66.0 | — | 15 Sept 2026 | LMArena |
24 benchmarks count, from 30 of 36 results. A grey row does not count. Too few models took that benchmark.