Hy3
Released Jul 6, 2026 · knowledge cutoff unpublished
- Status
- Superseded
- Location
- China
- Modality
- Text
- Context window
- 262K
- Max output
- 128K
- Speed
- —
- Price ($/MTok in / out)
- $0.13 / $0.53
- Cost per task
- $0.036
- Open weights
- Yes
Benchmarks5 of 8 reported
GPQA Diamond
SWE-bench Verified
Terminal-Bench 2.1
Humanity's Last Exam
LMArena Elo
MMLU-Pro
AIME (latest)
ARC-AGI-2
Strengths & weaknesses
Strengths
- Tool calls hold their shape across third-party agent scaffoldings instead of needing harness-specific patches
- Recovers from a failed tool call rather than looping on it — the reliability bug Tencent set out to fix after the preview
- Reasoning is a per-request switch, so quick answers and deep work come off one endpoint
- Apache 2.0 with no regional carve-outs, unlike the Hunyuan releases before it
Weaknesses
- Answers confidently where it should abstain — third-party knowledge probes catch it fabricating far more often than Tencent's own hallucination figures suggest
- Text-only: no images, screenshots or documents
- Thinks at length on high effort, so agentic runs consume far more output budget than the token price implies
Successor to April's Hy3 Preview; the licence moved to Apache 2.0 with this release. GPQA Diamond 90.4, SWE-bench Verified 78.0 and Terminal-Bench 2.1 71.7 are Tencent's own figures, published as an image and transcribed by launch coverage — Artificial Analysis independently measures GPQA 89.7 and Terminal-Bench 2.1 64.4. HLE 31.6% is Artificial Analysis' no-tools run; Tencent reports 53.2% with tools. Also reported: SWE-bench Pro 57.9%, SWE-bench Multilingual 75.8%, DeepSWE 28.0%, and a blind 270-expert evaluation at 2.67/4. MMLU-Pro, AIME and ARC-AGI-2 were not published. Pricing is the OpenRouter listing — Tencent publishes no first-party per-token price sheet.
For developers
API model strings
Not researched
Licence
Not researched
Retirement
No retirement announced
Lineage
No recorded predecessor