Hy4 preview
Released Aug 28, 2026 · knowledge cutoff unpublished
- Status
- Superseded
- Location
- China
- Modality
- Text
- Context window
- 1M
- Max output
- 64K
- Speed
- —
- Price ($/MTok in / out)
- $0.83 / $2.50
- Cost per task
- —
- Open weights
- Yes
Benchmarks3 of 10 reported
GPQA Diamond
SWE-bench Pro
Terminal-Bench 2.1
MMLU-Proretired
SWE-bench Verified
AIME (latest)retired
Humanity's Last Exam
LMArena Elo
GDPval-AA v2
ARC-AGI-2
Strengths & weaknesses
Strengths
- Gated sparse attention with a cross-layer index cache makes the full window usable in practice rather than just advertised — whole repositories fit in one prompt without chunking
- Took part in its own training run — proposed and tested changes to data strategy, evaluation and inference kernels, and the throughput gains it found were folded back into the release
- First-party endpoint speaks OpenAI Chat Completions, Responses and Anthropic Messages protocols, so existing agent harnesses point at it without a shim
Weaknesses
- Reasons longer than the task needs and re-verifies work that was already right — Tencent lists both as known issues, so agent runs burn far more tokens than the per-token price implies
- Text-only: image input is rejected outright, with vision deferred to the full Hy4 release
- Practically unrunnable at home — a 4-bit build is still around half a terabyte, so despite the open weights it is a rent-not-run model
Tencent's first 770B-total/49B-active MoE (256 routed + 1 shared expert, top-8), open-sourced 2026-08-28. GPQA Diamond 92.3 and Terminal-Bench 2.1 85.4 are Tencent's own chart figures, published as an image and transcribed by launch coverage — Artificial Analysis has not yet listed the model, so no independent run exists. HLE is left null: the with-tools figure (55.4%, text-only) is corroborated across trackers but the no-tools figure (43.4%) traces to a single transcription set. Tencent's chart does not report SWE-bench Verified; it reports SWE-bench Pro 65.7%, SWE-bench Multilingual 82.9%, DeepSWE 64.3%, and a 163-expert blind evaluation at 2.99/4 vs GLM-5.3 2.92 and Kimi K3 2.94. MMLU-Pro, AIME and ARC-AGI-2 were not published; the model is on arena.ai's Code Arena but not the text leaderboard. Pricing is Tencent's first-party TokenHub rate ($0.042 cache read), matched by OpenRouter; maxOutput 64K and a 960K max-input cap are from Tencent Cloud's docs. predecessorId is left null — Tencent extended free Hy3 access to 2026-09-30 alongside this launch and states no replacement.
For developers
API model strings
- tencent-cloud
hy4-preview
Licence
Apache License 2.0 · Permissive · commercial use permitted
Retirement
No retirement announced
Lineage
No recorded predecessor