Wait Which Model?
← Back to directory
Tencent

Hy4 preview

Released Aug 28, 2026 · knowledge cutoff unpublished

Status
Superseded
Location
China
Modality
Text
Context window
1M
Max output
64K
Speed
Price ($/MTok in / out)
$0.83 / $2.50
Cost per task
Open weights
Yes
Benchmarks3 of 10 reported

GPQA Diamond

92.3%

SWE-bench Pro

65.7%

Terminal-Bench 2.1

85.4%

MMLU-Proretired

SWE-bench Verified

AIME (latest)retired

Humanity's Last Exam

LMArena Elo

GDPval-AA v2

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Gated sparse attention with a cross-layer index cache makes the full window usable in practice rather than just advertised — whole repositories fit in one prompt without chunking
  • Took part in its own training run — proposed and tested changes to data strategy, evaluation and inference kernels, and the throughput gains it found were folded back into the release
  • First-party endpoint speaks OpenAI Chat Completions, Responses and Anthropic Messages protocols, so existing agent harnesses point at it without a shim

Weaknesses

  • Reasons longer than the task needs and re-verifies work that was already right — Tencent lists both as known issues, so agent runs burn far more tokens than the per-token price implies
  • Text-only: image input is rejected outright, with vision deferred to the full Hy4 release
  • Practically unrunnable at home — a 4-bit build is still around half a terabyte, so despite the open weights it is a rent-not-run model

Tencent's first 770B-total/49B-active MoE (256 routed + 1 shared expert, top-8), open-sourced 2026-08-28. GPQA Diamond 92.3 and Terminal-Bench 2.1 85.4 are Tencent's own chart figures, published as an image and transcribed by launch coverage — Artificial Analysis has not yet listed the model, so no independent run exists. HLE is left null: the with-tools figure (55.4%, text-only) is corroborated across trackers but the no-tools figure (43.4%) traces to a single transcription set. Tencent's chart does not report SWE-bench Verified; it reports SWE-bench Pro 65.7%, SWE-bench Multilingual 82.9%, DeepSWE 64.3%, and a 163-expert blind evaluation at 2.99/4 vs GLM-5.3 2.92 and Kimi K3 2.94. MMLU-Pro, AIME and ARC-AGI-2 were not published; the model is on arena.ai's Code Arena but not the text leaderboard. Pricing is Tencent's first-party TokenHub rate ($0.042 cache read), matched by OpenRouter; maxOutput 64K and a 960K max-input cap are from Tencent Cloud's docs. predecessorId is left null — Tencent extended free Hy3 access to 2026-09-30 alongside this launch and states no replacement.

For developers

API model strings

  • tencent-cloudhy4-preview

Licence

Apache License 2.0 · Permissive · commercial use permitted

Retirement

No retirement announced

Lineage

No recorded predecessor

News