ERNIE 5.1
Released May 8, 2026 · knowledge cutoff unpublished
- Status
- Unknown
- Location
- China
- Modality
- Text
- Context window
- 128K
- Max output
- 66K
- Speed
- —
- Price ($/MTok in / out)
- $0.59 / $2.65
- Cost per task
- —
- Open weights
- No
Benchmarks2 of 8 reported
AIME (latest)
LMArena Elo
MMLU-Pro
GPQA Diamond
SWE-bench Verified
Terminal-Bench 2.1
Humanity's Last Exam
ARC-AGI-2
Strengths & weaknesses
Strengths
- Claimed #1 Chinese model on both the LMArena Text and Search leaderboards at launch, reaching #4 globally on Search
- Reaches ERNIE 5.0-beating benchmark results while training at roughly 6% of the pre-training compute of comparable frontier models, by extracting the strongest sub-network from ERNIE 5.0's elastic sub-model matrix rather than training from scratch
- Agentic evaluation (tau-cubed-bench, SpreadsheetBench-Verified) reported ahead of DeepSeek-V4-Pro; math accuracy (AIME26 with tools) among the highest reported of any model
Weaknesses
- Trails GPT-5.5 and Claude Opus 4.7 on MMLU-Pro broad-knowledge testing despite leading on math and agentic evaluations
- Text-only — dropped ERNIE 5.0's native image/audio/video input, so multimodal work still requires pairing it with another model
- Developers flagged tool-call loops and rough edges in agentic workflows at launch, echoing ERNIE 5.0's production reliability complaints
Officially released 2026-05-08 (ernie.baidu.com/blog/posts/ernie-5.1-0508-release), with wider rollout at the Create 2026 developer conference May 13-14. AIME26-with-tools figure (99.6%) and the Arena Search leaderboard rank/score (4th globally, 1,223, #1 Chinese model) are Baidu's own self-reported figures — MMLU-Pro and GPQA Diamond claims ('approaches leading closed-source models') were not given as exact numbers in Baidu's release post, so left null rather than lifted from third-party trackers whose figures (e.g. one aggregator's MMLU-Pro 85.6/GPQA 82.1) could not be corroborated against an official source. lmarenaElo 1467 is the official arena.ai Text leaderboard score (independent measurement), ranked immediately above ERNIE-5.0-preview-1203 (1449) as the highest-scoring ERNIE model tracked there. SWE-bench, Terminal-Bench, HLE and ARC-AGI-2 were not reported by Baidu or found on any independent leaderboard/tracker. Third-party coverage frequently calls 5.1 a 'successor' to 5.0, but Baidu's own material positions ERNIE 5.0 as the continuing omni-modal flagship and 5.1 as a separate text specialist extracted from its sub-model matrix, not a stated replacement — predecessorId left null. Baidu discloses no knowledge-cutoff date. Closed weights; no license record.
For developers
API model strings
- baidu-qianfan
ernie-5.1
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor