前沿每周都在推进。我们追踪重要的模型——来自真正推动市场的八大实验室——附带来源数据和简明扼要的“为何重要”。
我们对公开基准测试的解读,每周更新。观点已清晰标注——点击任何模型可查看来源。
| 模型 | 实验室 | 发布时间 | 上下文 | 头条基准 |
|---|---|---|---|---|
| GLM-5.3 | Z.ai | 2026年8月14日 | — | Terminal-Bench 3.0 (Z.ai self-reported): 28.3, up from GLM-5.2's 4.6 (GPT-5.6 Sol 34.6) |
| Grok 4.6 | xAI | 2026年8月12日 | 500K | API price (per 1M in/out tokens, <200k prompt): $2 / $6 |
| DeepSeek V4 Pro (0813)开源权重 | DeepSeek | 2026年8月12日 | 1M | API price (per 1M in/out tokens, promo): $0.435 / $0.87 |
| Seedance 2.5视频模型 | ByteDance | 2026年7月31日 | — | Single-pass clip length: 30 seconds (up from 15) |
| DeepSeek V4 Flash (0731)开源权重 | DeepSeek | 2026年7月31日 | 1M | API price (per 1M in/out tokens): $0.14 / $0.28 |
| Claude Opus 5 | Anthropic | 2026年7月24日 | — | API price (per 1M in/out tokens): $5 / $25 |
| Kimi K3开源权重 | Moonshot AI | 2026年7月16日 | 1M | Total parameters (sparse MoE): 2.8T |
| GPT-5.6 (Sol · Terra · Luna) | OpenAI | 2026年7月9日 | — | API price, Sol (per 1M in/out tokens): $5 / $30 |
Z.ai's 14 Aug 2026 release reuses GLM-5.2's base model — every gain comes from scaled post-training on long-horizon RL environments. Large jumps on agentic coding (Terminal-Bench 3.0 4.6 → 28.3, SWE-Marathon 19.4 → 42.5) and an unplanned cyber result (CyberGym 84.5). All figures are Z.ai's own; no independent evaluation exists yet. Weights promised two weeks after launch.
为何重要: It reaches Opus 4.8's coding score on roughly 2.4× fewer output tokens, so the saving is on the bill rather than the leaderboard — but check your API calls first: `thinking.type: "disabled"` is no longer supported and will fail on glm-5.3.
xAI's 12 Aug 2026 successor to Grok 4.5: same $2/$6 API price, same 500K context, a longer post-training run aimed at long-running agents and visual/interactive work. A faster variant costs double. Knowledge cut-off is 1 Feb 2026.
为何重要: Frontier-class intelligence at the old 4.5 price — migrate off 4.5 with a one-line model swap. It is a knowledge-work pick, not the cheapest volume model and not the strongest terminal coding agent.
The deepseek-v4-pro API id now serves DeepSeek-V4-Pro-0813 weights. Calling method and promo pricing are unchanged. DeepSeek's changelog still has no Pro GA entry — the last note (31 Jul) said an official Pro release would follow soon — so this is a weight drop, not a new product.
为何重要: Harder reasoning stays cheap on the same endpoint. Keep Flash ($0.14/$0.28) for volume; do not treat 0813 as a priced relaunch until DeepSeek posts a changelog.
ByteDance's new video model generates 30-second audio-and-video clips in one pass, extends them over multiple rounds for multi-minute pieces, and edits by timestamp. At launch it runs inside Jimeng AI and Doubao Pro, with API access announced as coming via BytePlus ModelArk rather than available.
为何重要: The clearest jump yet in one-take video, but there is no general API to wire into a workflow — treat it as something to try inside a consumer video app, not a tool to rebuild your content process around until the API actually ships.
The official V4 Flash checkpoint entered public beta on the existing deepseek-v4-flash API id, replacing the April preview. The legacy deepseek-chat and deepseek-reasoner ids were retired on 2026-07-24.
为何重要: Near-frontier results at roughly a tenth of typical API prices, with MIT-licensed open weights — the budget and privacy-sensitive pick just got better without changing price.
Anthropic's new everyday flagship: close to Fable 5 on most benchmarks at half the price ($5/$25 vs $10/$50 per MTok), with a low/medium/high effort toggle to trade cost against capability.
为何重要: Frontier-level output at half the flagship price — re-check which Claude tier your workflows actually need before renewing.
Moonshot's 2.8-trillion-parameter sparse MoE — reportedly the largest open-weight model yet, with a 1M-token context window. Weights announced for late July; until then benchmark claims are vendor-reported.
为何重要: Open-weight models keep closing on the paid frontier — if you pay per-token for API work, the cheap tier just got stronger again.
OpenAI's GPT-5.6 family in three tiers — Sol (frontier), Terra (balanced), Luna (fast/cheap) — general availability across ChatGPT, Codex, and the API on Jul 9 after a government-vetted limited preview in late June.
为何重要: Three clear price tiers make it easier to match the model to the job — most small-business tasks belong on the cheap tier, not the flagship.
xAI's first model built specifically for coding and agentic work, priced aggressively under Anthropic and OpenAI flagships with a 500K context window.
为何重要: Agentic coding on a budget is now a three-way price war — worth re-testing your coding stack before renewing anything.
A Mythos-class model made safe for general use — Anthropic's most capable generally available model, sitting above the Opus tier.
为何重要: The frontier of general-purpose reasoning just moved again; capable assistants keep getting cheaper to match.
The Mythos-class model available to approved organizations without the general-use safety measures applied to Fable 5.
为何重要: Signals how fast the top tier is advancing — the same capability reaches everyone shortly after.
A strong all-round released model, widely cited as a top performer through mid-2026.
为何重要: A dependable default for hard reasoning, coding, and long-document work.
OpenAI's mid-2026 frontier update, trading the top spot with Claude Opus on many benchmarks.
为何重要: Keeps the price-for-capability race moving — good news for anyone paying per token.
A March 2026 frontier release with a 1M-token context window and strong computer-use scores.
为何重要: Million-token context means it can read whole manuals, contracts, or codebases at once.
An open-weights frontier model with a 1M+ token context window and strong coding scores.
为何重要: Open weights + very low cost make it the value pick for budget-conscious and privacy-sensitive setups.
Google's February 2026 frontier update to the Gemini 3 line, with a very large context window.
为何重要: Deep integration with Google Workspace makes it a natural fit if you live in Docs and Gmail.
A February 2026 iteration in the GPT-5 line ahead of the March 5.4 release.
为何重要: Part of the steady cadence keeping the mainstream assistant sharp.
A February 2026 Opus update (alongside Sonnet 4.6), continuing Anthropic's rapid iteration.
为何重要: Reliability gains at the same price point — worth re-testing your prompts on each bump.
The Gemini 3 flagship that opened the current generation for Google.
为何重要: Set the bar for long-context multimodal work heading into 2026.
A late-2025 Opus release that anchored Anthropic's top tier into 2026.
为何重要: The baseline many businesses standardized on before the 2026 wave.
我们每周更新一次,这与维护工具价格数据的例行工作同步。我们只列出真正影响市场的主要版本发布——前沿文本模型,现在也包括视频模型——每项数据都链接到来源。没有来源,就不列出数据。如果模型已发布但尚未可通过API使用,我们会说明,因为无法购买使用的模型不算推荐。