Dapols
方案AI 对比招聘定价工具基准测试博客
登录找出我的 AI 流程
Dapols

The job big companies pay a forward-deployed engineer six figures to do — as a tool, for two figures.

加入价格观察名单

已从供应商页面核实。我们目前尚未发送邮件——留下地址,开始时你会第一时间知道。

产品

  • 企业 AI 方案
  • AI计划查找器
  • AI 技能库
  • AI工具
  • 定价

探索

  • AI 对比招聘
  • 技术栈成本检查器
  • ROI计算器
  • 与您的工具兼容
  • 对比
  • 替代方案
  • 按行业
  • 按预算

产品与服务

  • AI 部署方案
  • Bigger or more complex? Tell us.
  • 规模更大或更复杂?

公司

  • 博客
  • AI 模型追踪器
  • 提交工具
  • 联系
  • 隐私
  • 服务条款

信任与方法论

  • 关于我们
  • 方法论
  • 我们如何排名 AI 工具
  • 前线部署的 AI
  • 附属链接披露
  • AI 工具定价更新
  • AI 价格指数
  • 安全与数据隐私

© 2026 Dapols. 保留所有权利。

support@dapols.comX

让 AI 真正干活——一次一条看得见成效的流程。

AI 模型追踪器

每一次重大 AI 模型发布。追踪、验证、解释。

前沿每周都在推进。我们追踪重要的模型——来自真正推动市场的八大实验室——附带来源数据和简明扼要的“为何重要”。

每周更新· 最新发布: 2026年8月14日

当前最佳

我们对公开基准测试的解读,每周更新。观点已清晰标注——点击任何模型可查看来源。

Best overall (general use)Claude Fable 5 Tops the LiveBench leaderboard (overall 83.0, 2026-07-31 snapshot). Grok 4.6 is not on that snapshot yet.Best frontier per dollarGrok 4.6 AA Intelligence Index 61, tied with GPT-5.6 Sol Max, at $2/$6 per 1M tokens (xAI, 12 Aug 2026). Weak on Terminal-Bench v3 (26%).Best for codingClaude Opus 5 Leads LiveBench agentic coding (65.2, 2026-07-31 snapshot) at half Fable 5's price; Fable 5 still holds the raw coding score (86.0).Best value / open-weightsDeepSeek V4 Flash (0731) Still the volume pick at $0.14/$0.28. V4 Pro-0813 is new weights on the same API id and promo price — not a LiveBench update yet.

近期前沿模型

模型实验室发布时间上下文头条基准
GLM-5.3Z.ai2026年8月14日—Terminal-Bench 3.0 (Z.ai self-reported): 28.3, up from GLM-5.2's 4.6 (GPT-5.6 Sol 34.6)
Grok 4.6xAI2026年8月12日500KAPI price (per 1M in/out tokens, <200k prompt): $2 / $6
DeepSeek V4 Pro (0813)开源权重DeepSeek2026年8月12日1MAPI price (per 1M in/out tokens, promo): $0.435 / $0.87
Seedance 2.5视频模型ByteDance2026年7月31日—Single-pass clip length: 30 seconds (up from 15)
DeepSeek V4 Flash (0731)开源权重DeepSeek2026年7月31日1MAPI price (per 1M in/out tokens): $0.14 / $0.28
Claude Opus 5Anthropic2026年7月24日—API price (per 1M in/out tokens): $5 / $25
Kimi K3开源权重Moonshot AI2026年7月16日1MTotal parameters (sparse MoE): 2.8T
GPT-5.6 (Sol · Terra · Luna)OpenAI2026年7月9日—API price, Sol (per 1M in/out tokens): $5 / $30

时间线

2026年8月

ZA
GLM-5.3Z.ai· 2026年8月14日

Z.ai's 14 Aug 2026 release reuses GLM-5.2's base model — every gain comes from scaled post-training on long-horizon RL environments. Large jumps on agentic coding (Terminal-Bench 3.0 4.6 → 28.3, SWE-Marathon 19.4 → 42.5) and an unplanned cyber result (CyberGym 84.5). All figures are Z.ai's own; no independent evaluation exists yet. Weights promised two weeks after launch.

为何重要: It reaches Opus 4.8's coding score on roughly 2.4× fewer output tokens, so the saving is on the bill rather than the leaderboard — but check your API calls first: `thinking.type: "disabled"` is no longer supported and will fail on glm-5.3.

Terminal-Bench 3.0 (Z.ai self-reported): 28.3, up from GLM-5.2's 4.6 (GPT-5.6 Sol 34.6)Z.ai Code Bench, High effort — output tokens per task: 31.4% at ~50K vs Claude Opus 4.8's 29.5% at ~120KGLM Coding Plan off-peak quota rate: 50% of standard points outside 14:00–18:00 UTC+8, Mon–Fri来源
XA
Grok 4.6xAI· 2026年8月12日

xAI's 12 Aug 2026 successor to Grok 4.5: same $2/$6 API price, same 500K context, a longer post-training run aimed at long-running agents and visual/interactive work. A faster variant costs double. Knowledge cut-off is 1 Feb 2026.

为何重要: Frontier-class intelligence at the old 4.5 price — migrate off 4.5 with a one-line model swap. It is a knowledge-work pick, not the cheapest volume model and not the strongest terminal coding agent.

500K 上下文API price (per 1M in/out tokens, <200k prompt): $2 / $6Artificial Analysis Intelligence Index: 61 (tied GPT-5.6 Sol Max; Fable 5 Max 62)Terminal-Bench v3.0: 26% (Sol Max 34.6%, Fable 5 Max 34.1%)来源
DS
DeepSeek V4 Pro (0813)DeepSeek· 2026年8月12日

The deepseek-v4-pro API id now serves DeepSeek-V4-Pro-0813 weights. Calling method and promo pricing are unchanged. DeepSeek's changelog still has no Pro GA entry — the last note (31 Jul) said an official Pro release would follow soon — so this is a weight drop, not a new product.

为何重要: Harder reasoning stays cheap on the same endpoint. Keep Flash ($0.14/$0.28) for volume; do not treat 0813 as a priced relaunch until DeepSeek posts a changelog.

1M 上下文API price (per 1M in/out tokens, promo): $0.435 / $0.87Model version on the existing API id: DeepSeek-V4-Pro-0813来源

2026年7月

BD
Seedance 2.5视频模型ByteDance· 2026年7月31日

ByteDance's new video model generates 30-second audio-and-video clips in one pass, extends them over multiple rounds for multi-minute pieces, and edits by timestamp. At launch it runs inside Jimeng AI and Doubao Pro, with API access announced as coming via BytePlus ModelArk rather than available.

为何重要: The clearest jump yet in one-take video, but there is no general API to wire into a workflow — treat it as something to try inside a consumer video app, not a tool to rebuild your content process around until the API actually ships.

Single-pass clip length: 30 seconds (up from 15)Reference material per pass: 30 images, 10 videos, 10 audio clips来源
DS
DeepSeek V4 Flash (0731)DeepSeek· 2026年7月31日

The official V4 Flash checkpoint entered public beta on the existing deepseek-v4-flash API id, replacing the April preview. The legacy deepseek-chat and deepseek-reasoner ids were retired on 2026-07-24.

为何重要: Near-frontier results at roughly a tenth of typical API prices, with MIT-licensed open weights — the budget and privacy-sensitive pick just got better without changing price.

1M 上下文API price (per 1M in/out tokens): $0.14 / $0.28Terminal Bench 2.1: 82.7SWE-bench Verified (thinking max): 79.0%来源
AN
Claude Opus 5Anthropic· 2026年7月24日

Anthropic's new everyday flagship: close to Fable 5 on most benchmarks at half the price ($5/$25 vs $10/$50 per MTok), with a low/medium/high effort toggle to trade cost against capability.

为何重要: Frontier-level output at half the flagship price — re-check which Claude tier your workflows actually need before renewing.

API price (per 1M in/out tokens): $5 / $25LiveBench overall (2026-07-31 snapshot): 80.1来源
KM
Kimi K3Moonshot AI· 2026年7月16日

Moonshot's 2.8-trillion-parameter sparse MoE — reportedly the largest open-weight model yet, with a 1M-token context window. Weights announced for late July; until then benchmark claims are vendor-reported.

为何重要: Open-weight models keep closing on the paid frontier — if you pay per-token for API work, the cheap tier just got stronger again.

1M 上下文Total parameters (sparse MoE): 2.8T来源
OA
GPT-5.6 (Sol · Terra · Luna)OpenAI· 2026年7月9日

OpenAI's GPT-5.6 family in three tiers — Sol (frontier), Terra (balanced), Luna (fast/cheap) — general availability across ChatGPT, Codex, and the API on Jul 9 after a government-vetted limited preview in late June.

为何重要: Three clear price tiers make it easier to match the model to the job — most small-business tasks belong on the cheap tier, not the flagship.

API price, Sol (per 1M in/out tokens): $5 / $30API price, Terra (per 1M in/out tokens): $2 / $12来源
XA
Grok 4.5xAI· 2026年7月8日

xAI's first model built specifically for coding and agentic work, priced aggressively under Anthropic and OpenAI flagships with a 500K context window.

为何重要: Agentic coding on a budget is now a three-way price war — worth re-testing your coding stack before renewing anything.

500K 上下文API price (per 1M in/out tokens): $2 / $6Terminal-Bench 2.1: 83.3%来源
AN
Claude Fable 5Anthropic· 2026年7月1日

A Mythos-class model made safe for general use — Anthropic's most capable generally available model, sitting above the Opus tier.

为何重要: The frontier of general-purpose reasoning just moved again; capable assistants keep getting cheaper to match.

来源

2026年6月

AN
Claude Mythos 5Anthropic· 2026年6月24日

The Mythos-class model available to approved organizations without the general-use safety measures applied to Fable 5.

为何重要: Signals how fast the top tier is advancing — the same capability reaches everyone shortly after.

来源

2026年5月

AN
Claude Opus 4.8Anthropic· 2026年5月1日

A strong all-round released model, widely cited as a top performer through mid-2026.

为何重要: A dependable default for hard reasoning, coding, and long-document work.

200K 上下文来源
OA
GPT-5.5OpenAI· 2026年5月1日

OpenAI's mid-2026 frontier update, trading the top spot with Claude Opus on many benchmarks.

为何重要: Keeps the price-for-capability race moving — good news for anyone paying per token.

来源

2026年3月

OA
GPT-5.4OpenAI· 2026年3月4日

A March 2026 frontier release with a 1M-token context window and strong computer-use scores.

为何重要: Million-token context means it can read whole manuals, contracts, or codebases at once.

1M 上下文OSWorld-Verified: 75.0%来源
DS
DeepSeek V4DeepSeek· 2026年3月3日

An open-weights frontier model with a 1M+ token context window and strong coding scores.

为何重要: Open weights + very low cost make it the value pick for budget-conscious and privacy-sensitive setups.

1M 上下文HumanEval: 94.7%来源

2026年2月

GO
Gemini 3.1 ProGoogle· 2026年2月1日

Google's February 2026 frontier update to the Gemini 3 line, with a very large context window.

为何重要: Deep integration with Google Workspace makes it a natural fit if you live in Docs and Gmail.

1M 上下文来源
OA
GPT-5.3OpenAI· 2026年2月1日

A February 2026 iteration in the GPT-5 line ahead of the March 5.4 release.

为何重要: Part of the steady cadence keeping the mainstream assistant sharp.

来源
AN
Claude Opus 4.6Anthropic· 2026年2月1日

A February 2026 Opus update (alongside Sonnet 4.6), continuing Anthropic's rapid iteration.

为何重要: Reliability gains at the same price point — worth re-testing your prompts on each bump.

200K 上下文来源

2025年11月

GO
Gemini 3 ProGoogle· 2025年11月1日

The Gemini 3 flagship that opened the current generation for Google.

为何重要: Set the bar for long-context multimodal work heading into 2026.

1M 上下文来源
AN
Claude Opus 4.5Anthropic· 2025年11月1日

A late-2025 Opus release that anchored Anthropic's top tier into 2026.

为何重要: The baseline many businesses standardized on before the 2026 wave.

200K 上下文来源

我们如何追踪

我们每周更新一次,这与维护工具价格数据的例行工作同步。我们只列出真正影响市场的主要版本发布——前沿文本模型,现在也包括视频模型——每项数据都链接到来源。没有来源,就不列出数据。如果模型已发布但尚未可通过API使用,我们会说明,因为无法购买使用的模型不算推荐。

模型每周都在变,你的方案跟得上。

你的方案是市场在某个时间点的快照,可选的按月订阅会在这些变化发生时重新核对一遍——让推荐的工具和价格始终反映当下的市场。

获取我的 AI 计划 查看方案