Grok 4.6
Same $2 / $6 API price as Grok 4.5. xAI reports an Artificial Analysis Intelligence Index of 61 — tied with GPT-5.6 Sol Max, one point behind Fable 5 Max. Strong on knowledge-work evals; 26% on Terminal-Bench v3.
Read the Grok 4.6 briefingAI基准测试
A practical view of published model scores that helps you compare capability, category strengths, and run cost without losing the detail.
Latest benchmark & model updates
The leaderboard on this page is LiveBench's 25 June release, re-captured on 22 September 2026. It now includes Grok 4.7 xHigh and DeepSeek V4.1 Flash Max Effort. The model notes below are from the vendors' own pages, labelled as such.
Same $2 / $6 API price as Grok 4.5. xAI reports an Artificial Analysis Intelligence Index of 61 — tied with GPT-5.6 Sol Max, one point behind Fable 5 Max. Strong on knowledge-work evals; 26% on Terminal-Bench v3.
Read the Grok 4.6 briefingDeepSeek V4.1 Flash replaced V4 Flash on 10 Sep 2026 on the new deepseek-flash API id, at $0.15 / $0.60 per 1M tokens off-peak and $0.30 / $1.20 at peak. From 14 Sep, V4 Pro requests are routed to it too, so switch model ids now. LiveBench now scores V4.1 Flash Max Effort at 81.1 overall.
See it in the model trackerLeaderboard rows come only from published LiveBench data. We do not invent scores.
Best overall
Claude Fable 5.1 Max Effort leads this snapshot at 83.4 overall, $1.212 per successful task.
Best value per task
Union Alpha runs at $0.000 per successful task and still scores 76.1 overall.
Best open weights
DeepSeek V4.1 Flash Max Effort is the strongest open-weight model here at 81.1 overall — the option you can host yourself.
Best per capability
If your work is mostly one kind of task, these lead their category in this snapshot.
Every number above comes from the 2026-06-25 release dated 2026-06-25, published by a third-party benchmark and reproduced unaltered.
第一张图回答“谁最强”。第二张图回答真正让您花钱的问题:其中哪些值得付费。
越长越好。条形显示总体得分(满分100);每个条形旁边的数字是精确值。选择一行可将该模型带入下方的成本图表。
显示54个中的前12个。切换到完整比较视图以查看每个模型。
向上更强,向左更便宜——因此最划算的位于左上角。阶梯线连接了没有更便宜选项能击败的模型;其下方的任何模型都被成本更低的模型超越。
Selected
Claude Fable 5.1 Max Effort
Anthropic
#5 · 处于价值前沿。
Best buys (cheapest → strongest)
按Tab键进入绘图,然后使用箭头键从最便宜到最贵遍历模型。
| 模型 | 组织 | 总体 | 每次成功任务 |
|---|---|---|---|
| Claude Fable 5.1 Max Effort#5 · 在其价格下最具价值 | Anthropic | 83.4 | US$1.212 |
| Claude Fable 5 Max Effort | Anthropic | 83.0 | US$1.439 |
| GPT-6 Astra Max Effort#4 · 在其价格下最具价值 | OpenAI | 82.2 | US$0.736 |
| Muse Spark 1.3 xHigh Effort#3 · 在其价格下最具价值 | Meta | 81.6 | US$0.219 |
| DeepSeek V4.1 Flash Max Effort#2 · 在其价格下最具价值 | DeepSeek | 81.1 | US$0.029 |
| GPT-5.6 Sol Max Effort | OpenAI | 81.0 | US$0.515 |
| GPT-5.5 Thinking xHigh Effort | OpenAI | 80.2 | US$0.435 |
| Claude 5 Opus Thinking Max Effort | Anthropic | 80.1 | US$0.699 |
| Kimi K3 | Moonshot AI | 79.2 | US$0.348 |
| Gemini 3.7 Flash High | 78.8 | US$0.157 | |
| Qwen 3.8 Max | Alibaba | 78.5 | US$0.275 |
| Grok 4.6 | xAI | 78.0 | US$0.207 |
| Muse Spark 1.2 xHigh Effort | Meta | 78.0 | US$0.375 |
| GPT-5.4 Thinking xHigh Effort | OpenAI | 78.0 | US$0.387 |
| GPT-5.6 Terra Max Effort | OpenAI | 77.9 | US$0.352 |
| DeepSeek V4 Pro 0813 | DeepSeek | 77.4 | US$0.044 |
| Grok 4.7 xHigh | xAI | 77.4 | US$0.718 |
| Gemini 3.1 Pro Preview High | 77.0 | US$0.286 | |
| DeepSeek V4 Flash Vision Exp | DeepSeek | 76.8 | US$0.051 |
| Claude 4.7 Opus Thinking xHigh Effort | Anthropic | 76.5 | US$0.528 |
| Qwen 3.8 Flash Next | Alibaba | 76.2 | US$0.042 |
| Claude 4.8 Opus Thinking Max Effort | Anthropic | 76.2 | US$0.983 |
| GLM-5.3 | Z.AI | 76.1 | US$0.450 |
| Claude Sonnet 5 xHigh Effort | Anthropic | 76.0 | US$0.505 |
| Grok 4.5 | xAI | 75.8 | US$0.131 |
| Gemini 3.8 Flash High | 75.8 | US$0.307 | |
| Qwen3.8 27B | Alibaba | 75.3 | US$0.094 |
| Muse Spark 1.1 xHigh Effort | Meta | 75.3 | US$0.198 |
| GPT-5.2 High | OpenAI | 74.6 | US$0.234 |
| Gemini 3.5 Flash High | 74.6 | US$0.249 | |
| Claude 4.6 Opus Thinking High Effort | Anthropic | 74.5 | US$0.404 |
| DeepSeek V4 Flash 0731 | DeepSeek | 74.2 | US$0.060 |
| GPT-5.2 Codex | OpenAI | 74.0 | US$0.187 |
| GPT-5.6 Luna Max Effort | OpenAI | 73.6 | US$0.169 |
| Gemini 3.6 Flash High | 73.6 | US$0.235 | |
| GLM-5.2 | Z.AI | 73.2 | US$0.225 |
| Qwen 3.7 Max | Alibaba | 73.1 | US$0.182 |
| Claude 4.6 Sonnet Thinking Medium Effort | Anthropic | 73.0 | US$0.306 |
| Claude 4.5 Opus Thinking High Effort | Anthropic | 72.6 | US$0.610 |
| Inkling xHigh Effort | Thinking Machines | 71.9 | US$0.310 |
| GLM-5.3 Flash | Z.AI | 71.6 | US$0.031 |
| Kimi K2.6 Thinking | Moonshot AI | 70.5 | US$0.169 |
| GPT-5.4 Nano xHigh | OpenAI | 69.6 | US$0.091 |
| Qwen 3.6 Plus | Alibaba | 68.9 | US$0.227 |
| Kimi K2.7 Code | Moonshot AI | 68.4 | US$0.100 |
| Grok Build 0.1#1 · 在其价格下最具价值 | xAI | 67.8 | US$0.024 |
| Nemotron 3 Ultra 550B A55B | NVIDIA | 67.4 | US$0.371 |
| Minimax M3 | Minimax | 67.3 | US$0.060 |
| GPT-5.4 Mini xHigh | OpenAI | 66.4 | US$0.334 |
| Qwen 3.6 27B | Alibaba | 64.0 | US$0.202 |
| Gemini 3.5 Flash-Lite High | 63.9 | US$0.069 | |
| Grok 4.3 | xAI | 62.3 | US$0.061 |
从排行榜中选择最多三个模型来比较它们的七个类别分数。
每个赛道从0到100。越靠右越强;粗体数字是该行的领先者。
| 类别 | Claude Fable 5.1 Max Effort | Union Alpha | DeepSeek V4.1 Flash Max Effort |
|---|---|---|---|
| 推理 | 91.7 | 80.8 | 86.7 |
| 编码 | 86.4 | 82.1 | 80.0 |
| 智能体编码 | 66.1 | 54.7 | 77.3 |
| 数学 | 97.0 | 95.3 | 93.3 |
| 数据分析 | 80.3 | 74.6 | 79.3 |
| 语言 | 89.5 | 85.9 | 81.2 |
| 指令遵循 | 73.0 | 59.5 | 70.0 |
对每个类别进行排序,缩小列表,并打开模型以查看其发布的子任务细分。
Claude Fable 5.1 Max Effort
Anthropic
83.4
总体
US$1.212
每次成功任务
Claude Fable 5 Max Effort
Anthropic
83.0
总体
US$1.439
每次成功任务
GPT-6 Astra Max Effort
OpenAI
82.2
总体
US$0.736
每次成功任务
Muse Spark 1.3 xHigh Effort
Meta
81.6
总体
US$0.219
每次成功任务
DeepSeek V4.1 Flash Max Effort
开源DeepSeek
81.1
总体
US$0.029
每次成功任务
GPT-5.6 Sol Max Effort
OpenAI
81.0
总体
US$0.515
每次成功任务
GPT-5.5 Thinking xHigh Effort
OpenAI
80.2
总体
US$0.435
每次成功任务
Claude 5 Opus Thinking Max Effort
Anthropic
80.1
总体
US$0.699
每次成功任务
Kimi K3
开源Moonshot AI
79.2
总体
US$0.348
每次成功任务
Gemini 3.7 Flash High
78.8
总体
US$0.157
每次成功任务
Qwen 3.8 Max
开源Alibaba
78.5
总体
US$0.275
每次成功任务
Muse Spark 1.2 xHigh Effort
Meta
78.0
总体
US$0.375
每次成功任务
GPT-5.4 Thinking xHigh Effort
OpenAI
78.0
总体
US$0.387
每次成功任务
Grok 4.6
xAI
78.0
总体
US$0.207
每次成功任务
GPT-5.6 Terra Max Effort
OpenAI
77.9
总体
US$0.352
每次成功任务
DeepSeek V4 Pro 0813
开源DeepSeek
77.4
总体
US$0.044
每次成功任务
Grok 4.7 xHigh
xAI
77.4
总体
US$0.718
每次成功任务
Gemini 3.1 Pro Preview High
77.0
总体
US$0.286
每次成功任务
DeepSeek V4 Flash Vision Exp
开源DeepSeek
76.8
总体
US$0.051
每次成功任务
Claude 4.7 Opus Thinking xHigh Effort
Anthropic
76.5
总体
US$0.528
每次成功任务
Qwen 3.8 Flash Next
开源Alibaba
76.2
总体
US$0.042
每次成功任务
Claude 4.8 Opus Thinking Max Effort
Anthropic
76.2
总体
US$0.983
每次成功任务
GLM-5.3
开源Z.AI
76.1
总体
US$0.450
每次成功任务
Union Alpha
Stealth
76.1
总体
US$0.000
每次成功任务
Claude Sonnet 5 xHigh Effort
Anthropic
76.0
总体
US$0.505
每次成功任务
Gemini 3.8 Flash High
75.8
总体
US$0.307
每次成功任务
Grok 4.5
xAI
75.8
总体
US$0.131
每次成功任务
Qwen3.8 27B
开源Alibaba
75.3
总体
US$0.094
每次成功任务
Muse Spark 1.1 xHigh Effort
Meta
75.3
总体
US$0.198
每次成功任务
GPT-5.2 High
OpenAI
74.6
总体
US$0.234
每次成功任务
Gemini 3.5 Flash High
74.6
总体
US$0.249
每次成功任务
Claude 4.6 Opus Thinking High Effort
Anthropic
74.5
总体
US$0.404
每次成功任务
DeepSeek V4 Flash 0731
开源DeepSeek
74.2
总体
US$0.060
每次成功任务
GPT-5.2 Codex
OpenAI
74.0
总体
US$0.187
每次成功任务
Gemini 3.6 Flash High
73.6
总体
US$0.235
每次成功任务
GPT-5.6 Luna Max Effort
OpenAI
73.6
总体
US$0.169
每次成功任务
GLM-5.2
开源Z.AI
73.2
总体
US$0.225
每次成功任务
Qwen 3.7 Max
Alibaba
73.1
总体
US$0.182
每次成功任务
Claude 4.6 Sonnet Thinking Medium Effort
Anthropic
73.0
总体
US$0.306
每次成功任务
Claude 4.5 Opus Thinking High Effort
Anthropic
72.6
总体
US$0.610
每次成功任务
Inkling xHigh Effort
开源Thinking Machines
71.9
总体
US$0.310
每次成功任务
GLM-5.3 Flash
开源Z.AI
71.6
总体
US$0.031
每次成功任务
Kimi K2.6 Thinking
开源Moonshot AI
70.5
总体
US$0.169
每次成功任务
GPT-5.4 Nano xHigh
OpenAI
69.6
总体
US$0.091
每次成功任务
ox-alpha-max
Stealth
69.2
总体
US$0.000
每次成功任务
Qwen 3.6 Plus
Alibaba
68.9
总体
US$0.227
每次成功任务
Kimi K2.7 Code
开源Moonshot AI
68.4
总体
US$0.100
每次成功任务
Grok Build 0.1
xAI
67.8
总体
US$0.024
每次成功任务
Nemotron 3 Ultra 550B A55B
开源NVIDIA
67.4
总体
US$0.371
每次成功任务
Minimax M3
Minimax
67.3
总体
US$0.060
每次成功任务
GPT-5.4 Mini xHigh
OpenAI
66.4
总体
US$0.334
每次成功任务
Qwen 3.6 27B
开源Alibaba
64.0
总体
US$0.202
每次成功任务
Gemini 3.5 Flash-Lite High
63.9
总体
US$0.069
每次成功任务
Grok 4.3
xAI
62.3
总体
US$0.061
每次成功任务
Cost is the source-provided cost-per-successful-task metric for the selected scope.
These scores say which model is strongest at a set of held-out tasks. They don’t say which one fits your workflow, your budget, or the tools you already pay for — and the cheapest model that clears your bar usually beats the highest scorer.
What changed recently
Dated releases from every major lab, with the vendor's own claims labelled as claims.
OpenWhat it costs to run
Our tool catalogue carries the price we last verified and the date we checked it.
OpenWhat to actually build
A deployment plan picks the stack for one workflow at your budget — models included.
Open