Grok 4.6
Same $2 / $6 API price as Grok 4.5. xAI reports an Artificial Analysis Intelligence Index of 61 — tied with GPT-5.6 Sol Max, one point behind Fable 5 Max. Strong on knowledge-work evals; 26% on Terminal-Bench v3.
Read the Grok 4.6 briefingAI benchmarks
A practical view of published model scores that helps you compare capability, category strengths, and run cost without losing the detail.
Latest benchmark & model updates
The leaderboard on this page is LiveBench's 25 June release, re-captured on 22 September 2026. It now includes Grok 4.7 xHigh and DeepSeek V4.1 Flash Max Effort. The model notes below are from the vendors' own pages, labelled as such.
Same $2 / $6 API price as Grok 4.5. xAI reports an Artificial Analysis Intelligence Index of 61 — tied with GPT-5.6 Sol Max, one point behind Fable 5 Max. Strong on knowledge-work evals; 26% on Terminal-Bench v3.
Read the Grok 4.6 briefingDeepSeek V4.1 Flash replaced V4 Flash on 10 Sep 2026 on the new deepseek-flash API id, at $0.15 / $0.60 per 1M tokens off-peak and $0.30 / $1.20 at peak. From 14 Sep, V4 Pro requests are routed to it too, so switch model ids now. LiveBench now scores V4.1 Flash Max Effort at 81.1 overall.
See it in the model trackerLeaderboard rows come only from published LiveBench data. We do not invent scores.
Best overall
Claude Fable 5.1 Max Effort leads this snapshot at 83.4 overall, $1.212 per successful task.
Best value per task
Union Alpha runs at $0.000 per successful task and still scores 76.1 overall.
Best open weights
DeepSeek V4.1 Flash Max Effort is the strongest open-weight model here at 81.1 overall — the option you can host yourself.
Best per capability
If your work is mostly one kind of task, these lead their category in this snapshot.
Every number above comes from the 2026-06-25 release dated 2026-06-25, published by a third-party benchmark and reproduced unaltered.
The first chart answers “who is strongest”. The second answers the question that actually costs you money: which of them is worth paying for.
Longer is better. Bars show the Overall score out of 100; the figure beside each bar is the exact value. Select a row to carry that model into the cost chart below.
Showing the top 12 of 54. Switch to the full comparison view for every model.
Up is stronger, left is cheaper — so the best buys sit top-left. Numbered points are the value frontier: nothing cheaper scores higher. Grey points are beaten on both score and price.
Selected
Claude Fable 5.1 Max Effort
Anthropic
#5 · On the value frontier.
Best buys (cheapest → strongest)
Hover or tap a numbered point (or its row on the right) to read the exact score and cost. Arrow keys walk models from cheapest to dearest.
| Model | Organization | Overall | per successful task |
|---|---|---|---|
| Claude Fable 5.1 Max Effort#5 · Best value at its price | Anthropic | 83.4 | $1.212 |
| Claude Fable 5 Max Effort | Anthropic | 83.0 | $1.439 |
| GPT-6 Astra Max Effort#4 · Best value at its price | OpenAI | 82.2 | $0.736 |
| Muse Spark 1.3 xHigh Effort#3 · Best value at its price | Meta | 81.6 | $0.219 |
| DeepSeek V4.1 Flash Max Effort#2 · Best value at its price | DeepSeek | 81.1 | $0.029 |
| GPT-5.6 Sol Max Effort | OpenAI | 81.0 | $0.515 |
| GPT-5.5 Thinking xHigh Effort | OpenAI | 80.2 | $0.435 |
| Claude 5 Opus Thinking Max Effort | Anthropic | 80.1 | $0.699 |
| Kimi K3 | Moonshot AI | 79.2 | $0.348 |
| Gemini 3.7 Flash High | 78.8 | $0.157 | |
| Qwen 3.8 Max | Alibaba | 78.5 | $0.275 |
| Grok 4.6 | xAI | 78.0 | $0.207 |
| Muse Spark 1.2 xHigh Effort | Meta | 78.0 | $0.375 |
| GPT-5.4 Thinking xHigh Effort | OpenAI | 78.0 | $0.387 |
| GPT-5.6 Terra Max Effort | OpenAI | 77.9 | $0.352 |
| DeepSeek V4 Pro 0813 | DeepSeek | 77.4 | $0.044 |
| Grok 4.7 xHigh | xAI | 77.4 | $0.718 |
| Gemini 3.1 Pro Preview High | 77.0 | $0.286 | |
| DeepSeek V4 Flash Vision Exp | DeepSeek | 76.8 | $0.051 |
| Claude 4.7 Opus Thinking xHigh Effort | Anthropic | 76.5 | $0.528 |
| Qwen 3.8 Flash Next | Alibaba | 76.2 | $0.042 |
| Claude 4.8 Opus Thinking Max Effort | Anthropic | 76.2 | $0.983 |
| GLM-5.3 | Z.AI | 76.1 | $0.450 |
| Claude Sonnet 5 xHigh Effort | Anthropic | 76.0 | $0.505 |
| Grok 4.5 | xAI | 75.8 | $0.131 |
| Gemini 3.8 Flash High | 75.8 | $0.307 | |
| Qwen3.8 27B | Alibaba | 75.3 | $0.094 |
| Muse Spark 1.1 xHigh Effort | Meta | 75.3 | $0.198 |
| GPT-5.2 High | OpenAI | 74.6 | $0.234 |
| Gemini 3.5 Flash High | 74.6 | $0.249 | |
| Claude 4.6 Opus Thinking High Effort | Anthropic | 74.5 | $0.404 |
| DeepSeek V4 Flash 0731 | DeepSeek | 74.2 | $0.060 |
| GPT-5.2 Codex | OpenAI | 74.0 | $0.187 |
| GPT-5.6 Luna Max Effort | OpenAI | 73.6 | $0.169 |
| Gemini 3.6 Flash High | 73.6 | $0.235 | |
| GLM-5.2 | Z.AI | 73.2 | $0.225 |
| Qwen 3.7 Max | Alibaba | 73.1 | $0.182 |
| Claude 4.6 Sonnet Thinking Medium Effort | Anthropic | 73.0 | $0.306 |
| Claude 4.5 Opus Thinking High Effort | Anthropic | 72.6 | $0.610 |
| Inkling xHigh Effort | Thinking Machines | 71.9 | $0.310 |
| GLM-5.3 Flash | Z.AI | 71.6 | $0.031 |
| Kimi K2.6 Thinking | Moonshot AI | 70.5 | $0.169 |
| GPT-5.4 Nano xHigh | OpenAI | 69.6 | $0.091 |
| Qwen 3.6 Plus | Alibaba | 68.9 | $0.227 |
| Kimi K2.7 Code | Moonshot AI | 68.4 | $0.100 |
| Grok Build 0.1#1 · Best value at its price | xAI | 67.8 | $0.024 |
| Nemotron 3 Ultra 550B A55B | NVIDIA | 67.4 | $0.371 |
| Minimax M3 | Minimax | 67.3 | $0.060 |
| GPT-5.4 Mini xHigh | OpenAI | 66.4 | $0.334 |
| Qwen 3.6 27B | Alibaba | 64.0 | $0.202 |
| Gemini 3.5 Flash-Lite High | 63.9 | $0.069 | |
| Grok 4.3 | xAI | 62.3 | $0.061 |
Pick up to three models from the leaderboard to compare their seven category scores.
Every track runs 0 to 100. Further right is stronger; the bold figure is the leader in that row.
| Category | Claude Fable 5.1 Max Effort | Union Alpha | DeepSeek V4.1 Flash Max Effort |
|---|---|---|---|
| Reasoning | 91.7 | 80.8 | 86.7 |
| Coding | 86.4 | 82.1 | 80.0 |
| Agentic coding | 66.1 | 54.7 | 77.3 |
| Mathematics | 97.0 | 95.3 | 93.3 |
| Data analysis | 80.3 | 74.6 | 79.3 |
| Language | 89.5 | 85.9 | 81.2 |
| Instruction following | 73.0 | 59.5 | 70.0 |
Sort every category, narrow the list, and open a model for its published subtask breakdown.
Claude Fable 5.1 Max Effort
Anthropic
83.4
Overall
$1.212
per successful task
Claude Fable 5 Max Effort
Anthropic
83.0
Overall
$1.439
per successful task
GPT-6 Astra Max Effort
OpenAI
82.2
Overall
$0.736
per successful task
Muse Spark 1.3 xHigh Effort
Meta
81.6
Overall
$0.219
per successful task
DeepSeek V4.1 Flash Max Effort
openDeepSeek
81.1
Overall
$0.029
per successful task
GPT-5.6 Sol Max Effort
OpenAI
81.0
Overall
$0.515
per successful task
GPT-5.5 Thinking xHigh Effort
OpenAI
80.2
Overall
$0.435
per successful task
Claude 5 Opus Thinking Max Effort
Anthropic
80.1
Overall
$0.699
per successful task
Kimi K3
openMoonshot AI
79.2
Overall
$0.348
per successful task
Gemini 3.7 Flash High
78.8
Overall
$0.157
per successful task
Qwen 3.8 Max
openAlibaba
78.5
Overall
$0.275
per successful task
Muse Spark 1.2 xHigh Effort
Meta
78.0
Overall
$0.375
per successful task
GPT-5.4 Thinking xHigh Effort
OpenAI
78.0
Overall
$0.387
per successful task
Grok 4.6
xAI
78.0
Overall
$0.207
per successful task
GPT-5.6 Terra Max Effort
OpenAI
77.9
Overall
$0.352
per successful task
DeepSeek V4 Pro 0813
openDeepSeek
77.4
Overall
$0.044
per successful task
Grok 4.7 xHigh
xAI
77.4
Overall
$0.718
per successful task
Gemini 3.1 Pro Preview High
77.0
Overall
$0.286
per successful task
DeepSeek V4 Flash Vision Exp
openDeepSeek
76.8
Overall
$0.051
per successful task
Claude 4.7 Opus Thinking xHigh Effort
Anthropic
76.5
Overall
$0.528
per successful task
Qwen 3.8 Flash Next
openAlibaba
76.2
Overall
$0.042
per successful task
Claude 4.8 Opus Thinking Max Effort
Anthropic
76.2
Overall
$0.983
per successful task
GLM-5.3
openZ.AI
76.1
Overall
$0.450
per successful task
Union Alpha
Stealth
76.1
Overall
$0.000
per successful task
Claude Sonnet 5 xHigh Effort
Anthropic
76.0
Overall
$0.505
per successful task
Gemini 3.8 Flash High
75.8
Overall
$0.307
per successful task
Grok 4.5
xAI
75.8
Overall
$0.131
per successful task
Qwen3.8 27B
openAlibaba
75.3
Overall
$0.094
per successful task
Muse Spark 1.1 xHigh Effort
Meta
75.3
Overall
$0.198
per successful task
GPT-5.2 High
OpenAI
74.6
Overall
$0.234
per successful task
Gemini 3.5 Flash High
74.6
Overall
$0.249
per successful task
Claude 4.6 Opus Thinking High Effort
Anthropic
74.5
Overall
$0.404
per successful task
DeepSeek V4 Flash 0731
openDeepSeek
74.2
Overall
$0.060
per successful task
GPT-5.2 Codex
OpenAI
74.0
Overall
$0.187
per successful task
Gemini 3.6 Flash High
73.6
Overall
$0.235
per successful task
GPT-5.6 Luna Max Effort
OpenAI
73.6
Overall
$0.169
per successful task
GLM-5.2
openZ.AI
73.2
Overall
$0.225
per successful task
Qwen 3.7 Max
Alibaba
73.1
Overall
$0.182
per successful task
Claude 4.6 Sonnet Thinking Medium Effort
Anthropic
73.0
Overall
$0.306
per successful task
Claude 4.5 Opus Thinking High Effort
Anthropic
72.6
Overall
$0.610
per successful task
Inkling xHigh Effort
openThinking Machines
71.9
Overall
$0.310
per successful task
GLM-5.3 Flash
openZ.AI
71.6
Overall
$0.031
per successful task
Kimi K2.6 Thinking
openMoonshot AI
70.5
Overall
$0.169
per successful task
GPT-5.4 Nano xHigh
OpenAI
69.6
Overall
$0.091
per successful task
ox-alpha-max
Stealth
69.2
Overall
$0.000
per successful task
Qwen 3.6 Plus
Alibaba
68.9
Overall
$0.227
per successful task
Kimi K2.7 Code
openMoonshot AI
68.4
Overall
$0.100
per successful task
Grok Build 0.1
xAI
67.8
Overall
$0.024
per successful task
Nemotron 3 Ultra 550B A55B
openNVIDIA
67.4
Overall
$0.371
per successful task
Minimax M3
Minimax
67.3
Overall
$0.060
per successful task
GPT-5.4 Mini xHigh
OpenAI
66.4
Overall
$0.334
per successful task
Qwen 3.6 27B
openAlibaba
64.0
Overall
$0.202
per successful task
Gemini 3.5 Flash-Lite High
63.9
Overall
$0.069
per successful task
Grok 4.3
xAI
62.3
Overall
$0.061
per successful task
Cost is the source-provided cost-per-successful-task metric for the selected scope.
These scores say which model is strongest at a set of held-out tasks. They don’t say which one fits your workflow, your budget, or the tools you already pay for — and the cheapest model that clears your bar usually beats the highest scorer.
What changed recently
Dated releases from every major lab, with the vendor's own claims labelled as claims.
OpenWhat it costs to run
Our tool catalogue carries the price we last verified and the date we checked it.
OpenWhat to actually build
A deployment plan picks the stack for one workflow at your budget — models included.
Open