Grok 4.6
Same $2 / $6 API price as Grok 4.5. xAI reports an Artificial Analysis Intelligence Index of 61 — tied with GPT-5.6 Sol Max, one point behind Fable 5 Max. Strong on knowledge-work evals; 26% on Terminal-Bench v3.
Read the Grok 4.6 briefingAI 벤치마크
A practical view of published model scores that helps you compare capability, category strengths, and run cost without losing the detail.
Latest benchmark & model updates
The leaderboard on this page is LiveBench's 25 June release, re-captured on 22 September 2026. It now includes Grok 4.7 xHigh and DeepSeek V4.1 Flash Max Effort. The model notes below are from the vendors' own pages, labelled as such.
Same $2 / $6 API price as Grok 4.5. xAI reports an Artificial Analysis Intelligence Index of 61 — tied with GPT-5.6 Sol Max, one point behind Fable 5 Max. Strong on knowledge-work evals; 26% on Terminal-Bench v3.
Read the Grok 4.6 briefingDeepSeek V4.1 Flash replaced V4 Flash on 10 Sep 2026 on the new deepseek-flash API id, at $0.15 / $0.60 per 1M tokens off-peak and $0.30 / $1.20 at peak. From 14 Sep, V4 Pro requests are routed to it too, so switch model ids now. LiveBench now scores V4.1 Flash Max Effort at 81.1 overall.
See it in the model trackerLeaderboard rows come only from published LiveBench data. We do not invent scores.
Best overall
Claude Fable 5.1 Max Effort leads this snapshot at 83.4 overall, $1.212 per successful task.
Best value per task
Union Alpha runs at $0.000 per successful task and still scores 76.1 overall.
Best open weights
DeepSeek V4.1 Flash Max Effort is the strongest open-weight model here at 81.1 overall — the option you can host yourself.
Best per capability
If your work is mostly one kind of task, these lead their category in this snapshot.
Every number above comes from the 2026-06-25 release dated 2026-06-25, published by a third-party benchmark and reproduced unaltered.
첫 번째 차트는 '누가 가장 강한가'에 답합니다. 두 번째 차트는 실제로 비용이 드는 질문, 즉 어떤 모델에 비용을 지불할 가치가 있는지에 답합니다.
길수록 좋습니다. 막대는 100점 만점의 전체 점수를 나타내며, 각 막대 옆의 숫자는 정확한 값입니다. 행을 선택하면 해당 모델이 아래 비용 차트에 반영됩니다.
전체 54개 중 상위 12개를 표시합니다. 전체 비교 보기로 전환하면 모든 모델을 볼 수 있습니다.
위로 갈수록 강하고, 왼쪽으로 갈수록 저렴합니다. 따라서 최고의 선택은 왼쪽 위에 위치합니다. 계단식 선은 더 저렴한 모델이 능가하지 못하는 모델을 연결하며, 선 아래의 모델은 더 저렴한 모델에 점수가 밀리는 모델입니다.
Selected
Claude Fable 5.1 Max Effort
Anthropic
#5 · 가치 프런티어에 위치
Best buys (cheapest → strongest)
플롯으로 탭 이동한 후 화살표 키를 사용하여 가장 저렴한 모델부터 비싼 모델까지 탐색하세요.
| 모델 | 조직 | 전체 | 성공 태스크당 |
|---|---|---|---|
| Claude Fable 5.1 Max Effort#5 · 가격 대비 최고 가치 | Anthropic | 83.4 | US$1.212 |
| Claude Fable 5 Max Effort | Anthropic | 83.0 | US$1.439 |
| GPT-6 Astra Max Effort#4 · 가격 대비 최고 가치 | OpenAI | 82.2 | US$0.736 |
| Muse Spark 1.3 xHigh Effort#3 · 가격 대비 최고 가치 | Meta | 81.6 | US$0.219 |
| DeepSeek V4.1 Flash Max Effort#2 · 가격 대비 최고 가치 | DeepSeek | 81.1 | US$0.029 |
| GPT-5.6 Sol Max Effort | OpenAI | 81.0 | US$0.515 |
| GPT-5.5 Thinking xHigh Effort | OpenAI | 80.2 | US$0.435 |
| Claude 5 Opus Thinking Max Effort | Anthropic | 80.1 | US$0.699 |
| Kimi K3 | Moonshot AI | 79.2 | US$0.348 |
| Gemini 3.7 Flash High | 78.8 | US$0.157 | |
| Qwen 3.8 Max | Alibaba | 78.5 | US$0.275 |
| Grok 4.6 | xAI | 78.0 | US$0.207 |
| Muse Spark 1.2 xHigh Effort | Meta | 78.0 | US$0.375 |
| GPT-5.4 Thinking xHigh Effort | OpenAI | 78.0 | US$0.387 |
| GPT-5.6 Terra Max Effort | OpenAI | 77.9 | US$0.352 |
| DeepSeek V4 Pro 0813 | DeepSeek | 77.4 | US$0.044 |
| Grok 4.7 xHigh | xAI | 77.4 | US$0.718 |
| Gemini 3.1 Pro Preview High | 77.0 | US$0.286 | |
| DeepSeek V4 Flash Vision Exp | DeepSeek | 76.8 | US$0.051 |
| Claude 4.7 Opus Thinking xHigh Effort | Anthropic | 76.5 | US$0.528 |
| Qwen 3.8 Flash Next | Alibaba | 76.2 | US$0.042 |
| Claude 4.8 Opus Thinking Max Effort | Anthropic | 76.2 | US$0.983 |
| GLM-5.3 | Z.AI | 76.1 | US$0.450 |
| Claude Sonnet 5 xHigh Effort | Anthropic | 76.0 | US$0.505 |
| Grok 4.5 | xAI | 75.8 | US$0.131 |
| Gemini 3.8 Flash High | 75.8 | US$0.307 | |
| Qwen3.8 27B | Alibaba | 75.3 | US$0.094 |
| Muse Spark 1.1 xHigh Effort | Meta | 75.3 | US$0.198 |
| GPT-5.2 High | OpenAI | 74.6 | US$0.234 |
| Gemini 3.5 Flash High | 74.6 | US$0.249 | |
| Claude 4.6 Opus Thinking High Effort | Anthropic | 74.5 | US$0.404 |
| DeepSeek V4 Flash 0731 | DeepSeek | 74.2 | US$0.060 |
| GPT-5.2 Codex | OpenAI | 74.0 | US$0.187 |
| GPT-5.6 Luna Max Effort | OpenAI | 73.6 | US$0.169 |
| Gemini 3.6 Flash High | 73.6 | US$0.235 | |
| GLM-5.2 | Z.AI | 73.2 | US$0.225 |
| Qwen 3.7 Max | Alibaba | 73.1 | US$0.182 |
| Claude 4.6 Sonnet Thinking Medium Effort | Anthropic | 73.0 | US$0.306 |
| Claude 4.5 Opus Thinking High Effort | Anthropic | 72.6 | US$0.610 |
| Inkling xHigh Effort | Thinking Machines | 71.9 | US$0.310 |
| GLM-5.3 Flash | Z.AI | 71.6 | US$0.031 |
| Kimi K2.6 Thinking | Moonshot AI | 70.5 | US$0.169 |
| GPT-5.4 Nano xHigh | OpenAI | 69.6 | US$0.091 |
| Qwen 3.6 Plus | Alibaba | 68.9 | US$0.227 |
| Kimi K2.7 Code | Moonshot AI | 68.4 | US$0.100 |
| Grok Build 0.1#1 · 가격 대비 최고 가치 | xAI | 67.8 | US$0.024 |
| Nemotron 3 Ultra 550B A55B | NVIDIA | 67.4 | US$0.371 |
| Minimax M3 | Minimax | 67.3 | US$0.060 |
| GPT-5.4 Mini xHigh | OpenAI | 66.4 | US$0.334 |
| Qwen 3.6 27B | Alibaba | 64.0 | US$0.202 |
| Gemini 3.5 Flash-Lite High | 63.9 | US$0.069 | |
| Grok 4.3 | xAI | 62.3 | US$0.061 |
리더보드에서 최대 3개의 모델을 선택하여 7개 카테고리 점수를 비교하세요.
모든 트랙은 0~100으로 표시됩니다. 오른쪽으로 갈수록 강하며, 굵은 숫자는 해당 행의 1위입니다.
| 카테고리 | Claude Fable 5.1 Max Effort | Union Alpha | DeepSeek V4.1 Flash Max Effort |
|---|---|---|---|
| 추론 | 91.7 | 80.8 | 86.7 |
| 코딩 | 86.4 | 82.1 | 80.0 |
| 에이전트 코딩 | 66.1 | 54.7 | 77.3 |
| 수학 | 97.0 | 95.3 | 93.3 |
| 데이터 분석 | 80.3 | 74.6 | 79.3 |
| 언어 | 89.5 | 85.9 | 81.2 |
| 지시 따르기 | 73.0 | 59.5 | 70.0 |
모든 카테고리 정렬, 목록 좁히기, 모델 열어서 게시된 서브태스크 분석 보기
Claude Fable 5.1 Max Effort
Anthropic
83.4
전체
US$1.212
성공 태스크당
Claude Fable 5 Max Effort
Anthropic
83.0
전체
US$1.439
성공 태스크당
GPT-6 Astra Max Effort
OpenAI
82.2
전체
US$0.736
성공 태스크당
Muse Spark 1.3 xHigh Effort
Meta
81.6
전체
US$0.219
성공 태스크당
DeepSeek V4.1 Flash Max Effort
오픈DeepSeek
81.1
전체
US$0.029
성공 태스크당
GPT-5.6 Sol Max Effort
OpenAI
81.0
전체
US$0.515
성공 태스크당
GPT-5.5 Thinking xHigh Effort
OpenAI
80.2
전체
US$0.435
성공 태스크당
Claude 5 Opus Thinking Max Effort
Anthropic
80.1
전체
US$0.699
성공 태스크당
Kimi K3
오픈Moonshot AI
79.2
전체
US$0.348
성공 태스크당
Gemini 3.7 Flash High
78.8
전체
US$0.157
성공 태스크당
Qwen 3.8 Max
오픈Alibaba
78.5
전체
US$0.275
성공 태스크당
Muse Spark 1.2 xHigh Effort
Meta
78.0
전체
US$0.375
성공 태스크당
GPT-5.4 Thinking xHigh Effort
OpenAI
78.0
전체
US$0.387
성공 태스크당
Grok 4.6
xAI
78.0
전체
US$0.207
성공 태스크당
GPT-5.6 Terra Max Effort
OpenAI
77.9
전체
US$0.352
성공 태스크당
DeepSeek V4 Pro 0813
오픈DeepSeek
77.4
전체
US$0.044
성공 태스크당
Grok 4.7 xHigh
xAI
77.4
전체
US$0.718
성공 태스크당
Gemini 3.1 Pro Preview High
77.0
전체
US$0.286
성공 태스크당
DeepSeek V4 Flash Vision Exp
오픈DeepSeek
76.8
전체
US$0.051
성공 태스크당
Claude 4.7 Opus Thinking xHigh Effort
Anthropic
76.5
전체
US$0.528
성공 태스크당
Qwen 3.8 Flash Next
오픈Alibaba
76.2
전체
US$0.042
성공 태스크당
Claude 4.8 Opus Thinking Max Effort
Anthropic
76.2
전체
US$0.983
성공 태스크당
GLM-5.3
오픈Z.AI
76.1
전체
US$0.450
성공 태스크당
Union Alpha
Stealth
76.1
전체
US$0.000
성공 태스크당
Claude Sonnet 5 xHigh Effort
Anthropic
76.0
전체
US$0.505
성공 태스크당
Gemini 3.8 Flash High
75.8
전체
US$0.307
성공 태스크당
Grok 4.5
xAI
75.8
전체
US$0.131
성공 태스크당
Qwen3.8 27B
오픈Alibaba
75.3
전체
US$0.094
성공 태스크당
Muse Spark 1.1 xHigh Effort
Meta
75.3
전체
US$0.198
성공 태스크당
GPT-5.2 High
OpenAI
74.6
전체
US$0.234
성공 태스크당
Gemini 3.5 Flash High
74.6
전체
US$0.249
성공 태스크당
Claude 4.6 Opus Thinking High Effort
Anthropic
74.5
전체
US$0.404
성공 태스크당
DeepSeek V4 Flash 0731
오픈DeepSeek
74.2
전체
US$0.060
성공 태스크당
GPT-5.2 Codex
OpenAI
74.0
전체
US$0.187
성공 태스크당
Gemini 3.6 Flash High
73.6
전체
US$0.235
성공 태스크당
GPT-5.6 Luna Max Effort
OpenAI
73.6
전체
US$0.169
성공 태스크당
GLM-5.2
오픈Z.AI
73.2
전체
US$0.225
성공 태스크당
Qwen 3.7 Max
Alibaba
73.1
전체
US$0.182
성공 태스크당
Claude 4.6 Sonnet Thinking Medium Effort
Anthropic
73.0
전체
US$0.306
성공 태스크당
Claude 4.5 Opus Thinking High Effort
Anthropic
72.6
전체
US$0.610
성공 태스크당
Inkling xHigh Effort
오픈Thinking Machines
71.9
전체
US$0.310
성공 태스크당
GLM-5.3 Flash
오픈Z.AI
71.6
전체
US$0.031
성공 태스크당
Kimi K2.6 Thinking
오픈Moonshot AI
70.5
전체
US$0.169
성공 태스크당
GPT-5.4 Nano xHigh
OpenAI
69.6
전체
US$0.091
성공 태스크당
ox-alpha-max
Stealth
69.2
전체
US$0.000
성공 태스크당
Qwen 3.6 Plus
Alibaba
68.9
전체
US$0.227
성공 태스크당
Kimi K2.7 Code
오픈Moonshot AI
68.4
전체
US$0.100
성공 태스크당
Grok Build 0.1
xAI
67.8
전체
US$0.024
성공 태스크당
Nemotron 3 Ultra 550B A55B
오픈NVIDIA
67.4
전체
US$0.371
성공 태스크당
Minimax M3
Minimax
67.3
전체
US$0.060
성공 태스크당
GPT-5.4 Mini xHigh
OpenAI
66.4
전체
US$0.334
성공 태스크당
Qwen 3.6 27B
오픈Alibaba
64.0
전체
US$0.202
성공 태스크당
Gemini 3.5 Flash-Lite High
63.9
전체
US$0.069
성공 태스크당
Grok 4.3
xAI
62.3
전체
US$0.061
성공 태스크당
Cost is the source-provided cost-per-successful-task metric for the selected scope.
These scores say which model is strongest at a set of held-out tasks. They don’t say which one fits your workflow, your budget, or the tools you already pay for — and the cheapest model that clears your bar usually beats the highest scorer.
What changed recently
Dated releases from every major lab, with the vendor's own claims labelled as claims.
OpenWhat it costs to run
Our tool catalogue carries the price we last verified and the date we checked it.
OpenWhat to actually build
A deployment plan picks the stack for one workflow at your budget — models included.
Open