Score leader
GPT-5.6 Sol · Max effort
82.4 · Overall
A source-attributed view of LiveBench scores that helps you compare capability, category strengths, and run cost without losing the detail.
Score leader
GPT-5.6 Sol · Max effort
82.4 · Overall
Lowest run cost
DeepSeek V4 Pro
$0.050 · per successful task
Top open weights
Kimi K3
78.5 · Overall
Sort every category, narrow the list, and open a model for its published subtask breakdown.
| Model | Compare | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| OpenAI | 82.4 | 91.7 | 83.9 | 65.6 | 96.2 | 79.8 | 87.7 | 71.8 | $0.589 | |
| Anthropic | 80.8 | 89.7 | 86.0 | 46.9 | 96.0 | 80.5 | 90.7 | 75.8 | $1.573 | |
| Anthropic | 80.3 | 90.2 | 83.2 | 61.3 | 94.8 | 77.9 | 87.3 | 67.5 | $0.487 | |
| OpenAI | 79.8 | 90.6 | 78.2 | 68.0 | 94.9 | 79.3 | 82.9 | 64.6 | $0.497 | |
| Moonshot AI | 78.5 | 90.7 | 81.4 | 57.6 | 84.4 | 78.7 | 85.5 | 71.4 | $0.379 | |
| 77.1 | 84.0 | 76.5 | 45.4 | 91.0 | 78.5 | 85.4 | 79.1 | $0.262 | ||
| xAI | 76.3 | 87.2 | 68.6 | 59.8 | 90.8 | 73.0 | 82.8 | 71.5 | $0.128 | |
| DeepSeek | 71.6 | 82.7 | 70.0 | 42.6 | 90.7 | 74.5 | 78.1 | 62.4 | $0.050 |
Cost is the source-provided cost-per-successful-task metric for the selected LiveBench scope.
Scores are clearer beside cost and a model’s category profile. Use comparison to keep the trade-off visible.
Cost view
Higher is stronger. Lower cost sits further left. Points use the source-provided cost metric.
Compare
Pick up to three models from the leaderboard to compare their seven category scores.
Use the compare buttons in the table to add up to three models.