Skip to main content
All articles

English · 8 min read

Best AI model for coding in October 2026: GPT-6.1 Sol vs Claude Opus 5.5 vs Sonnet 5.5 (and Gemini 4 Argon, which you can't buy yet)

Claude Opus 5.5 leads the independent coding tests we track. GPT-6.1 Sol and Claude Sonnet 5.5 cost half as much at $2 / $10 per 1M tokens, and Google's Gemini 4 Argon is not on sale yet. Prices, sourced scores, which coding plan to pair with each, and what to pick for which job.

By Dapols ·

The short answer: for the hardest coding work, use Claude Opus 5.5. It has the highest coding and agentic coding scores of any closed model on LiveBench, and Artificial Analysis ranks it first of 223 models overall. For everyday coding, GPT-6.1 Sol or Claude Sonnet 5.5 costs half as much per token ($2 / $10 per 1M) and gets you most of the way there. Of the three, GPT-6.1 Sol was the cheapest per finished task on Artificial Analysis's run. If you want open weights or very cheap agentic coding, DeepSeek V4.1 Flash tops LiveBench's agentic coding table at about three cents per successful task. Gemini 4 Argon looks strong on Google's own numbers, but it is not available to buy yet.

Every number below is labelled with who measured it: the lab itself, or one of two independent benchmark sites, Artificial Analysis and LiveBench. LiveBench figures are from our 1 Oct 2026 capture. Checked on 1 October 2026. Prices are list prices per 1 million tokens on each lab's own API.

The coding models side by side

Claude Opus 5.5Claude Sonnet 5.5GPT-6.1 SolDeepSeek V4.1 FlashGemini 4 Argon
LabAnthropicAnthropicOpenAIDeepSeekGoogle
DateReleased 22 Sep 2026Released 28 Sep 2026Released 29 Sep 2026Released 10 Sep 2026Announced 30 Sep 2026
Price, input / output$4 / $20$2 / $10$2 / $10$0.15 / $0.60 off-peak, double at peak$2 / $10 introductory, then $4 / $20 (announced)
Artificial Analysis Intelligence Index (independent)58, #1 of 2235652, #11 of 22339 (23 Sep capture)53, #8 of 223
LiveBench coding (independent, 1 Oct 2026 capture)89.388.980.4——
LiveBench agentic coding (independent, 1 Oct 2026 capture)71.739.354.577.3—
LiveBench cost per successful task (independent, 1 Oct 2026 capture)about $0.80about $0.14about $0.14about $0.03—
Can you buy it today?Yes, Claude API and the big cloudsYes, Claude API, Claude apps and Claude CodeYes, API, plus Codex and ChatGPT Work on paid plansYes, API, or run the MIT-licensed weights yourselfNo. Vetted cyber defenders only

A dash means we have no figure from that source to quote. Index scores are from the Artificial Analysis pages for Opus 5.5, Sonnet 5.5, GPT-6.1 Sol and Gemini 4 Argon, read on 1 October 2026 unless a different capture date is shown. The previous GPT-6 Sol, which GPT-6.1 Sol replaces at the same price, scored 48 on that index and 81.8 on LiveBench coding.

Two things stand out. Opus 5.5 leads both coding columns among closed models. And GPT-6.1 Sol is not better than Sonnet 5.5 at everything: Sonnet 5.5 scores higher on LiveBench coding, GPT-6.1 Sol scores higher on agentic coding, and both cost about $0.14 per successful LiveBench task.

What each lab reports, and why you cannot line them up

These are the labs' own figures, not independent tests. Each lab picked different benchmarks, so a number from one list cannot be compared with a number from another.

Anthropic's own figures (Opus 5.5, Sonnet 5.5):

TestOpus 5.5Sonnet 5.5
Terminal-Bench 4.0, agentic coding in a command line66.4%70.6%
CursorBench 4.0, coding inside an editor57.8%55.5%
FrontierCode 1.1 (main, xhigh effort)54.4%52.1%

OpenAI's own figures for GPT-6.1 Sol (source). OpenAI mostly published comparisons rather than raw scores:

  • DeepSWE v1.1, long-horizon coding: matches GPT-6 Astra at roughly one-fifth of the cost, and 6.4 points above GPT-6 Sol's best at lower effort.
  • AutomationBench, business workflows: 2.2 points above Claude Opus 5.5 at medium effort, at roughly a third of the cost.
  • Terminal-Bench Science 0.1: $5.47 per task at max effort, against $23.21 for Opus 5.5 and $23.80 for GPT-6 Astra. Astra still scores highest there, at 68.1%.

Google-reported figures for Gemini 4 Argon (source): 77.9% on DeepSWE v1.1, which Google calls a new state of the art; 51.3% on AutomationBench, ranked first; and 68% on CWE-bench v1 vulnerability fixing, tied for first. Bloomberg has reported internal skepticism at Google about Argon's coding performance, which is one more reason to wait for independent scores.

Terminal-Bench 4.0 and Terminal-Bench Science 0.1 are different tests. OpenAI did not publish a DeepSWE percentage for GPT-6.1 Sol, so its result cannot be set next to Google's 77.9%. That is why the table above leans on the two independent sites, where every model runs the same tests.

Cheaper per token is not always cheaper per job

GPT-6.1 Sol and Sonnet 5.5 have the same list price, $2 / $10. On Artificial Analysis's index run, at each model's highest effort, the average cost per task came out at about $0.72 for GPT-6.1 Sol (1 Oct capture), $5.98 for Opus 5.5 and $7.60 for Sonnet 5.5 (both 29 Sep capture). Artificial Analysis labels Sonnet 5.5 "very verbose": it wrote far more output than the others, and output is the expensive part.

LiveBench tells a milder story. There, GPT-6.1 Sol and Sonnet 5.5 both cost about $0.14 per successful task, and Opus 5.5 about $0.80. The gap depends on the effort setting. Claude Code defaults to medium effort, so a real Sonnet 5.5 bill in Claude Code will not look like the max-effort run. Start at the default and raise it only when a real task fails. More on this in our Sonnet 5.5 vs Opus 5.5 vs GPT-6 Sol comparison.

The 272K rule on GPT-6.1 Sol

GPT-6.1 Sol has a 1,050,000-token context window, but a prompt over 272K input tokens bills the whole request at twice the input rate and 1.5 times the output rate. That is $4 / $15 per 1M instead of $2 / $10. Coding agents that load a large repository into one prompt can cross that line without you noticing. Anthropic bills both Claude models at the standard rate across their full 1M window.

Caching helps on the OpenAI side: cached input on GPT-6.1 Sol is $0.10 per 1M, half of GPT-6 Sol's $0.20.

Which coding plan to pair with each model

If you write code by hand with an AI assistant, a flat monthly plan is usually cheaper than paying per token. These are list prices from our catalog. Usage limits change often, so check them before you buy.

If you choosePair it withList price
Claude Opus 5.5 or Sonnet 5.5Claude Code, included with Claude Pro or MaxPro $20 a month ($200 a year billed annually); Max from $100 a month for 5x Pro's usage, with a higher 20x tier
GPT-6.1 SolCodex, included with ChatGPT Plus or Pro. OpenAI says GPT-6.1 Sol is now in Codex for Plus, Pro, Business, Enterprise and Edu usersPlus $20 a month; Pro from $100 a month (tiers at $100, $200 and $500, monthly only)
More than one lab's models in an editorCursor Pro$20 a month ($192 a year billed annually)
Automation or your own appThe lab's API, pay per tokenSee the price row in the table above

A few notes:

  • Claude Code runs in your terminal and IDE extensions. Max mainly buys more usage, so start on Pro and move up only if you hit the limits. Our Cursor vs Claude Code guide covers which one fits how your team works.
  • Codex is where OpenAI put GPT-6.1 Sol first. It is not yet in regular ChatGPT chat. OpenAI says a faster "GPT-6.1 Sol Ultrafast", with up to 8x faster token generation in Codex, is coming "in the coming days".
  • Cursor is an editor that lets you pick models from several labs. Check its model list for the models above before you commit.
  • DeepSeek V4.1 Flash has no subscription. You pay per token, and DeepSeek's API has an Anthropic-format endpoint, so it can plug into Claude Code.

Which model for which job

JobOur pickWhy
Large refactors, hard bugs, long agent runsClaude Opus 5.5Highest LiveBench coding and agentic coding scores of any closed model, and #1 on Artificial Analysis
Everyday coding where cost per finished job matters mostGPT-6.1 SolCheapest of the three per task on Artificial Analysis, $2 / $10 list price
Everyday coding on a Claude planClaude Sonnet 5.588.9 on LiveBench coding at half Opus 5.5's per-token price; keep effort at the default
Very cheap agent loops, or code that must stay on your own serversDeepSeek V4.1 Flash#1 on LiveBench agentic coding at about $0.03 per successful task, MIT-licensed weights
Anything that needs Gemini 4 ArgonWaitNot on sale; re-test when it reaches the Gemini API or Google AI Ultra

One caution on DeepSeek V4.1 Flash: it tops LiveBench's agentic coding table but scores 39 on Artificial Analysis's broader index (23 Sep capture). Benchmarks disagree, so test it on your own code before moving a team to it.

Gemini 4 Argon: what we know

Google announced Gemini 4 Argon on 30 September 2026 as its new frontier model for long-horizon software engineering, legal and finance work, and cyber defence, with a 1M-token output limit, up from 64K. At launch it is only rolling out to vetted cyber defenders through Google's Fairwind Program, while Google goes through the US government's voluntary pre-release review. Google says paid API customers and Google AI Ultra subscribers come first when it opens up, "as soon as possible", with no date. The announced price is $2 / $10 per 1M at first, then $4 / $20. Until you can actually buy it, it should not change your plans.

Our model tracker keeps the current pick for coding and every other category, with sources. If you want a coding setup chosen for your own team and budget, that is what a Business AI Plan does. The free 2-minute AI plan finder is the quick first step.

Frequently asked questions

What is the best AI model for coding in October 2026? Claude Opus 5.5, if you want the strongest results: it has the highest coding (89.3) and agentic coding (71.7) scores of any closed model on LiveBench (1 Oct 2026 capture) and ranks first on Artificial Analysis. For everyday work at half the per-token price, use GPT-6.1 Sol or Claude Sonnet 5.5.

Is GPT-6.1 Sol better than Claude Opus 5.5 for coding? Not on the independent tests. Opus 5.5 scores higher on LiveBench coding and agentic coding and on Artificial Analysis (58 against 52). GPT-6.1 Sol costs half as much per token and much less per task. OpenAI's own figures put it 2.2 points above Opus 5.5 on AutomationBench, which is a business-workflow test, not a coding one.

Can I use Gemini 4 Argon yet? No. Google announced it on 30 September 2026, but it is only available to vetted cyber defenders. Google says paid API customers and Google AI Ultra subscribers come first, with no date.

How much does GPT-6.1 Sol cost? $2 per 1M input tokens and $10 per 1M output tokens, with cached input at $0.10. A prompt over 272K input tokens bills the whole request at $4 / $15.

Which coding plan should a solo developer buy? Start with a $20 plan that matches the model you prefer: Claude Pro for Claude Code, ChatGPT Plus for Codex with GPT-6.1 Sol, or Cursor Pro if you want an editor with models from several labs. Move up only when you hit the usage limits.

Has Claude Haiku 5.5 been released? No. Haiku 5.5 is not out yet. Anthropic has announced it for "the coming weeks", and Claude Haiku 4.5 is still the current Haiku.

Sources: OpenAI — Introducing GPT-6.1 Sol, OpenAI — GPT-6.1 Sol model page, Google — Gemini 4 Argon, Anthropic — Claude Opus 5.5, Anthropic — Claude Sonnet 5.5, DeepSeek pricing, Artificial Analysis — GPT-6.1 Sol, Artificial Analysis — Gemini 4 Argon, Artificial Analysis — Claude Opus 5.5, Artificial Analysis — Claude Sonnet 5.5, LiveBench. Checked 1 October 2026.