Every major AI model release. Tracked, verified, explained.
The frontier moves weekly. We track the models that matter — from the eight labs that actually move the market — with sourced numbers and a plain-English 'why it matters'.
Best right now
Our read of public benchmarks, updated weekly. Opinion, clearly labeled — click any model for its source.
Recent frontier models
The timeline
September 2026
Google's 30 Sep 2026 flagship, built for long multi-step work in software engineering, legal and finance research, and cyber defence, with a 1M-token output limit (up from 64K). Google also reports 51.3% on Zapier's AutomationBench (ranked #1), 91.7% on LVBench long-video understanding and 68% on CWE-bench v1 vulnerability fixing (tied first); those are Google's own figures. At launch it is only rolling out to vetted cyber defenders through Google's Fairwind Program while it goes through the US government's voluntary pre-release review; Google says paid API customers and Google AI Ultra subscribers come first when it opens, with no date.
Why it matters: On paper it is the strongest coding model of the week at Claude Sonnet 5.5's price, but you cannot buy it yet — keep your current model, and re-test when it reaches the Gemini API or Google AI Ultra.
OpenAI's 29 Sep 2026 DevDay upgrade to GPT-6 Sol, on the API as gpt-6.1-sol with a 1,050,000-token window, 128,000-token output and an April 2026 knowledge cutoff, and in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu (not yet in regular ChatGPT chat). OpenAI reports it within 2.1 points of GPT-6 Astra on OSWorld 2.0 computer use at about one-seventh the cost per task, 2.2 points above Claude Opus 5.5 on AutomationBench at about a third of the cost, and $5.47 per Terminal-Bench Science task against $23.21 for Opus 5.5; those are OpenAI's own figures, and Astra still scores highest on the hardest science tasks (68.1%). A faster Ultrafast tier is promised "in the coming days".
Why it matters: It is now the best frontier model per dollar on this tracker: same $2/$10 as GPT-6 Sol, and an independent test puts it at about $0.72 per task against $1.60 for Muse Spark 1.3 and $7.60 for Claude Sonnet 5.5 — so anyone on GPT-6 Sol should switch the model id, and anyone paying Astra rates for coding should re-test on it first.
Anthropic's 28 Sep 2026 mid-tier model, on the Claude API as claude-sonnet-5-5 with a 1M-token context window and 128K output, at the same $2 / $10 as Sonnet 5. Anthropic reports it 30%+ faster and up to 30% cheaper per task than Sonnet 5, and within 2 points of Opus 5.5 on GDPval-AA knowledge work (1844 against 1846) but behind it on CursorBench 4.0 (55.5% against 57.8%); those are Anthropic's own figures. Anthropic says Opus 5.5 "remains clearly stronger at complex, open-ended work requiring sustained judgment". Haiku 5.5 is announced for "the coming weeks" and is not released; Haiku 4.5 is still the current Haiku.
Why it matters: It is the Claude to try first for coding and everyday agent work at half of Opus 5.5's per-token price, but an independent test found it writes far more output than Opus 5.5 and so cost more per task at max effort — compare a real bill on your own task before switching.
Anthropic's 22 Sep 2026 successor to Opus 5, on the Claude API as claude-opus-5-5 and on AWS, Google Cloud and Azure the same day, with a 1M-token context window and 128K output. Anthropic also reports 57.8% on CursorBench 4.0 and 1846 Elo on GDPval-AA v2.1 knowledge work (Fable 5.1: 51.8% and 1735); those are Anthropic's own figures. Anthropic now recommends it as the starting model for most workloads and lists Opus 5, Fable 5 and Opus 4.8 as legacy models, still available. Claude Sonnet 5.5 followed on 28 Sep 2026; Haiku 5.5 is announced for the coming weeks.
Why it matters: It costs less than Opus 5 per token and less than half of Fable 5.1, yet an independent index now ranks it first overall — so anyone paying for Opus 5 or Fable 5.1 should re-test on it before the next bill, and a switch is a one-line model-id change.
OpenAI's 22 Sep 2026 mid-tier model, trained like GPT-6 Astra and replacing GPT-5.6 Sol at half its price. It is on the API as gpt-6-sol with a 1,050,000-token window and 128,000-token output, and in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users (not yet in regular ChatGPT chat). OpenAI also reports 68.8% on DeepSWE v1.1 at max effort, within 1.1 points of Claude Fable 5's best; those are OpenAI's own figures. Astra stays OpenAI's flagship.
Why it matters: It is the cheaper everyday OpenAI model for agent and coding work, but keep prompts under 272K tokens: above that the whole request is billed at double the input rate and 1.5x the output rate.
OpenAI's 22 Sep 2026 low-cost model for focused, high-volume work, released alongside GPT-6 Sol and replacing GPT-5.6 Luna at half its price. It is on the API as gpt-6-luna with a 1,050,000-token window and 128,000-token output, in ChatGPT Work and Codex for paid plans, and free and Go users can use it in the ChatGPT desktop app.
Why it matters: At $0.10/$0.50 it is one of the lowest list prices from a major lab — below DeepSeek V4.1 Flash's off-peak $0.15/$0.60 — but a prompt over 272K tokens is billed at $0.20/$0.75, and there V4.1 Flash off-peak is cheaper again.
xAI's 21 Sep 2026 successor to Grok 4.6 at the same $2/$6 price and the same 500K context, built on a new larger base model with a longer RL run aimed at multi-hour tasks. Reasoning effort is selectable (low, medium, high, xhigh). It is on the xAI API (grok-4.7), OpenRouter and Cursor; the faster grok-4-7-fast variant is only inside Cursor and Grok Build.
Why it matters: Moving from 4.6 is a one-line model swap at the same rate card, but it is still behind Claude Fable 5.1 and GPT-6 Astra on independent coding-agent rankings, and it is very verbose: Artificial Analysis measured about 81k output tokens per task against 27k for GPT-6 Astra, so the cost per finished job is higher than the price suggests.
DeepSeek's 10 Sep 2026 replacement for both V4 Flash and V4 Pro: a 552B-parameter MoE on a new causal encoder–decoder design that activates only 8B parameters on input and 16B on output, with native image understanding and a KV cache about a quarter the size of V4 Flash. It is served on the new deepseek-flash API id; deepseek-v4-flash and the Vision Exp id now route to it, and from 04:00 UTC on 14 Sep 2026 every deepseek-v4-pro call is routed to V4.1 Flash at Flash rates until a V4.1 Pro ships (no date). Weights are MIT-licensed on Hugging Face. DeepSeek's own table puts it ahead of V4 Pro on every agentic benchmark but behind it on pure knowledge (GPQA Diamond 90.9 vs 92.4); Artificial Analysis has since scored it independently.
Why it matters: The volume tier just got cheaper and smarter at once — $0.15/$0.60 off-peak with vision included — but the old model ids are gone, so anything still configured with deepseek-v4-flash, deepseek-v4-flash-vision-exp or deepseek-v4-pro should be switched to deepseek-flash before 14 Sep rather than silently rerouted.
OpenAI's 3 Sep 2026 flagship, built for long multi-step work it finishes end to end: reasoning, coding, and driving a computer or browser. Text and images in, text out, with a 922,000-token input window, a 128,000-token output limit and an April 2026 knowledge cutoff. It shipped first to enterprises in OpenAI's Trusted Access Program, with the API and the Plus, Pro, Business and Enterprise plans following over the days after. OpenAI reports 57.7% on Terminal-Bench 4.0 (Fable 5.1 55.8%, Opus 5 52.3%) and 41.4% on AutomationBench (Fable 5.1 31.4%); those are OpenAI's own figures. Its advanced cybersecurity capabilities are deliberately gated to a tester group rather than shipped to everyone.
Why it matters: It is the first model that is clearly better at operating a computer than at answering a question — so it earns its $10/$50 only if your bottleneck is a long chain of clicks and forms, not writing. For drafting and answering, an independent index ranks it 6th, behind Claude Opus 5.5 at less than half the price.
Google's 2 Sep 2026 Flash release, its third in six weeks, announced alongside a security-tuned Gemini 3.8 Flash Cyber variant. Text, image, video, audio and PDF in, text out, with a 1,048,576-token input window and 65,536-token output limit. Google reports it completes more than three times as many tasks as Gemini 3.7 Flash on long-horizon document work; that figure is Google's own.
Why it matters: The headline price is introductory and Google has already published the date it doubles — 1 Jan 2027 — so budget against $1.50/$7.50 rather than the launch rate if you are choosing a model to build on.
Meta's 2 Sep 2026 flagship, available the same day in Muse Code and the Meta Model API. Text, images, video, PDFs and audio in, text out, with a 1M-token window. Meta reports 88.8 on Terminal-Bench 2.1 (tying GPT-5.6 Sol, ahead of Opus 5's 86.7) and 59.4 on SWEAtlas CodeBase QnA; those scores are published as a scorecard image on Meta's own post rather than as text, and are self-reported. Artificial Analysis has since scored it independently at 48 on its current Intelligence Index scale. Open weights are promised on the roadmap with no date.
Why it matters: The contributor tier is roughly 12x cheaper on input and 21x on output, but the payment is your prompts and outputs becoming Meta training data — fine for drafting a menu, wrong for anything carrying client details.
Anthropic's 1 Sep 2026 flagship for coding and knowledge work, available the same day on the Claude API, AWS, Google Cloud and Azure. Anthropic also announced Claude Mythos 5.1 — the same underlying model with lighter cyber safeguards — but it is gated behind Cyber and Life Sciences verification programmes and is currently limited to a set of US organisations, so most buyers cannot get it. All benchmark figures are Anthropic's own; no independent evaluation exists yet.
Why it matters: The per-token rate card did not move — $10/$50 is exactly what Fable 5 cost — so the real saving is the 75% cut to cached input, at $0.25 per 1M: worth having only if your workload re-sends the same context repeatedly, and worth nothing if it doesn't.
August 2026
Alibaba's 26 Aug 2026 open-weight multimodal MoE, released as an early preview of the architecture intended for Qwen4. 125B total parameters plus 51B n-gram embeddings, with only ~6B active per token. The open checkpoint supports 262,144 tokens natively and extends to 1M with YaRN; the production endpoint, served as Qwen3.8-Flash on QwenCloud, ships 1M by default. Weights are on Hugging Face under the qwen-community-1.0 licence. All benchmark figures are Qwen's own; no independent evaluation exists yet.
Why it matters: At $0.16/$0.47 per 1M it undercuts DeepSeek V4 Flash's post-repricing rates while scoring higher on Qwen's own agentic-coding and tool-use tests — currently the cheapest credible option in the volume tier if you pay per token.
Z.ai's 26 Aug 2026 open-weight MoE: 320B total parameters with just 18B active, context up to 1M tokens. Z.ai reports large gains over GLM-5.2 across six coding and agentic benchmarks and says it nearly matches Claude Opus 4.8 on their in-house Z.ai Code Bench v1.0 at max effort (29.0 vs 29.5). It circulated anonymously on OpenRouter as "Ox Alpha" for about a week before Z.ai claimed authorship. All figures are Z.ai's own.
Why it matters: Z.ai puts it on the cost/intelligence frontier at $0.045 per task, roughly a tenth of GLM-5.2 — so if the claim survives independent testing it is the cheapest route to near-frontier coding help. Watch-and-verify: the numbers are vendor-reported and no hosted rate card shipped with the announcement.
Z.ai's 14 Aug 2026 release reuses GLM-5.2's base model — every gain comes from scaled post-training on long-horizon RL environments. Large jumps on agentic coding (Terminal-Bench 3.0 4.6 → 28.3, SWE-Marathon 19.4 → 42.5) and an unplanned cyber result (CyberGym 84.5). All figures are Z.ai's own; no independent evaluation exists yet. Weights promised two weeks after launch.
Why it matters: It reaches Opus 4.8's coding score on roughly 2.4× fewer output tokens, so the saving is on the bill rather than the leaderboard — but check your API calls first: `thinking.type: "disabled"` is no longer supported and will fail on glm-5.3.
xAI's 12 Aug 2026 successor to Grok 4.5: same $2/$6 API price, same 500K context, a longer post-training run aimed at long-running agents and visual/interactive work. A faster variant costs double. Knowledge cut-off is 1 Feb 2026.
Why it matters: Frontier-class intelligence at the old 4.5 price — migrate off 4.5 with a one-line model swap. It is a knowledge-work pick, not the cheapest volume model and not the strongest terminal coding agent.
The deepseek-v4-pro API id serves DeepSeek-V4-Pro-0813 weights. DeepSeek marked V4 Pro generally available on 13 Aug 2026, citing Terminal Bench 2.1 87.9 and NL2Repo 61.5 and adding native Responses API support for Codex. The same announcement replaced flat promo pricing with peak/off-peak billing from 16:00 UTC on 16 Aug 2026: peak hours are 01:00–04:00 and 06:00–10:00 UTC Mon–Fri, and off-peak is half of peak. On 10 Sep 2026 DeepSeek announced that all deepseek-v4-pro traffic would route to V4.1 Flash from 04:00 UTC on 14 Sep; it then reversed that and confirmed V4 Pro API service continues with billing unchanged.
Why it matters: NOT retired after all. DeepSeek said V4 Pro would be routed to V4.1 Flash from 14 Sep 2026, then reversed it and is still serving it at unchanged prices — re-verified on the live pricing page on 18 Sep, four days after the announced cutover. It stays a usable step-up, but a narrow one: V4.1 Flash beats it on DeepSeek's own agentic benchmarks at about a quarter of the price, so reach for Pro only when something measurably needs it.
Qwen's 2 Aug 2026 flagship and the first Qwen-Max-class model with open weights. Built on the Qwen3.5 architecture and scaled to 2.4 trillion parameters (95B active), served on QwenCloud through OpenAI-, Anthropic- and DashScope-compatible endpoints with low/medium/xhigh reasoning levels. The launch post promised weights "next week"; a checkpoint is published as Qwen/Qwen3.8-2.4T-A95B. Qwen published no per-token price with the announcement.
Why it matters: The first frontier-scale model you can both rent and self-host, which matters if you want a top-tier model without sending data to a US vendor — but it is a heavyweight, so for most buyers Qwen3.8-Flash-Next is the practical pick.
17 older releases, more than 60 days before the latest
July 2026
ByteDance's new video model generates 30-second audio-and-video clips in one pass, extends them over multiple rounds for multi-minute pieces, and edits by timestamp. At launch it runs inside Jimeng AI and Doubao Pro, with API access announced as coming via BytePlus ModelArk rather than available.
Why it matters: The clearest jump yet in one-take video, but there is no general API to wire into a workflow — treat it as something to try inside a consumer video app, not a tool to rebuild your content process around until the API actually ships.
Anthropic's new everyday flagship: close to Fable 5 on most benchmarks at half the price ($5/$25 vs $10/$50 per MTok), with a low/medium/high effort toggle to trade cost against capability. Since 22 Sep 2026 Anthropic lists it as a legacy model — still available, succeeded by Opus 5.5 at $4/$20.
Why it matters: Frontier-level output at half the flagship price — re-check which Claude tier your workflows actually need before renewing.
Moonshot's 2.8-trillion-parameter sparse MoE — reportedly the largest open-weight model yet, with a 1M-token context window. Weights announced for late July; until then benchmark claims are vendor-reported.
Why it matters: Open-weight models keep closing on the paid frontier — if you pay per-token for API work, the cheap tier just got stronger again.
OpenAI's GPT-5.6 family in three tiers — Sol (frontier), Terra (balanced), Luna (fast/cheap) — general availability across ChatGPT, Codex, and the API on Jul 9 after a government-vetted limited preview in late June. Not the same models as GPT-6 Sol and GPT-6 Luna (22 Sep 2026), which succeed GPT-5.6 Sol and Luna at half their prices; GPT-5.6 Terra has no GPT-6 counterpart yet.
Why it matters: Anything still calling gpt-5.6-sol or gpt-5.6-luna now pays double what gpt-6-sol and gpt-6-luna cost — worth switching the model id rather than waiting for the Sol promotion to end.
xAI's first model built specifically for coding and agentic work, priced aggressively under Anthropic and OpenAI flagships with a 500K context window.
Why it matters: Agentic coding on a budget is now a three-way price war — worth re-testing your coding stack before renewing anything.
A Mythos-class model made safe for general use, sitting above the Opus tier. Succeeded by Fable 5.1 on 1 Sep 2026; Anthropic now lists it as a legacy model, still available.
Why it matters: The frontier of general-purpose reasoning just moved again; capable assistants keep getting cheaper to match.
Traffic now goes to DeepSeek V4.1 Flash DeepSeek V4.1 Flash
The official V4 Flash checkpoint entered public beta on the existing deepseek-v4-flash API id, replacing the April preview. The legacy deepseek-chat and deepseek-reasoner ids were retired on 2026-07-24.
Why it matters: Retired on 10 Sep 2026: the deepseek-v4-flash id is still accepted but is now served by V4.1 Flash and billed at the V4.1 Flash price, which is lower. Update any configuration to the new deepseek-flash id; the prices above are this checkpoint's historical rates.
June 2026
The Mythos-class model available to approved organizations without the general-use safety measures applied to Fable 5.
Why it matters: Signals how fast the top tier is advancing — the same capability reaches everyone shortly after.
May 2026
A strong all-round released model, widely cited as a top performer through mid-2026. Anthropic now lists it as a legacy model, still available.
Why it matters: A dependable default for hard reasoning, coding, and long-document work.
OpenAI's mid-2026 frontier update, trading the top spot with Claude Opus on many benchmarks.
Why it matters: Keeps the price-for-capability race moving — good news for anyone paying per token.
March 2026
A March 2026 frontier release with a 1M-token context window and strong computer-use scores.
Why it matters: Million-token context means it can read whole manuals, contracts, or codebases at once.
Traffic now goes to DeepSeek V4.1 Flash DeepSeek V4.1 Flash
The open-weights V4 preview with a 1M-token context window, served on the deepseek-v4-flash and deepseek-v4-pro API ids until the official checkpoints replaced it.
Why it matters: Historical: open weights plus very low cost made it the value pick for budget and privacy-sensitive setups. Its successor on the API is V4.1 Flash.
February 2026
Google's February 2026 frontier update to the Gemini 3 line, with a very large context window.
Why it matters: Deep integration with Google Workspace makes it a natural fit if you live in Docs and Gmail.
A February 2026 iteration in the GPT-5 line ahead of the March 5.4 release.
Why it matters: Part of the steady cadence keeping the mainstream assistant sharp.
A February 2026 Opus update (alongside Sonnet 4.6), continuing Anthropic's rapid iteration.
Why it matters: Reliability gains at the same price point — worth re-testing your prompts on each bump.
November 2025
The Gemini 3 flagship that opened the current generation for Google.
Why it matters: Set the bar for long-context multimodal work heading into 2026.
A late-2025 Opus release that anchored Anthropic's top tier into 2026.
Why it matters: The baseline many businesses standardized on before the 2026 wave.
How we track this
We update this weekly as part of the same routine that keeps our tool prices current. We only list genuinely major releases from the labs that move the market — frontier text models, and now video models too — and every stat links to its source. No source, no number. Where a model is announced but not yet reachable through an API, we say so, because a model you cannot buy yet is not a recommendation.
Models change weekly. Your plan keeps up.
Your plan is a dated snapshot of the market, and the optional monthly subscription re-checks it as these shifts land — so your recommended tools and prices keep describing the market that exists.