La frontiera si muove ogni settimana. Tracciamo i modelli che contano — dagli otto laboratori che muovono realmente il mercato — con numeri fontati e un 'perché è importante' in linguaggio semplice.
La nostra lettura dei benchmark pubblici, aggiornata settimanalmente. Opinione, chiaramente etichettata — clicca su qualsiasi modello per la sua fonte.
| Modello | Laboratorio | Rilasciato | Contesto | Benchmark principale |
|---|---|---|---|---|
| GLM-5.3 | Z.ai | 14 ago 2026 | — | Terminal-Bench 3.0 (Z.ai self-reported): 28.3, up from GLM-5.2's 4.6 (GPT-5.6 Sol 34.6) |
| Grok 4.6 | xAI | 12 ago 2026 | 500K | API price (per 1M in/out tokens, <200k prompt): $2 / $6 |
| DeepSeek V4 Pro (0813)pesi aperti | DeepSeek | 12 ago 2026 | 1M | API price (per 1M in/out tokens, promo): $0.435 / $0.87 |
| Seedance 2.5video | ByteDance | 31 lug 2026 | — | Single-pass clip length: 30 seconds (up from 15) |
| DeepSeek V4 Flash (0731)pesi aperti | DeepSeek | 31 lug 2026 | 1M | API price (per 1M in/out tokens): $0.14 / $0.28 |
| Claude Opus 5 | Anthropic | 24 lug 2026 | — | API price (per 1M in/out tokens): $5 / $25 |
| Kimi K3pesi aperti | Moonshot AI | 16 lug 2026 | 1M | Total parameters (sparse MoE): 2.8T |
| GPT-5.6 (Sol · Terra · Luna) | OpenAI | 9 lug 2026 | — | API price, Sol (per 1M in/out tokens): $5 / $30 |
Z.ai's 14 Aug 2026 release reuses GLM-5.2's base model — every gain comes from scaled post-training on long-horizon RL environments. Large jumps on agentic coding (Terminal-Bench 3.0 4.6 → 28.3, SWE-Marathon 19.4 → 42.5) and an unplanned cyber result (CyberGym 84.5). All figures are Z.ai's own; no independent evaluation exists yet. Weights promised two weeks after launch.
Perché è importante: It reaches Opus 4.8's coding score on roughly 2.4× fewer output tokens, so the saving is on the bill rather than the leaderboard — but check your API calls first: `thinking.type: "disabled"` is no longer supported and will fail on glm-5.3.
xAI's 12 Aug 2026 successor to Grok 4.5: same $2/$6 API price, same 500K context, a longer post-training run aimed at long-running agents and visual/interactive work. A faster variant costs double. Knowledge cut-off is 1 Feb 2026.
Perché è importante: Frontier-class intelligence at the old 4.5 price — migrate off 4.5 with a one-line model swap. It is a knowledge-work pick, not the cheapest volume model and not the strongest terminal coding agent.
The deepseek-v4-pro API id now serves DeepSeek-V4-Pro-0813 weights. Calling method and promo pricing are unchanged. DeepSeek's changelog still has no Pro GA entry — the last note (31 Jul) said an official Pro release would follow soon — so this is a weight drop, not a new product.
Perché è importante: Harder reasoning stays cheap on the same endpoint. Keep Flash ($0.14/$0.28) for volume; do not treat 0813 as a priced relaunch until DeepSeek posts a changelog.
ByteDance's new video model generates 30-second audio-and-video clips in one pass, extends them over multiple rounds for multi-minute pieces, and edits by timestamp. At launch it runs inside Jimeng AI and Doubao Pro, with API access announced as coming via BytePlus ModelArk rather than available.
Perché è importante: The clearest jump yet in one-take video, but there is no general API to wire into a workflow — treat it as something to try inside a consumer video app, not a tool to rebuild your content process around until the API actually ships.
The official V4 Flash checkpoint entered public beta on the existing deepseek-v4-flash API id, replacing the April preview. The legacy deepseek-chat and deepseek-reasoner ids were retired on 2026-07-24.
Perché è importante: Near-frontier results at roughly a tenth of typical API prices, with MIT-licensed open weights — the budget and privacy-sensitive pick just got better without changing price.
Anthropic's new everyday flagship: close to Fable 5 on most benchmarks at half the price ($5/$25 vs $10/$50 per MTok), with a low/medium/high effort toggle to trade cost against capability.
Perché è importante: Frontier-level output at half the flagship price — re-check which Claude tier your workflows actually need before renewing.
Moonshot's 2.8-trillion-parameter sparse MoE — reportedly the largest open-weight model yet, with a 1M-token context window. Weights announced for late July; until then benchmark claims are vendor-reported.
Perché è importante: Open-weight models keep closing on the paid frontier — if you pay per-token for API work, the cheap tier just got stronger again.
OpenAI's GPT-5.6 family in three tiers — Sol (frontier), Terra (balanced), Luna (fast/cheap) — general availability across ChatGPT, Codex, and the API on Jul 9 after a government-vetted limited preview in late June.
Perché è importante: Three clear price tiers make it easier to match the model to the job — most small-business tasks belong on the cheap tier, not the flagship.
xAI's first model built specifically for coding and agentic work, priced aggressively under Anthropic and OpenAI flagships with a 500K context window.
Perché è importante: Agentic coding on a budget is now a three-way price war — worth re-testing your coding stack before renewing anything.
A Mythos-class model made safe for general use — Anthropic's most capable generally available model, sitting above the Opus tier.
Perché è importante: The frontier of general-purpose reasoning just moved again; capable assistants keep getting cheaper to match.
The Mythos-class model available to approved organizations without the general-use safety measures applied to Fable 5.
Perché è importante: Signals how fast the top tier is advancing — the same capability reaches everyone shortly after.
A strong all-round released model, widely cited as a top performer through mid-2026.
Perché è importante: A dependable default for hard reasoning, coding, and long-document work.
OpenAI's mid-2026 frontier update, trading the top spot with Claude Opus on many benchmarks.
Perché è importante: Keeps the price-for-capability race moving — good news for anyone paying per token.
A March 2026 frontier release with a 1M-token context window and strong computer-use scores.
Perché è importante: Million-token context means it can read whole manuals, contracts, or codebases at once.
An open-weights frontier model with a 1M+ token context window and strong coding scores.
Perché è importante: Open weights + very low cost make it the value pick for budget-conscious and privacy-sensitive setups.
Google's February 2026 frontier update to the Gemini 3 line, with a very large context window.
Perché è importante: Deep integration with Google Workspace makes it a natural fit if you live in Docs and Gmail.
A February 2026 iteration in the GPT-5 line ahead of the March 5.4 release.
Perché è importante: Part of the steady cadence keeping the mainstream assistant sharp.
A February 2026 Opus update (alongside Sonnet 4.6), continuing Anthropic's rapid iteration.
Perché è importante: Reliability gains at the same price point — worth re-testing your prompts on each bump.
The Gemini 3 flagship that opened the current generation for Google.
Perché è importante: Set the bar for long-context multimodal work heading into 2026.
A late-2025 Opus release that anchored Anthropic's top tier into 2026.
Perché è importante: The baseline many businesses standardized on before the 2026 wave.
Aggiorniamo questa pagina settimanalmente come parte della stessa routine che mantiene aggiornati i prezzi dei nostri strumenti. Elenchiamo solo le uscite davvero importanti dai laboratori che muovono il mercato: modelli di testo all'avanguardia e ora anche modelli video. Ogni dato statistico rimanda alla sua fonte. Nessuna fonte, nessun numero. Quando un modello è annunciato ma non ancora raggiungibile tramite API, lo diciamo, perché un modello che non puoi ancora acquistare non è una raccomandazione.
Il tuo piano è una fotografia datata del mercato e l'abbonamento mensile facoltativo la rivede man mano che questi cambiamenti arrivano, così gli strumenti consigliati e i prezzi continuano a rispecchiare il mercato reale.