सीमा हर हफ्ते आगे बढ़ती है। हम उन models को ट्रैक करते हैं जो वाकई मायने रखते हैं — उन आठ labs से जो बाज़ार को असल में चलाती हैं — स्रोत सहित आंकड़ों और आसान भाषा में “यह क्यों मायने रखता है” के साथ।
सार्वजनिक बेंचमार्क पर हमारी राय, साप्ताहिक अद्यतन। राय, स्पष्ट रूप से लेबल की गई — किसी भी मॉडल पर उसके स्रोत के लिए क्लिक करें।
| मॉडल | प्रयोगशाला | रिलीज़ हुआ | संदर्भ | हेडलाइन बेंचमार्क |
|---|---|---|---|---|
| GLM-5.3 | Z.ai | 14 अग॰ 2026 | — | Terminal-Bench 3.0 (Z.ai self-reported): 28.3, up from GLM-5.2's 4.6 (GPT-5.6 Sol 34.6) |
| Grok 4.6 | xAI | 12 अग॰ 2026 | 500K | API price (per 1M in/out tokens, <200k prompt): $2 / $6 |
| DeepSeek V4 Pro (0813)खुले वजन | DeepSeek | 12 अग॰ 2026 | 1M | API price (per 1M in/out tokens, promo): $0.435 / $0.87 |
| Seedance 2.5वीडियो | ByteDance | 31 जुल॰ 2026 | — | Single-pass clip length: 30 seconds (up from 15) |
| DeepSeek V4 Flash (0731)खुले वजन | DeepSeek | 31 जुल॰ 2026 | 1M | API price (per 1M in/out tokens): $0.14 / $0.28 |
| Claude Opus 5 | Anthropic | 24 जुल॰ 2026 | — | API price (per 1M in/out tokens): $5 / $25 |
| Kimi K3खुले वजन | Moonshot AI | 16 जुल॰ 2026 | 1M | Total parameters (sparse MoE): 2.8T |
| GPT-5.6 (Sol · Terra · Luna) | OpenAI | 9 जुल॰ 2026 | — | API price, Sol (per 1M in/out tokens): $5 / $30 |
Z.ai's 14 Aug 2026 release reuses GLM-5.2's base model — every gain comes from scaled post-training on long-horizon RL environments. Large jumps on agentic coding (Terminal-Bench 3.0 4.6 → 28.3, SWE-Marathon 19.4 → 42.5) and an unplanned cyber result (CyberGym 84.5). All figures are Z.ai's own; no independent evaluation exists yet. Weights promised two weeks after launch.
यह क्यों मायने रखता है: It reaches Opus 4.8's coding score on roughly 2.4× fewer output tokens, so the saving is on the bill rather than the leaderboard — but check your API calls first: `thinking.type: "disabled"` is no longer supported and will fail on glm-5.3.
xAI's 12 Aug 2026 successor to Grok 4.5: same $2/$6 API price, same 500K context, a longer post-training run aimed at long-running agents and visual/interactive work. A faster variant costs double. Knowledge cut-off is 1 Feb 2026.
यह क्यों मायने रखता है: Frontier-class intelligence at the old 4.5 price — migrate off 4.5 with a one-line model swap. It is a knowledge-work pick, not the cheapest volume model and not the strongest terminal coding agent.
The deepseek-v4-pro API id now serves DeepSeek-V4-Pro-0813 weights. Calling method and promo pricing are unchanged. DeepSeek's changelog still has no Pro GA entry — the last note (31 Jul) said an official Pro release would follow soon — so this is a weight drop, not a new product.
यह क्यों मायने रखता है: Harder reasoning stays cheap on the same endpoint. Keep Flash ($0.14/$0.28) for volume; do not treat 0813 as a priced relaunch until DeepSeek posts a changelog.
ByteDance's new video model generates 30-second audio-and-video clips in one pass, extends them over multiple rounds for multi-minute pieces, and edits by timestamp. At launch it runs inside Jimeng AI and Doubao Pro, with API access announced as coming via BytePlus ModelArk rather than available.
यह क्यों मायने रखता है: The clearest jump yet in one-take video, but there is no general API to wire into a workflow — treat it as something to try inside a consumer video app, not a tool to rebuild your content process around until the API actually ships.
The official V4 Flash checkpoint entered public beta on the existing deepseek-v4-flash API id, replacing the April preview. The legacy deepseek-chat and deepseek-reasoner ids were retired on 2026-07-24.
यह क्यों मायने रखता है: Near-frontier results at roughly a tenth of typical API prices, with MIT-licensed open weights — the budget and privacy-sensitive pick just got better without changing price.
Anthropic's new everyday flagship: close to Fable 5 on most benchmarks at half the price ($5/$25 vs $10/$50 per MTok), with a low/medium/high effort toggle to trade cost against capability.
यह क्यों मायने रखता है: Frontier-level output at half the flagship price — re-check which Claude tier your workflows actually need before renewing.
Moonshot's 2.8-trillion-parameter sparse MoE — reportedly the largest open-weight model yet, with a 1M-token context window. Weights announced for late July; until then benchmark claims are vendor-reported.
यह क्यों मायने रखता है: Open-weight models keep closing on the paid frontier — if you pay per-token for API work, the cheap tier just got stronger again.
OpenAI's GPT-5.6 family in three tiers — Sol (frontier), Terra (balanced), Luna (fast/cheap) — general availability across ChatGPT, Codex, and the API on Jul 9 after a government-vetted limited preview in late June.
यह क्यों मायने रखता है: Three clear price tiers make it easier to match the model to the job — most small-business tasks belong on the cheap tier, not the flagship.
xAI's first model built specifically for coding and agentic work, priced aggressively under Anthropic and OpenAI flagships with a 500K context window.
यह क्यों मायने रखता है: Agentic coding on a budget is now a three-way price war — worth re-testing your coding stack before renewing anything.
A Mythos-class model made safe for general use — Anthropic's most capable generally available model, sitting above the Opus tier.
यह क्यों मायने रखता है: The frontier of general-purpose reasoning just moved again; capable assistants keep getting cheaper to match.
The Mythos-class model available to approved organizations without the general-use safety measures applied to Fable 5.
यह क्यों मायने रखता है: Signals how fast the top tier is advancing — the same capability reaches everyone shortly after.
A strong all-round released model, widely cited as a top performer through mid-2026.
यह क्यों मायने रखता है: A dependable default for hard reasoning, coding, and long-document work.
OpenAI's mid-2026 frontier update, trading the top spot with Claude Opus on many benchmarks.
यह क्यों मायने रखता है: Keeps the price-for-capability race moving — good news for anyone paying per token.
A March 2026 frontier release with a 1M-token context window and strong computer-use scores.
यह क्यों मायने रखता है: Million-token context means it can read whole manuals, contracts, or codebases at once.
An open-weights frontier model with a 1M+ token context window and strong coding scores.
यह क्यों मायने रखता है: Open weights + very low cost make it the value pick for budget-conscious and privacy-sensitive setups.
Google's February 2026 frontier update to the Gemini 3 line, with a very large context window.
यह क्यों मायने रखता है: Deep integration with Google Workspace makes it a natural fit if you live in Docs and Gmail.
A February 2026 iteration in the GPT-5 line ahead of the March 5.4 release.
यह क्यों मायने रखता है: Part of the steady cadence keeping the mainstream assistant sharp.
A February 2026 Opus update (alongside Sonnet 4.6), continuing Anthropic's rapid iteration.
यह क्यों मायने रखता है: Reliability gains at the same price point — worth re-testing your prompts on each bump.
The Gemini 3 flagship that opened the current generation for Google.
यह क्यों मायने रखता है: Set the bar for long-context multimodal work heading into 2026.
A late-2025 Opus release that anchored Anthropic's top tier into 2026.
यह क्यों मायने रखता है: The baseline many businesses standardized on before the 2026 wave.
हम इसे साप्ताहिक रूप से उसी दिनचर्या के हिस्से के रूप में अपडेट करते हैं जो हमारे टूल की कीमतों को अद्यतित रखती है। हम केवल उन लैब्स की वास्तविक प्रमुख रिलीज़ सूचीबद्ध करते हैं जो बाजार को आगे बढ़ाती हैं — फ्रंटियर टेक्स्ट मॉडल, और अब वीडियो मॉडल भी — और हर आँकड़ा अपने स्रोत से जुड़ता है। कोई स्रोत नहीं, कोई संख्या नहीं। जहाँ एक मॉडल घोषित किया गया है लेकिन अभी तक API के माध्यम से उपलब्ध नहीं है, हम ऐसा कहते हैं, क्योंकि जो मॉडल आप अभी तक नहीं खरीद सकते, वह सिफारिश नहीं है।
आपका प्लान बाज़ार की एक तारीख़ वाली तस्वीर है, और वैकल्पिक मासिक सब्सक्रिप्शन इन बदलावों के आने पर उसे फिर से जाँचता है — ताकि आपके सुझाए गए टूल और कीमतें मौजूदा बाज़ार के मुताबिक बनी रहें।