Dapols
PlansAI vs hiringPricingToolsBenchmarksBlog
Sign inFind my AI workflow
Dapols

The job big companies pay a forward-deployed engineer six figures to do — as a tool, for two figures.

Join the price-watch list

Verified from vendor pages. We're not sending emails yet — leave your address and you'll be first when we do.

Product

  • AI Deployment Plans
  • AI Plan Finder
  • AI skills library
  • AI tools
  • Pricing

Explore

  • AI vs hiring
  • Stack cost checker
  • ROI calculator
  • Works with your tools
  • Comparisons
  • Alternatives
  • By industry
  • By budget

Offerings

  • AI Deployment Plans
  • Bigger or more complex? Tell us.
  • Bigger or more complex?

Company

  • Blog
  • AI model tracker
  • Submit a tool
  • Contact
  • Privacy
  • Terms of Service

Trust & methodology

  • About
  • Methodology
  • How we rank AI tools
  • Forward-deployed AI
  • Affiliate disclosure
  • AI tool pricing updates
  • AI price index
  • Security & data privacy

© 2026 Dapols. All rights reserved.

support@dapols.comX

Put AI to work—one measurable workflow at a time.

All articles

English · 8 min read

GLM-5.3: the same base model, a much better coder — and a cyber result nobody planned

Z.ai shipped GLM-5.3 on 14 August 2026. Every gain comes from post-training, not a bigger model. Terminal-Bench 3.0 went 4.6 to 28.3. Here is what is real, what it costs, and the migration that will break your API calls.

Read in another language:Françaisहिन्दीPortuguêsEspañolDeutsch

The short answer: GLM-5.3 is GLM-5.2's base model with a month of extra post-training on top — and it is a large jump in agentic coding. It is live now on the GLM Coding Plan and the API. Weights are promised in two weeks. If your code sends thinking.type: "disabled", your requests will fail on GLM-5.3 — that migration is the only urgent thing in this launch.

Verified 14 August 2026 against Z.ai's launch post, "GLM-5.3: Frontier Coding with Emergent Cyber Capabilities", including its benchmark table and methodology footnotes. All numbers below are Z.ai's own self-reported figures unless stated. There is no independent evaluation of GLM-5.3 yet, so it is not on our LiveBench snapshot — a lab's own benchmark is a claim, not a third-party score.

What actually shipped

The same base model. This is the headline most coverage will miss. Z.ai states plainly that GLM-5.3 "uses the same base model as GLM-5.2 — every gain comes from post-training." No bigger model, no new pre-training run. The gains come from scaling reinforcement learning on long-horizon task environments.

That is a genuinely interesting result, and it is also why the improvements are lopsided: they are enormous where long-horizon agentic work is measured, and modest elsewhere.

A live API, model id glm-5.3, rolled out to all GLM Coding Plan subscribers.

Open weights in two weeks, once "safety evaluation and hardening are complete." A date-bounded commitment, not an indefinite gate.

The benchmark jumps that matter

Z.ai's self-reported figures, GLM-5.3 vs GLM-5.2:

BenchmarkGLM-5.3GLM-5.2
Terminal-Bench 3.028.34.6
SWE-Marathon v1.142.519.4
ExploitBench54.424.4
AutomationBench v1.0.648.226.2
DeepSWE v1.166.946.2
Toolathlon Verified73.059.9
FrontierSWE78.167.5
CyberGym84.577.2
Terminal-Bench 2.188.281.0

Terminal-Bench 3.0 going from 4.6 to 28.3 is the clearest signal: on the hardest long-horizon agentic tasks, GLM-5.2 was essentially failing, and GLM-5.3 is not.

Note the shape of it. Terminal-Bench 2.1 moved 81.0 → 88.2, a normal increment. Terminal-Bench 3.0 moved 4.6 → 28.3, a 6× jump. The harder and longer the task, the bigger the gain — which is exactly what you would expect if the improvement is long-horizon RL rather than raw capability.

Where it still loses. On Terminal-Bench 3.0, Z.ai's own table puts GPT-5.6 Sol at 34.6 and Fable 5 at 33.7, both ahead of GLM-5.3's 28.3. On its in-house Z.ai Code Bench, GLM-5.3 reaches 34.5% at Max effort against Claude Fable 5's 39.5%. Z.ai says so directly, which is worth some credit.

The token-efficiency result is the business story

Buried under the benchmark table is the number that actually affects a bill.

On Z.ai Code Bench at Max effort, GLM-5.3 reaches 34.5% using roughly 75K output tokens per task — against GLM-5.2's 23.4% at 96K. Better results, fewer tokens.

Against a closed model: at High effort GLM-5.3 hits 31.4% at around 50K output tokens, versus Claude Opus 4.8 at 29.5% with 120K. Slightly better score, roughly 2.4× fewer output tokens.

For anyone running coding agents at volume, output tokens are the bill. A model that scores the same for a fraction of the tokens is a cost story, not a leaderboard story.

The cyber result, and why it is double-edged

Z.ai added vulnerability-discovery data to the training mix expecting modest gains, and reports being surprised: GLM-5.3 began reasoning across full exploitation chains, not just spotting isolated flaws. On ExploitGym it completes 105 tasks in two hours and 130 in six, against GLM-5.2's 29 and 39.

Run against real codebases with Chinese security teams, the model surfaced 2,436 vulnerabilities across 269 open-source projects — 107 critical and 990 high. The average flaw had survived 26.6 years undiscovered; the oldest dated to 1981. Z.ai has published a Security Disclosure Ledger tracking these, with 53 disclosed and 2,383 still under embargo.

Two honest caveats. First, Z.ai's own comparisons show it well behind the closed frontier here — the page cites competitors at 181 and 247 ExploitGym tasks against GLM-5.3's 105 and 130. Z.ai's summary of its own gap is candid: capability is growing fastest exactly where it is furthest behind.

Second, a capable open-weights vulnerability-discovery model is dual-use by construction. The two-week weights delay for "safety evaluation and hardening" is presumably about precisely that.

(A labelling inconsistency worth flagging: Z.ai's prose refers to "Mythos 5" for the cyber comparisons while its own table column is headed "Fable 5". We have quoted GLM-5.3's own figures rather than guess which competitor is meant.)

The migration that will break your calls

GLM-5.3 supports three thinking effort levels — low, high, max. Disabling thinking is no longer supported.

If your application currently sends thinking.type: "disabled", change it to enabled and set reasoning_effort to low before switching the model id to glm-5.3. Otherwise the request fails. Z.ai recommends max for coding tasks.

This is the one item in the launch with a deadline attached to it.

What it costs

Z.ai did not publish per-token pricing. The GLM Coding Plan has moved to a points-based quota system, with points counted separately for input, cached input, and output tokens.

The pricing detail that matters most: calls outside peak hours consume 50% of the standard points. Peak is 14:00–18:00 UTC+8, Monday to Friday. Everything else — including all weekend — is half price.

If you run batch work, that is a scheduling decision worth making deliberately. Overnight in Europe or the Americas is off-peak in UTC+8.

Through ZCode, Z.ai also advertises a 98%+ cache hit rate on repeated context and a 1.5× limited-time quota boost until 31 August.

What to do this week

If you are on a GLM Coding Plan. GLM-5.3 has already been rolled out to you. Check your code for thinking.type: "disabled" first, then try it — the token-efficiency gain is the thing to measure, not the benchmark number.

If you self-host GLM-5.2. Wait two weeks for weights. Nothing to pull yet.

If you use a coding harness. GLM-5.3 works in ZCode, Claude Code, OpenCode and others — you are not locked into Z.ai's own client.

If you were comparing GLM against Claude or GPT. GLM-5.3 narrows the gap on agentic coding substantially and wins clearly on tokens-per-result. It does not take the top spot on Z.ai's own table. Decide on the cost-per-outcome, not the ranking.

The honest caveats

These are the lab's own numbers. Every figure here is self-reported. Z.ai documents its methodology unusually thoroughly — harness versions, temperatures, context lengths, timeouts, anti-cheat measures — which is more than most labs publish, and it still is not an independent evaluation. GLM-5.3 is on our model tracker with its figures labelled as Z.ai's own, and it will not join our benchmark page until a third party scores it.

Several benchmarks are in-house. Z.ai Code Bench is private and unauditable by design. Its contamination argument is reasonable; it is still Z.ai grading Z.ai.

"Open weights" has not happened yet. Two weeks is a stated intention. GLM-5.2 shipped weights on day one; this did not.

Data and jurisdiction. Unchanged: know where your tokens go before routing customer data through a hosted endpoint.

The move this week

Do the migration check. Measure output tokens, not scores. If you batch, move the work off-peak and take the 50%.

If you want the "which model for which job" split decided against your actual workflow and bill, a Business AI Plan ($99 one-time, optional $29/month to keep it current) does that. The free 2-minute AI plan finder is the fast first pass.

Sources: Z.ai — GLM-5.3: Frontier Coding with Emergent Cyber Capabilities, Z.ai developer docs, ZCode. Retrieved 14 August 2026.

Contents9 sections
  1. 01What actually shipped
  2. 02The benchmark jumps that matter
  3. 03The token-efficiency result is the business story
  4. 04The cyber result, and why it is double-edged
  5. 05The migration that will break your calls
  6. 06What it costs
  7. 07What to do this week
  8. 08The honest caveats
  9. 09The move this week

Join the price-watch list

Verified from vendor pages. We're not sending emails yet — leave your address and you'll be first when we do.

Get your AI plan

Your best tools, quick wins, and budget — in two minutes.

Take the quiz