Glm Coding Plan · Research

GLM Coding Plan Pricing 2026: Z.ai Lite vs Pro vs Max Explained

Z.ai GLM Coding Plan Lite, Pro, and Max tiers compared for 2026 agent coding tools
AK

Threat intelligence editor · Updated Jul 16, 2026, 2:38 AM EDT

Conflating the subscription with the metered API produces Insufficient Balance errors, surprise charges, or a hard stop with no fallback mid-task.

Z.ai sells two products under one brand, and mixing them up is the fastest way to waste a 2026 engineering budget. Z.ai is one of several Chinese labs reinventing coding-agent token pricing, and the two products get confused across all of them. The GLM Coding Plan is a flat monthly subscription built for agentic coding tools. The separate pay-as-you-go API meters tokens for any application. They use different keys, different base URLs, and different rules when you hit a limit. Developers wiring Claude Code, ZCode, Cline, OpenCode, Roo Code, Cursor, and similar clients need the Coding Plan. Cursor prices its own agent access on separate Ultra vs Pro credit tiers. Production scripts, custom agents, and unsupported tools need the metered API. Open-model routes such as Ollama Cloud, OpenRouter and OpenCode Go price similar workloads by the token instead of a flat plan. Conflating the two produces “Insufficient Balance” errors, unexpected charges, or a hard stop with no fallback.

This guide maps live GLM Coding Plan pricing, credit quotas, model multipliers, hard-stop behaviour, supported tooling, and which tier—or pure API spend—fits real usage profiles. Figures below were checked against Z.ai's own documentation on 14 September 2026.

What changed since mid-2026: on 30 July 2026 Z.ai replaced the plan's prompt-based quotas with a credits-based system metered on tokens, and reset the tier prices at the same time. On 14 August 2026 GLM-5.3 shipped and became the flagship on every tier, followed by GLM-5.3-Flash on 26 August 2026. If you are working from a plan guide written before August, essentially every quota number in it is denominated in a unit Z.ai no longer uses.

Two products, one brand

ProductBilling modelWhat it unlocksWhere it works
GLM Coding PlanFixed monthly subscriptionCredit pools + exclusive MCP stackOfficially supported coding clients only
Pay-as-you-go APIPer 1M tokensOpenAI-style model accessAny app, agent, or production workload

Coding Plan calls never drain platform cash or prepaid API balance once the plan quota is gone. General API keys never unlock Coding Plan MCP exclusives or the coding-specific endpoints. Team and Individual Coding Plan keys are not interchangeable with platform API keys. Domestic bigmodel.cn accounts and international z.ai accounts often sit on different domains and credentials; do not assume one key works on both.

GLM Coding Plan pricing and quotas: Lite vs Pro vs Max

List monthly prices and official credit allowances, as published on the subscribe page and in the DevPack overview:

TierList monthlyYearly billing (30% off, effective /mo)5-hour creditsWeekly creditsRelative vs Lite
Lite$18$12.60 ($151.20/yr)2,00010,0001×
Pro$80$56 ($672/yr)12,00060,0006×
Max$168$117.60 ($1,411.20/yr)28,000140,00014×

Billing-cycle discounts are 20% quarterly and 30% yearly; monthly billing pays list. That puts quarterly at roughly $14.40 / $64 / $134.40 per month. Live figures on the subscribe page can change; verify before purchase at https://z.ai/subscribe.

Two things worth flagging for anyone pricing this against an older write-up. First, Pro and Max are more expensive than they were at the start of the summer—$80 and $168, not $72 and $160. Second, the tier spread is no longer the one the old prompt tables implied: on credits, Pro is 6× Lite and Max is 14× Lite, so Max costs 9.3× Lite's sticker for 14× the weekly allowance, while Pro costs 4.4× Lite for 6×.

How credits are counted

Credits are deducted from token usage, not from turns. The published formula is:

credits = (input tokens × input multiplier + cached input tokens × cached multiplier + output tokens × output multiplier) ÷ 10,000

Model or toolInputCached inputOutput
GLM-5.36.91.724
GLM-5.3-Flash2.30.568
Web Search / Web Reader / Zread MCP——1.2 flat per call

Practical consequences:

  • Output dominates. GLM-5.3 output is charged at roughly 3.5× its input multiplier, so verbose agent loops and large diffs burn the pool far faster than large prompts do.
  • Cached input is cheap. At 1.7 against 6.9, keeping a stable repository context across turns is the single easiest saving inside one session.
  • MCP calls no longer have their own pool. Web Search, Web Reader, Zread and Vision Understanding all draw from the same credit balance as model use, at a flat 1.2 credits per call. There is no separate monthly MCP allowance and no state where MCP goes dark while model quota remains.
  • Both windows apply at once. The 5-hour pool and the weekly pool are independent ceilings; hitting either stops work even if the other has room.

What the credits are worth

Z.ai does not publish a headline value multiple, but the multipliers make the arithmetic straightforward. At the standard rate, 2,400 credits buys 1M GLM-5.3 output tokens, which costs $4.40 on the metered API—about $0.0018 of API-equivalent spend per credit, and about $0.0037 off-peak, where the same tokens cost half the credits.

Run that against a fully consumed weekly cap (roughly 4.33 weeks per month) and the ceiling lands at approximately:

  • Lite — about $79 of GLM-5.3 API value per month at standard rates, about $159 if most work lands off-peak. Against $18, roughly 4× to 9×.
  • Pro — about $476 standard, about $952 off-peak. Against $80, roughly 6× to 12×.
  • Max — about $1,111 standard, about $2,222 off-peak. Against $168, roughly 7× to 13×.

That is a ceiling, not an expectation: it assumes you actually exhaust the weekly cap every week and that your mix is GLM-5.3 output-heavy. Route work to GLM-5.3-Flash and the credit side gets cheaper but so does the API side you are comparing against. Treat these as an upper bound on plan value, not a guarantee.

Models and peak rates

All tiers get GLM-5.3 and GLM-5.3-Flash. The older names still resolve, but not to the models they name:

Requested modelWhat actually serves it
GLM-5.2, GLM-5.1GLM-5.3
GLM-4.7GLM-5.3-Flash

The routing is silent and needs no configuration change, so an agent config written in July against glm-4.7 is today running on GLM-5.3-Flash at that model's multipliers. GLM-5.3 shipped on 14 August 2026 reusing the GLM-5.2 base model, with its gains coming from post-training; Z.ai reports roughly a 50% improvement over GLM-5.2 on its in-house code benchmark.

Peak and off-peak. Peak is Monday to Friday, 14:00–18:00 Singapore Standard Time (UTC+8), and peak is simply the standard credit rate. Everything outside that window—including all day on both weekend days—is charged at 50% of the standard rate. This is the permanent structure of the credits system, not a promotion with an expiry date. US and EU working hours fall almost entirely in the off-peak window, which is worth an effective doubling of capacity for anyone outside Asia.

Routing. The live cost lever is model class, not generation. GLM-5.3-Flash charges an 8× output multiplier against GLM-5.3's 24×—a 3× difference on the dimension that dominates agent workloads. Route routine edits, test runs, and file-by-file refactors to GLM-5.3-Flash; reserve GLM-5.3 for long-horizon agentic tasks where the quality difference actually pays for itself.

Context and thinking. GLM-5.3 documents a 1M-token context window with 128K maximum output as its standard configuration; in Claude Code the long window is selected with the glm-5.3[1m] model tag and a matching 1M compaction window. Note one breaking change from the GLM-5.2 era: GLM-5.3 always reasons. It exposes effort levels of low, high and max, with max as the default, and a request passing a thinking type of disabled now fails rather than falling back to a non-thinking mode.

For comparison only, current metered API list prices (per 1M tokens; confirm live on docs.z.ai/guides/overview/pricing):

ModelInputCached inputOutput
GLM-5.3$1.4$0.26$4.4
GLM-5.3-Flash$0.15$0.03$0.50
GLM-5.2$1.4$0.26$4.4
GLM-4.7$0.6$0.11$2.2
GLM-4.5-Air$0.2$0.03$1.1

GLM-5-Turbo is still named on the subscribe page and still has a model guide, but no current price for it appears in the public pricing table—confirm with Z.ai before building against it. Those rates matter for break-even math; they do not unlock Coding Plan tools or MCP exclusives.

Hard-stop behaviour: no overage, no balance drain

EventWhat happens
5-hour credit pool exhaustedWait for dynamic refresh (quota resets 5 hours after consumption). No auto overage.
Weekly pool exhaustedWait for 7-day reset counted from order/subscribe time.
Coding Plan tool call after quota = 0Does not pull from account balance or API credits.
MCP callsWeb Search / Web Reader / Zread / Vision draw from the same credit pool at 1.2 credits per call; no separate allowance to run down.
Soft signalsUsage Statistics / Plan Overview progress bars; risk notices if policy flags fire.
ConcurrencyDynamic; principle Max, then Pro, then Lite; higher off-peak. Recommended concurrent projects: Lite 1 · Pro 1–2 · Max 2+.
RefundsNon-refundable once purchased.
Auto-renewOn by default. FAQ says cancel at least 24 hours before the next billing date; the usage policy has cited 3 days—confirm in the console UI.

Wrong base URL or unsupported client can still surface error 1113 Insufficient Balance or charge the account even after a plan purchase. Use the Coding Plan key from Plan Overview and the coding endpoints only.

Migrating from a legacy plan. Existing subscriptions ran to the end of their billing cycle under the old rules. Legacy Plan V1 and Team Plan holders must wait for expiry before moving to a credits tier; legacy Plan V2 holders can upgrade to a higher new tier immediately, but same-tier switches and downgrades wait for the cycle to end.

Supported tools and practical limits

Officially called out clients include ZCode (Z.ai's own harness for GLM-5.3), Claude Code, Cline, OpenCode, Roo Code, Kilo Code, OpenClaw, Crush, Goose, Cursor, and “Other Tools.” Pricing pages claim 20+ integrations. If you run GLM through OpenClaw, read OpenClaw vs Hermes Agent security compared first.

Rules that matter in production:

  • All supported tools share one subscription credit pool.
  • Must use the Coding Plan key (Team key ≠ general Z.AI API key).
  • Base URLs (confirm in current quick-start): Anthropic-compatible https://api.z.ai/api/anthropic for Claude Code and Goose; OpenAI-compatible coding path https://api.z.ai/api/coding/paas/v4 for other tools.
  • Account sharing and multi-user access are prohibited.
  • Unsupported tools risk rate limits, freezes, or bans.
  • Vision, Web Search, Web Reader, and Zread MCP features are Coding Plan package exclusives—docs state no alternate access path outside the package.

One live campaign is worth building a schedule around: during the GLM-5.3-Flash Usage Campaign, paid plan users get unlimited GLM-5.3-Flash via ZCode daily between 23:00 and 09:00, plus doubled quota on other agents in that window. Z.ai has not published an end date, so check the DevPack overview before you plan a sprint around it.

Decision guide: Lite vs Pro vs Max vs API-only

ProfileRational buyWhy
Light/occasional agent use, small repos, side projectsLite or API-only2,000 credits per 5 hours is roughly 830K GLM-5.3 output tokens at the standard rate (about 1.67M off-peak), and far less once input and tool calls are counted; rare bursts often cheaper on metered API
Full-time AI-assisted solo dev, mid repos, mostly one toolPro6× Lite; positioned as the everyday tier; still watch the 60,000 weekly cap
Multi-hour agents, multi-repo work, peak priority, heavy MCPMax14× Lite plus higher concurrency/priority claims
Production apps, custom agents, multi-tenant products, non-supported toolsAPI onlyPlan forbids general use; keys and endpoints are separate
All-day Claude Max / Codex-class volumeOften not Max aloneField reports: Pro weekly can land mid-week; Max can burn a large weekly slice in one long 5-hour window
Multi-vendor stack (Kimi / Qwen / DeepSeek + GLM)Providers / OpenRouter + APIAvoid pure Coding Plan lock-in

Cost framing for eng managers

  • A few tens of millions of GLM-5.3 tokens per month on the API can exceed Lite or Pro stickers. MiniMax prices its own Plus, Max and Ultra token plans on a similarly tiered credit ladder.
  • Coding Plan is a cost ceiling for supported IDE agents, not an unlimited Opus replacement. Xiaomi's MiMo token plan sets a comparable ceiling for its own coding agents.
  • Break-even tilts toward a tier when projected daily tool usage would exceed (monthly price ÷ 30) in equivalent API spend—or when a hard ceiling matters more than flexibility.
  • The cheapest structural lever is scheduling: the same work costs half the credits outside Monday–Friday 14:00–18:00 UTC+8, and weekends are off-peak all day.

Community feedback is mixed: GLM-5.3 quality for front-end and agent context draws praise, but heavy users often call direct Coding Plan value weaker than subsidized Claude/Codex or flexible multi-model providers, and Lite can burn quickly under aggressive agent loops. Treat any marketing framing of the plan as “tens of billions of tokens” as model-mix and multiplier-dependent, not a guarantee.

Verification before you buy

Re-check https://z.ai/subscribe and https://docs.z.ai/devpack/ for current prices, promo banners, model lists, and cancel timing, and read https://docs.z.ai/devpack/notice/usage-revision if you are migrating from a pre-August subscription. Confirm your tool is on the supported list, configure the Coding Plan key and correct base URL, and assume non-refundable once purchased. Z.ai has changed both the metering unit and the tier prices once already in 2026, so treat any figure older than a month as needing a re-check.

Bottom line: Pick the product first—Coding Plan for supported agent IDEs, metered API for everything else—then pick the tier from real credit burn under your own model mix and working hours, not from sticker envy. Lite, Pro, and Max are hard-capped credit pools with multiplier-adjusted flagship cost, MCP billed from the same pool, and zero overage. That ceiling is the feature for 2026 budgets; it only works if you never treat the plan as a general API key.

Related reading