Novita AI LLM, AI Cost

Z.ai GLM Coding Plan Pricing: What Each Tier Costs and When the API Is Cheaper

Z.ai GLM Coding Plan Pricing: What Each Tier Costs and When the API Is Cheaper

Z.ai’s GLM Coding Plan has three tiers — Lite at $18/month, Pro at $80/month, and Max at $168/month — each with a fixed 5-hour and weekly credit quota for GLM-5.3 and GLM-5.3-Flash. There’s no overage billing: once a quota is used up, calls through supported tools stop working until the next reset, and the plan can’t be used to call the GLM API directly outside those tools. For developers whose usage is uneven or who need direct API access, paying per token through a provider like Novita AI avoids that cap entirely.

GLM Coding Plan Pricing and Quotas

PlanMonthlyAnnual (per month)5-Hour CreditsWeekly Credits
Lite$18$12.602,00010,000
Pro$80$5612,00060,000
Max$168$117.6028,000140,000

Pricing and quotas from z.ai’s subscription page and developer docs, checked October 10, 2026. Annual billing cuts the effective monthly price by 30%; quarterly billing cuts it by 20%. Pro gives roughly 6x Lite’s usage and Max roughly 14x, which lines up with the price steps.

Both the 5-hour and weekly limits apply at the same time. The 5-hour credits are dynamically refreshed: each block of usage resets 5 hours after it was consumed, rather than on one fixed clock for the whole plan. The weekly credits are activated when you subscribe and reset on a 7-day cycle from that point. Running out of either one blocks further calls in supported tools until the next reset — there’s no automatic fallback to pay-as-you-go billing.

How a Credit Converts to Tokens

Z.ai prices usage in credits, not tokens directly. The formula for GLM-5.3:

Credits = (input_tokens x 6.9 + cached_input_tokens x 1.7 + output_tokens x 24) / 10,000

GLM-5.3-Flash uses lower multipliers (input 2.3, cached input 0.56, output 8), so the same token volume costs fewer credits on Flash. Calls made during off-peak hours (outside Monday-Friday, 14:00-18:00 Singapore time) are charged at half the standard credit rate, which is where z.ai’s claim of saving up to 92% versus metered API pricing comes from — it assumes near-max cache hit rates and off-peak timing stacked together.

Here’s what that formula means for an actual coding session. Take a typical agentic turn in a mid-sized repo: 15,000 fresh input tokens, 150,000 cached input tokens (the rest of the repo context already in cache), and 3,000 output tokens.

Credits = (15,000 x 6.9 + 150,000 x 1.7 + 3,000 x 24) / 10,000
        = (103,500 + 255,000 + 72,000) / 10,000
        = 43.05 credits per turn

On the Lite plan’s 2,000 credit 5-hour quota, that’s about 46 turns before you hit the limit; on the 10,000 weekly quota, about 232 turns per week. Run that same token volume through Novita AI’s pay-as-you-go GLM-5.3 pricing ($1.4/Mt input, $0.26/Mt cache read, $4.4/Mt output) and one turn costs roughly $0.073. At 232 turns a week, that’s about $73-74 a month metered — compared to $18 a month on Lite. If you consistently use most of your weekly quota, the subscription is clearly the cheaper path.

The gap narrows fast if usage is lower or spikier. A developer who only runs 50-60 turns a week is paying $18 for quota they’re using a quarter of, where the metered equivalent would be under $20. And nothing carries over: unused credits don’t roll into the next cycle, so a slow week doesn’t bank credits for a busy one.

Converting Any Remaining Credit Balance to Tokens

The turn-based estimate above assumes one specific token mix. Because a credit blends three token types at different weights, “tokens per credit” isn’t fixed — it depends on how much of a session is cached. Using GLM-5.3’s multipliers, one credit buys:

  • Up to 5,882 tokens if that usage were entirely cached input (10,000 / 1.7)
  • As few as 417 tokens if that usage were entirely output (10,000 / 24)

Real sessions mix input, cached input, and output, so actual token counts land somewhere between those two bounds. Applying that range to each plan’s full quota gives a reusable reference — and the same multiply-by-417 (low) or multiply-by-5,882 (high) math works on any remaining credit balance, not just a full quota:

Plan5-Hour CreditsApprox. Tokens (5-hour)Weekly CreditsApprox. Tokens (weekly)
Lite2,0000.8M – 11.8M10,0004.2M – 58.8M
Pro12,0005.0M – 70.6M60,00025.0M – 352.9M
Max28,00011.7M – 164.7M140,00058.4M – 823.5M

Token ranges calculated from GLM-5.3’s credit multipliers (input 6.9, cached input 1.7, output 24), checked against docs.z.ai October 10, 2026.

When the Quota Runs Out

Z.ai’s FAQ is explicit about this: “Users subscribed to the Coding Plan can only make calls via the plan’s quota in supported tools. API calls outside the plan are not available.” Two consequences follow from that:

  • No overdraft. When your 5-hour or weekly credits hit zero, calls through ZCode, Claude Code, Codex, OpenCode, and the plan’s other supported tools stop. The plan does not fall back to deducting your account balance, so you aren’t surprised by a bill — but you also can’t push through a deadline by paying more.
  • No API access. The Coding Plan quota only works inside the officially supported tools. It cannot be used to call the GLM API directly, so it’s not an option for custom scripts, CI pipelines, or any workflow outside those specific clients.

Both of these are structural, not something a bigger plan tier fixes — Max still has a hard weekly cap, just a higher one.

When Pay-As-You-Go on the GLM API Makes More Sense

If your usage doesn’t fit a flat weekly allowance — bursty workloads, CI jobs, custom tooling, or multiple projects sharing one budget — metered API pricing removes the cap instead of raising it. GLM-5.3 is available on Novita AI’s serverless API at $1.4/Mt input, $0.26/Mt cached input, and $4.4/Mt output, with GLM-5.3-Flash at $0.15/Mt input and $0.5/Mt output for lighter workloads. There’s no monthly cap, no tool allowlist, and no weekly reset to plan around — you pay for what you use and can call the model from any codebase or CI step that can hit an OpenAI-compatible endpoint.

This fits a different usage pattern than the Coding Plan rather than a strictly cheaper one: it’s the right choice when quota predictability matters less than being able to call the model from anywhere, or when your weekly usage swings too much for a fixed tier to make sense.

Matching Your Usage to a Subscription or the API

Which option costs less comes down to how much you use GLM-5.3 or GLM-5.3-Flash each week and whether that usage has to run outside the Coding Plan’s supported tools. Three common patterns, using the per-turn example above as a reference point:

Heavy daily coding, one developer. Running agentic coding sessions for several hours a day inside Claude Code, ZCode, or another supported tool pushes you toward the 200+ turns a week used in the earlier example — well above the roughly 50-60 turns a week where Lite’s $18/month breaks even against metered pricing. Past that point, a subscription is the cheaper choice, and the same logic holds for Pro and Max: because their per-credit price drops as the tier goes up, the usage level where the subscription pays for itself is an even smaller share of their (larger) weekly quota.

Occasional or light use. Below roughly 50-60 turns a week — a few sessions rather than daily use — metered pricing tends to cost about the same as Lite’s $18/month or less, since credits that go unused each week don’t carry over or refund. This fits someone evaluating GLM-5.3 before committing to a plan, or using it a few times a week rather than every day.

Team or multi-project usage. The Coding Plan’s quota is tied to one account and one list of supported tools — it can’t move between team members on different schedules, and it can’t be called from a CI pipeline or a custom script. A team with uneven usage across people, or any workflow that needs direct API access alongside the coding tools, gets more predictable cost control from metered billing against one shared budget than from stacking individual subscriptions or running into the “API calls outside the plan are not available” restriction covered above.

Conclusion

The GLM Coding Plan’s $18-$168/month tiers are a good deal if your usage is steady enough to use most of a weekly quota and you only need the model inside supported coding tools. Once usage gets bursty, needs direct API access, or regularly exceeds a tier’s cap, the quota stops being a discount and starts being a wall — at that point, metered pricing through a provider like Novita AI’s GLM-5.3 API scales with actual usage instead of a fixed weekly limit.

FAQ

How much does the Z.ai GLM Coding Plan cost?

Lite is $18/month, Pro is $80/month, and Max is $168/month, billed monthly. Annual billing drops the effective price by 30% (to $12.60, $56, and $117.60 per month respectively), and quarterly billing drops it by 20%.

What happens when I run out of credits on the GLM Coding Plan?

Calls through supported tools stop until the next 5-hour or weekly reset. The plan does not charge your account balance for overage, but it also doesn’t let you pay to continue — you either wait for the reset or upgrade to a higher tier.

Can I use my GLM Coding Plan quota to call the GLM API directly?

No. Z.ai’s FAQ states the Coding Plan quota only works inside officially supported tools (ZCode, Claude Code, Codex, OpenCode, and similar). It cannot be used for direct API calls, custom scripts, or CI pipelines.

Is GLM-5.3 available outside the Coding Plan?

Yes. GLM-5.3 and GLM-5.3-Flash are available as pay-as-you-go models through providers like Novita AI, billed per token with no subscription or weekly cap.

Do unused Coding Plan credits roll over?

No. Both the 5-hour and weekly credit allowances reset on their own schedule regardless of how much you used, so unused credits don’t carry into the next cycle.

Related Posts