Aug 19, 2026
AI

GLM-5.3 API pricing is $1.40 input and $4.40 output per million tokens

Z.ai has listed GLM-5.3 on its API at the same standard token rates as GLM-5.2, with lower-priced cached input.

Renata Fuchs

By Renata Fuchs · Policy Reporter

· 3 min read

Z.ai’s GLM-5.3 API pricing is $1.40 per million input tokens and $4.40 per million output tokens, according to the company’s developer documentation. The coding-focused model is available through Z.ai’s API, giving developers a pay-as-you-go option alongside the company’s separate GLM Coding Plan subscriptions.

The posted rate is unchanged from GLM-5.2 and GLM-5.1. Z.ai also lists cached input at $0.26 per million tokens and says cached-input storage is free for a limited time. A defined request volume of 1 million non-cached input tokens plus 1 million output tokens would cost $5.80 at those list rates. That is an arithmetic illustration, not a bundled price or an estimate of a typical workload.

What does GLM-5.3 API pricing include?

Z.ai’s price table separates standard input, cached input and output. Actual charges therefore depend on the mix of those token types. Z.ai says its context-caching system can identify repeated or highly similar content from prior requests, reuse prior computation and report cached-token counts in API usage data.

The company’s documentation lists GLM-5.3 as supporting OpenAI Chat Completion, OpenAI Response and Anthropic Message protocols, each with its own endpoint. There is a limitation for subscribers: developers who currently have, or previously had, a GLM Coding Plan can access the model API only through the OpenAI Chat Completion-compatible protocol, Z.ai says.

GLM-5.3 accepts text inputs, has a 1 million-token context window and a 128,000-token maximum output length, according to Z.ai. Reasoning is mandatory, with low, high and max effort settings. The documented default is max. Applications that previously set thinking.type to disabled must switch it to enabled and set reasoning effort to low before adopting the new model ID, or requests will fail, the company says.

How does the Coding Plan differ from API billing?

The GLM Coding Plan is a subscription product with five-hour and weekly credit limits, rather than a substitute for standard API list pricing. For GLM-5.3 usage under that plan, Z.ai assigns credit multipliers of 6.9 for input, 1.7 for cached input and 24 for output. During off-peak periods, plan usage consumes 50% of the usual credit rate; Z.ai defines peak hours as 2 p.m. to 6 p.m. Singapore time on weekdays.

Z.ai introduced GLM-5.3 on August 14 and said it shares GLM-5.2’s base model, with its improvements coming from post-training. The company claims improved coding, agent and cybersecurity performance, including a 50% gain over GLM-5.2 on its private Z.ai Code Bench. Those performance claims are from Z.ai, and the company has said it plans to release the model’s weights after safety evaluation and hardening rather than indicating that they are already available.

For teams evaluating the model, the immediate commercial change is access at an existing price point, not a price cut. Z.ai’s listed rates provide a starting point, while protocol compatibility, mandatory reasoning settings and the distinction between API billing and subscription credits determine how an integration is configured.

This story draws on original reporting from VentureBeat.

More from AI

All AI →