GLM 5.3 Pricing in 2026: Every Route Compared
GLM 5.3 costs $1.40 per million input tokens at Z.ai and $0.81 on SayGm, or $18 a month on the Coding Plan. Here's what each route actually costs, where the subscription stops being a bargain, and which host is cheapest for 5.3, 5.2 and Flash.

On this page
For heavy daily coding in one tool, the Z AI Coding Plan at $18 a month beats every per-token route once you pass about 15 million tokens. Everything else: a per-token host, where SayGm is the cheapest standing rate on GLM 5.3 ($0.81 and $2.55) and 5.3 Flash ($0.13 and $0.45).
If the prompt has to stay private, SayGm's TEE routes for 5.1 and 5.2 are the cheapest confidential option on the market.
Table of Contents
- Three ways to pay for GLM
- What does Z.ai charge for GLM?
- Is the GLM Coding Plan cheaper than the API?
- GLM 5.3 pricing by host: the cheapest routes
- GLM 5.2 API pricing: the cheapest hosts
- GLM 5.3 Flash: the budget tier
- What 100 million GLM tokens costs
- Who hosts GLM, and what each one is good at
- Can you run GLM without anyone reading your prompts?
- Switching to SayGm for GLM
- Common questions about GLM pricing
- TL;DR
- About SayGm
GLM 5.3 pricing has become a live question this year for a simple reason: Z.ai has shipped four flagship releases since February and now sells the same model three different ways, through a subscription, its own API and a couple of dozen third-party hosts. Those three routes can differ by 3x for the same tokens.
This guide lays out what Z.ai charges, where the GLM Coding Plan stops being a bargain, what the cheapest way to use GLM 5.2 and 5.3 actually is, and what each of the six hosts that matter is good at. It follows the same method as our Kimi K2.6 and K3 pricing comparison, so the two are directly compara
Three ways to pay for GLM
Most model families give you two choices: the maker's API or a third-party host. GLM has a third, and it is the one most people search for. The GLM Coding Plan is a monthly subscription that plugs into Claude Code, Cline, Roo Code and about twenty other coding tools, with a prompt quota instead of a token meter. Z.ai's pay-as-you-go API is the second route, priced per million tokens. The third is the open-weight market: because GLM 5.x ships under an MIT licence, hosts from DeepInfra to Cloudflare serve it, usually at or below Z.ai's rate.
Which one is cheapest depends almost entirely on how you use the model. The rest of this post puts numbers on that.
What does Z.ai charge for GLM?
Z.ai publishes its rates in the developer docs. Per million tokens:
| Model | Input | Cached input | Output | Notes |
|---|---|---|---|---|
| GLM-5.3 | $1.40 | $0.26 | $4.40 | Flagship, 1M context |
| GLM-5.3-Flash | $0.15 | $0.03 | $0.50 | Small, fast tier |
| GLM-5.2 | $1.40 | $0.26 | $4.40 | Previous flagship |
| GLM-5.1 | $1.40 | $0.26 | $4.40 | April release |
| GLM-5 | $1.00 | $0.20 | $3.20 | February release |
| GLM-4.7-Flash, GLM-4.5-Flash | Free | Free | Free | Older small models |
Three things to notice. GLM 5.1, 5.2 and 5.3 all carry the same list price, so upgrading costs nothing at Z.ai's rate; the current rate arrived with 5.1 in April and has held since. The Flash tier is almost ten times cheaper than the flagship, which is why it deserves its own section below. And cached input at $0.26 is a real discount but not a dramatic one; GLM rewards prefix caching less than Kimi does, where cache hits run at a sixth of the fresh rate.
Is the GLM Coding Plan cheaper than the API?
For a heavy daily coder, usually yes. For everyone else, usually no. Here is the arithmetic.
The Coding Plan comes in three tiers: Lite at $18 a month, Pro at $72 and Max at $160, with Z.ai running 30 percent promotional discounts on and off through the year. Lite is quoted at roughly 80 prompts per five-hour window and around 400 a week; Pro and Max scale that quota by about 5x and 20x. By MLQ's account GLM 5.3 shipped into the plan on 14 August, two weeks before the weights were public, so the plan is also where new releases land first.
Now the API side. At a typical agentic mix of four input tokens per output token, GLM 5.3 through SayGm costs about $1.16 per million tokens, and through Z.ai's own API about $2.00. So $18 of credit buys roughly 15 million tokens on SayGm or 9 million direct. A coding agent that re-reads a large repository on every turn can burn 30,000 tokens a prompt, which means 400 prompts a week is in the region of 50 million tokens a month. At that usage the Lite plan is cheaper than any per-token route, by a wide margin.
The plan loses in three situations. If your usage is bursty, a quiet month still costs $18. If you are building something for other people, the plan's per-user quota and tool-only access make it the wrong shape; it is a per-user seat rather than an API. And if you want to route between GLM and other models, or fall back when one degrades, you need a per-token endpoint. The usual advice holds: light or bursty use is cheaper on the API, heavy daily interactive use is cheaper on the plan.
GLM 5.3 pricing by host: the cheapest routes

Per the release coverage on DEV, GLM-5.3 is a 744-billion-parameter mixture-of-experts model with about 40 billion active, the same base as 5.2 with the gains coming entirely from post-training. Z.ai claims a 50 percent improvement on coding benchmarks over 5.2, and the same coverage lists 28.3 on Terminal-Bench 3.0 against 4.6 for 5.2, plus a vulnerability-discovery capability that set off a safety debate of its own. OpenRouter lists 28 hosts for it, which is more than any other open-weight model this year. Prices per million tokens, from each host's page or OpenRouter's GLM 5.3 listing, as of mid September:
| Host | Input | Output | Cache read | Confidential? |
|---|---|---|---|---|
| SayGm (glm-5.3) | $0.81 | $2.55 | $0.15 | No |
| OpenRouter (cheapest route + 5.5% fee) | from $0.93 | from $3.13 | varies | No |
| Reka AI | $0.88 | $2.97 | $0.18 | No |
| DeepInfra (FP4) | $0.90 | $3.00 | $0.15 | No |
| Phala | $0.98 | $3.08 | $0.18 | Yes |
| Z.ai (direct) | $1.40 | $4.40 | $0.26 | No |
| Fireworks, Together, Cloudflare | $1.40 | $4.40 | $0.26 | No |
The market for 5.3 has already split into two bands. Around a dozen hosts mirror Z.ai's $1.40 and $4.40, and a smaller group undercuts it by 30 to 40 percent. SayGm's rate is the floor of that second group. Speed is a separate axis: Artificial Analysis measures Makora and Wafer at over 200 tokens per second, DeepInfra at 138, and Z.ai's own endpoint at 73, so paying list price does not buy you the fastest route.
GLM 5.2 API pricing: the cheapest hosts
GLM-5.2 arrived on 16 June (per OpenRouter's listing) with the jump to a 1-million-token context, and it is the version most people mean by "glm 5.2 price" searches. Because it is now the previous flagship, several hosts are discounting it to clear capacity, and this is the one table in this post where SayGm is not the cheapest.
| Host | Input | Output | Cache read | Confidential? |
|---|---|---|---|---|
| SayGm (glm-5.2) | $0.66 | $2.07 | $0.12 | No |
| SayGm (glm-5.2-tee) | $0.70 | $2.20 | $0.15 | Yes |
| DeepInfra (35% promo, FP4) | $0.49 | $1.56 | $0.09 | No |
| OpenRouter (cheapest route + 5.5% fee) | from $0.51 | from $1.65 | varies | No |
| Ambient | $0.60 | $2.00 | $0.15 | No |
| DigitalOcean | $0.70 | $2.20 | $0.11 | No |
| Phala | $1.26 | $3.00 | $0.22 | Yes |
| Z.ai (direct) | $1.40 | $4.40 | $0.26 | No |
DeepInfra's promotional rate is about 25 percent below SayGm's standard 5.2 route while it lasts, and Ambient at $0.60 and $2.00 also comes in under it. If the cheapest way to use GLM 5.2 is the only criterion, DeepInfra is the answer today, with the caveat that promotional pricing is temporary and the FP4 quant is worth checking against your evals. If you want 5.2 inside confidential hardware, SayGm's TEE route at $0.70 and $2.20 is the cheapest of the two hosts that offer one; Phala's is $1.26 and $3.00. OpenRouter also carries a rate-limited free 5.2 endpoint with a 33K context, which is fine for trying the model and not much else.
GLM 5.3 Flash: the budget tier
Flash is the GLM model that gets overlooked in pricing comparisons and shouldn't be. At Z.ai it is $0.15 in and $0.50 out; through SayGm it is $0.13 and $0.45. That is roughly a tenth of the flagship, and for classification, extraction, routing decisions and first-pass drafts it is often the right model. A million tokens of 5.3 Flash through SayGm costs about 20 cents at a 4:1 mix. If you are running an agent that makes many cheap calls and a few expensive ones, pointing the cheap calls at glm-5.3-flash and the hard ones at glm-5.3 is the single biggest lever on your GLM bill.
What 100 million GLM tokens costs
Same assumptions as the Kimi comparison: 100 million tokens a month, 80 million in and 20 million out, no cache discount applied.
| Host | GLM 5.3 | GLM 5.2 | GLM 5.3 Flash |
|---|---|---|---|
| SayGm | $116 | $94 (TEE $100) | $20 |
| DeepInfra | $132 | $70 (promo) | not listed |
| OpenRouter (cheapest route + fee) | $137 | $74 | varies |
| Phala (TEE) | $140 | $161 | not listed |
| Z.ai (direct) | $200 | $200 | $22 |
At a billion tokens a month the gap between SayGm and Z.ai direct on 5.3 is about $840; between SayGm and the list-price hosts, the same. Flash at that volume is $200 on SayGm, which is less than 100 million tokens of the flagship anywhere.
Who hosts GLM
Twenty-eight hosts serve GLM 5.3. Here are the most popular ones, and the pros and cons of each.
SayGm

Good at: the lowest price on 5.3 and Flash, and the only way to run GLM privately inside a sealed enclave for less than Z.ai charges for the ordinary version. Credit is prepaid and drawn down per request, so there is no subscription and no fee for topping up, and prices are published as data rather than on a page, so an app can read them. Existing OpenAI, Anthropic or Gemini code works by changing the base URL.
Missing: dedicated deployments, fine-tuning, and a sealed route for 5.3, which is still to come.
Z.ai

Good at: being first - they built the damn thing! New versions appear here a fortnight before anyone else can serve them, it is the only route to the Coding Plan, the free GLM-4.7-Flash and 4.5-Flash endpoints are Z.ai-only, and repeated context is discounted more heavily here than anywhere else.
Missing: value at volume & privacy. It is the most expensive route to every paid model in the lineup, and Artificial Analysis measures it as slower than most of the hosts charging less.
Note: Your usage here may also go towards training future Z.ai models - whether that is a positive depends on your perspective and use case.
OpenRouter

Good at: not going down. One key reaches 28 GLM routes, requests can be sorted by price, speed or tool-calling accuracy, and a failing host is swapped out automatically. Filters will exclude any host that trains on your prompts, and promotional prices are labelled as such, which is the quickest way to tell a real rate from a temporary one.
Missing: privacy of its own. A router reads every prompt that passes through it, and what happens next is whatever the host at the other end says in its policy. Credit purchases also carry a 5.5 percent fee. Our OpenRouter alternatives comparison covers the trade.
DeepInfra

Good at: giving you a dial. The same model comes in three speeds, cheaper if your work can wait in a queue and dearer if it cannot, and a reasoning_effort setting decides how hard the model thinks before answering, which on a reasoning model is a real lever on the bill. It is also the fastest of the cheaper hosts.
Missing: a confidential option. It also serves a compressed build of the weights, which is part of how the price stays low and worth testing against your own evals before you commit.
Fireworks

Good at: passing a security review. Fine-tuning, dedicated deployments with no rate limits, and function calling are all supported, which is usually what an enterprise checklist is asking for.
Missing: a discount. You pay Z.ai's list price for the privilege, and more if you want the faster tier.
Together

Good at: consolidation. If GLM is one of several open-weight models you run, a catalog past 200 of them with dedicated endpoints and reserved capacity is a reasonable place to put all of it.
Missing: any GLM-specific advantage. The price is list, and nothing here is tuned for this model in particular.
Which hosts keep your prompts private?
SayGm and Phala. Everywhere else, whoever runs the GPU can read your prompt, and the only thing stopping them is their own policy.
That matters more for GLM than for most models because of what people use it for: repository-scale coding, internal tools, and agents reading documents you would not email to a stranger. It is also the question a security review will ask about Z.ai, which is a Chinese company processing requests in China.
Switching to SayGm for GLM
GLM is OpenAI-compatible everywhere, so the change is a base URL and a model ID:
from openai import OpenAI
client = OpenAI(
base_url="https://api.saygm.com/v1",
api_key="YOUR_SAYGM_KEY",
)
response = client.chat.completions.create(
model="glm-5.3",
messages=[{"role": "user", "content": "Write the migration for this schema change."}],
)
print(response.choices[0].message.content)Use glm-5.3-flash for the cheap calls, glm-5.2-tee when the prompt has to stay sealed. The quickstart covers key setup, and the current rate for every GLM model is in the live catalog at api.saygm.com/v1/models, which is also where this post's SayGm figures come from.
Common questions about GLM pricing
How much does GLM 5.3 cost?
GLM 5.3 pricing is $1.40 per million input tokens and $4.40 output at Z.ai, $0.81 and $2.55 on SayGm, and between those two figures on most other hosts. GLM 5.3 Flash is $0.15 and $0.50 at Z.ai, $0.13 and $0.45 on SayGm. At 100 million tokens a month that is $200 direct versus $116 on SayGm for the flagship.
Is GLM 5.3 open source?
The 5.x line is released under the MIT licence, and 5.3's weights followed its Coding Plan launch by about two weeks. That licence is why so many hosts can serve it and why the third-party price sits below Z.ai's own.
What is GLM 5.3 Flash for?
Flash is the small, fast member of the 5.3 family, priced at about a tenth of the flagship. Use it for classification, extraction, summarisation and any agent step where a frontier model is overkill. For multi-step coding, stay on the full 5.3.
Is GLM cheaper than Kimi?
At list price, GLM 5.3 at $1.40 and $4.40 is a little dearer than Kimi K2.6 at $0.95 and $4.00, and much cheaper than Kimi K3 at $3.00 and $15.00. Through SayGm, GLM 5.3 at $0.81 and $2.55 sits between Kimi K2.6 at $0.28 and $1.67 and K3 at $1.80 and $8.99. Which is cheaper for your task depends on how many tokens each needs to finish it.
Why did GLM API prices go up in 2026?
Z.ai's list price moved from $1.00 and $3.20 on GLM 5 to $1.40 and $4.40 when GLM 5.1 launched in April, and has held that price across 5.2 and 5.3. Third-party hosts absorbed most of the increase, which is why the gap between Z.ai direct and the cheap band is wider now than it was in February.
About SayGm
SayGm is a drop-in inference gateway for teams who don't want to just take a company's word that their prompts are private. Every request is routed through a hardware-verified confidential environment - not even SayGm can see what's inside it. That's not a policy, it's provable. Swap in your existing OpenAI, Anthropic, or Gemini code and you're covered in minutes, at transparent, published rates with no hidden markup.
Say gm to AI at saygm.com.
- GLM
- Pricing


