Questions? hello@viberation.devGet supportBlogDocsChangelog
Get started

Z.ai's open-weight family, behind the cheap coding subscriptions.

GLM is the model family from Z.ai, and the reason its coding plan turns up in every "cheaper than Claude Code" thread — the weights are published, plenty of providers serve them, and the Flash models are some of the cheapest usable coding models anywhere. The 5.3 line runs to a million tokens of context, with Flash variants that also read images and video. Z.ai's own chat app is listed separately under Chats.

At a glance — GLM 5.3

Cost
$1.40 in · $4.40 out per 1M tokensMid-range
Memory
1.3M tokens — about 1,950 pages
Reads
Text
Can do
  • Use tools (agents)
  • Structured output
  • Thinking mode
Good at
Coding: better than 87% · Agentic work: better than 95% · Overall: better than 86%
Reliability
99.98% uptime, last 24h
Released
Aug 18, 2026
Advanced details
API model ID
z-ai/glm-5.3
Max output
944K tokens
Cached input
$0.26 per 1M
Tokenizer
Other
Coding index
74.8
Agentic work index
53.1
Overall index
44.8

Providers

The same model, hosted by different companies. OpenRouter picks one per request and falls back to the next if it fails.

ProviderInput /1MOutput /1MCached /1MMax outputUptime 24h
Baidu · fp8$0.561$1.764$0.104131K99.98%
DeepInfra · fp4$0.563$2.50$0.125131K98.65%
Morph$0.583$1.833$0.096944K99.76%
Inceptron · fp4$0.597$2.502$0.152944K99.29%
InferenceNet$0.684$2.28$0.114131K99.74%
Relace$0.70$2.20$0.13131K—
Reka$0.761$2.574$0.152236K99.81%
Io Net · fp8$0.77$2.62$0.1466K99.74%
Sail Research · us$0.77$4.00$0.19944K99.93%
Sail Research · fp8$0.77$4.00$0.19944K99.84%
Novita · fp8$0.783$2.46$0.145131K99.85%
Phala$0.84$2.64$0.156131K99.50%
DigitalOcean$0.91$2.86$0.169128K99.27%
GMICloud · fp8$0.98$3.08$0.182944K99.73%
Makora · fp4$1.05$4.20$0.19128K98.69%
SiliconFlow · fp8$1.12$3.52$0.208262K99.40%
Alibaba$1.19$3.74$0.238131K99.82%
Decart · fp4$1.19$3.74$0.196944K99.76%
Friendli$1.26$3.96$0.234944K99.98%
AkashML · fp8$1.30$4.40$0.26131K99.73%
Wafer · us$1.40$4.40$0.26944K99.86%
BaseTen · fp4Degraded$1.40$4.40$0.14262K99.27%
Mistral · nvfp4$1.40$4.40$0.14131K99.91%
Crusoe · fp4$1.40$4.40$0.26944K99.65%
PrimeIntellect$1.40$4.40$0.26131K99.97%
Wafer$1.40$4.40$0.26944K99.87%
Venice$1.40$4.40$0.26131K98.67%
Together$1.40$4.40$0.26944K99.83%
Parasail · fp8$1.40$4.40$0.26944K99.57%
Modal$1.40$4.40$0.26944K99.38%
BaseTen · fp4Degraded$1.40$4.40$0.14262K98.06%
Fireworks$1.40$4.40$0.26944K98.81%
Cloudflare$1.40$4.40$0.261.2M99.96%
AtlasCloud · fp8$1.40$4.40$0.26131K99.70%
Z.AI · fp8$1.40$4.40$0.26131K99.93%
BaseTen · fast$2.10$6.60$0.21262K99.71%
BaseTen · fast$2.10$6.60$0.21262K99.86%
Alibaba · fast$2.80$8.80$0.56131K99.70%

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • parallel_tool_calls
  • presence_penalty
  • reasoning
  • reasoning_effort
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Live from OpenRouter, refreshed hourly. Scores from Artificial Analysis.

Key info

Pricing
Open source
Category
Models
Best for
Intermediate
Made by
Z.ai
Weights
Open
Good for
Coding on a budget

More in Models

  • France's frontier lab, with the small models everyone self-hosts.

    Models
  • Meta's current line, built for long multi-agent runs.

    Models
  • The open-weight family everyone else gets benchmarked against.

    Models
  • Google's open models — Gemini's research, small enough to self-host.

    Models