Questions? hello@viberation.devGet supportBlogDocsChangelog
Get started

Z.ai's open-weight family, behind the cheap coding subscriptions.

GLM is the model family from Z.ai, and the reason its coding plan turns up in every "cheaper than Claude Code" thread — the weights are published, plenty of providers serve them, and the Flash models are some of the cheapest usable coding models anywhere. The 5.3 line runs to a million tokens of context, with Flash variants that also read images and video. Z.ai's own chat app is listed separately under Chats.

At a glance — GLM 4.7

Cost
$0.60 in · $2.20 out per 1M tokensMid-range
Memory
205K tokens — about 300 pages
Reads
Text
Can do
  • Use tools (agents)
  • Structured output
  • Thinking mode
Good at
Coding: better than 46%
Reliability
99.98% uptime, last 24h
Released
Dec 22, 2025
Advanced details
API model ID
z-ai/glm-4.7
Max output
131K tokens
Cached input
$0.11 per 1M
Tokenizer
Other
Coding index
45.3

Providers

The same model, hosted by different companies. OpenRouter picks one per request and falls back to the next if it fails.

ProviderInput /1MOutput /1MCached /1MMax outputUptime 24h
DeepInfra · fp4$0.40$1.75$0.08131K99.87%
Venice · fp4$0.40$1.929$0.0816K99.50%
AtlasCloud · fp8Degraded$0.52$1.85$0.12182K85.55%
Novita · fp8$0.54$1.98$0.099131K99.79%
Google$0.60$2.20—128K99.98%
Z.AI · fp4$0.60$2.20$0.11131K99.86%
Mancer 2 · fp4$0.70$2.50—118K99.21%

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_a
  • top_k
  • top_p

Live from OpenRouter, refreshed hourly. Scores from Artificial Analysis.

Key info

Pricing
Open source
Category
Models
Best for
Intermediate
Made by
Z.ai
Weights
Open
Good for
Coding on a budget

More in Models

  • France's frontier lab, with the small models everyone self-hosts.

    Models
  • Meta's current line, built for long multi-agent runs.

    Models
  • The open-weight family everyone else gets benchmarked against.

    Models
  • Google's open models — Gemini's research, small enough to self-host.

    Models