744B MoE with 40B active under MIT, 200K context and 128K output — the only model in this group you can get on both Bedrock and Vertex.
Why this oneWhat it is actually for
You reach for GLM-5 when you want an open-weight agentic model that your procurement team can actually buy through a US cloud. It is the unusual case in this group: MIT-licensed weights on Hugging Face and a first-party listing on Amazon Bedrock as zai.glm-5 across twelve regions, plus Google Vertex AI and Alibaba Cloud Model Studio. That combination means you can prototype against Z.ai's API, deploy through Bedrock for compliance, and fall back to self-hosting if either relationship sours — without changing models. At $1.00 in and $3.20 out on Z.ai it is not the cheapest model here, but it is the one with the fewest single points of failure. Its stated strength is long-horizon coding and engineering agents.
What it isArchitecture, lineage, training
GLM-5 is a sparse mixture-of-experts model from Z.ai, formerly Zhipu. The Hugging Face card for zai-org/GLM-5 gives 744B total parameters with 40B active, up from GLM-4.5's 355B total and 32B active, pretrained on 28.5T tokens against GLM-4.5's 23T, with post-training on an asynchronous RL infrastructure the team calls "slime". The card credits DeepSeek Sparse Attention for the deployment-cost reduction.
The architecture numbers come from config.json in that repository. The class is GlmMoeDsaForCausalLM — the "Dsa" is the sparse-attention indexer, visible as index_n_heads 32, index_head_dim 128 and index_topk 2048. hidden_size is 6144 across num_hidden_layers 78, with first_k_dense_replace 3, so the first three layers are dense feed-forward at intermediate_size 12288 and the remaining 75 are MoE. num_attention_heads is 64 with num_key_value_heads also 64, but the head geometry is multi-head latent attention: kv_lora_rank 512, q_lora_rank 2048, qk_nope_head_dim 192, qk_rope_head_dim 64 and v_head_dim 256. The MoE block has n_routed_experts 256 plus n_shared_experts 1, with num_experts_per_tok 8 and moe_intermediate_size 2048; routing uses a sigmoid scoring function with topk_method noaux_tc and routed_scaling_factor 2.5. vocab_size is 154880 and rope_theta is 1000000.
One number is worth reading carefully. max_position_embeddings is 202752 — about 198K, not a round 200K. The Z.ai docs and the Amazon Bedrock model card both say "200K", which is a rounding of the real ceiling. If you are building a chunker that assumes exactly 204800 positions, it will overflow. Licence is MIT, and a GLM-5-FP8 variant exists alongside the bf16 release.
At a glanceSee it
GLM-5 per token: sparse-attention indexing, three dense layers, then eight of 256 routed experts.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| Model id | glm-5 on Z.ai; zai.glm-5 on Amazon Bedrock | Z.ai GLM-5 guide; Bedrock model card |
| Context window | 200K stated; max_position_embeddings is exactly 202752 | Z.ai GLM-5 guide; config.json |
| Max output tokens | 128K. The API caps max_tokens at 131072 | Z.ai GLM-5 guide; chat-completion reference |
| Input modalities | Text only | Z.ai GLM-5 guide; Bedrock card marks image, audio and video input as unsupported |
| Output modalities | Text only | Bedrock model card |
| Total / active parameters | 744B total, 40B active | Hugging Face model card |
| Pretraining tokens | 28.5T | Hugging Face model card |
| Knowledge cutoff | Not disclosed | Not stated on any Z.ai page read |
| Weights available | Yes — zai-org/GLM-5, plus a GLM-5-FP8 variant | Hugging Face |
| Licence | MIT | Hugging Face model card; Bedrock links to the GLM-5 repository LICENSE |
| Bedrock launch date | 2026-02-11, lifecycle Active, no EOL date | Amazon Bedrock model card |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
model | Selects the model | glm-5 | Z.ai also serves glm-5.1 and glm-5.2 on the same endpoint at higher prices; glm-5 is the cheapest of the three. |
messages | Conversation history | Required | Cached input is billed at $0.20 against $1.00, so a stable prefix is worth five times its length in savings. |
max_tokens | Caps generated tokens | Range [1, 131072] | The ceiling is the model's full 128K output. Long agent turns need this raised well above any default. |
temperature | Sampling randomness | Range [0.0, 1.0], default 1.0 for this series | Note the ceiling is 1.0, not 2.0 — half the range most APIs give you. There is no way to go hotter than 1.0. |
top_p | Nucleus sampling mass | Range [0.01, 1.0], default 0.95 | The model card's per-task recommendations move this: 0.95 for reasoning and code, 1.0 for agent work. |
do_sample | Switches sampling off entirely | Boolean, default true | Setting it false gives greedy decoding directly, which is cleaner than trying to reach determinism by driving temperature toward 0. |
top_k | Top-k truncation | Not documented on the Z.ai chat-completion reference | Do not send it. Use top_p or do_sample. |
thinking | Turns the reasoning pass on or off | Object; {"type": "enabled"} is the default, {"type": "disabled"} turns it off | Unlike Qwen3-Max, thinking is on by default here. Disabling it is the fast path for classification and extraction. |
reasoning_effort | Reasoning level | GLM-5.2 only. Not accepted by glm-5 | On glm-5 your reasoning control is binary — on or off — with no intermediate levels and no numeric budget. |
tools | Function, retrieval and web-search tools | Array, maximum 128 | Z.ai's tool array covers more than functions; retrieval and web search are declared the same way. |
tool_choice | Forces or forbids tool use | Default auto | auto is the sane default for an agentic model; forcing calls tends to produce fabricated arguments. |
tool_stream | Streams function-call deltas | Boolean, default false | Lets a UI show a tool call forming rather than waiting for the whole call object. |
response_format | Output shape | text default, or json_object | Structured output is supported; a strict JSON-schema mode is not documented on this endpoint. |
stop | Halt sequences | Array, maximum 4 items | Only four — much tighter than DeepSeek's sixteen. Budget them. |
stream | Server-sent events | Boolean, default false | With thinking on by default, streaming is close to required for interactive use. |
request_id / user_id | Request and end-user identifiers | 6–64 and 6–128 characters | Note the minimum lengths — a short id is rejected, which trips up integrations passing numeric user keys. |
Three knobs earn their keep. thinking is the first: on by default, and switching it off is the single largest latency and cost change available on this model. do_sample is the second and is under-used — it gives you clean greedy decoding without the awkwardness of a temperature range that stops at 1.0. Third is max_tokens: GLM-5 is sold as a long-horizon agentic model and the card's own reasoning recipe uses max_new_tokens of 131072, so leaving this at a small default quietly defeats the thing you bought it for.
SamplingShaping the output distribution
GLM-5's sampling range is narrower than most. temperature is documented over [0.0, 1.0] with a default of 1.0 for this series, so unlike the OpenAI-shaped APIs there is no headroom above 1.0 at all — 1.0 is the hot end. top_p runs [0.01, 1.0] with a default of 0.95. Because the temperature ceiling is the default, essentially all of your tuning is downward, and the cleaner way to reach determinism is do_sample: false for greedy decoding rather than pushing temperature to zero. The model card publishes per-task recipes, which is the best guidance available: reasoning at temperature=1.0, top_p=0.95, max_new_tokens=131072; code at temperature=0.7, top_p=0.95; agent tasks at temperature=1.0, top_p=1.0. Those are worth following literally — the agent recipe deliberately disables nucleus truncation, which suggests the model's tool-selection behaviour depends on tail mass that top_p 0.95 would cut. top_k is not documented on the Z.ai endpoint. Start with the card's recipe for your task shape and change nothing else until you have a baseline.
ReasoningThinking, effort and budgets
GLM-5 has a thinking mode that is enabled by default, controlled by a thinking object taking {"type": "enabled"} or {"type": "disabled"}. That default is the opposite of Qwen3-Max's and matters for anyone benchmarking the two side by side. The control is binary: reasoning_effort exists on the Z.ai endpoint but the reference marks it as GLM-5.2 only, so on glm-5 there is no intermediate level and no numeric token budget. If you need finer control over how much the model deliberates, your options are to disable thinking entirely, to move to GLM-5.2, or to self-host and cap generation yourself. Z.ai's documentation for glm-5 does not state whether reasoning tokens are billed at the output rate or whether the reasoning trace is returned in a separate field on this model — both are worth verifying empirically against usage before you build cost models on assumptions. The Bedrock model card describes the model as "chat-completion" and does not enumerate a reasoning-content field either.
ToolsFunction calling and server tools
Tool calling is a headline capability — Z.ai positions GLM-5 for "long-horizon planning" agentic engineering and the card claims 77.8 on SWE-bench Verified, which is a vendor-reported number. The tools array takes up to 128 entries and covers function, retrieval and web-search tool types; tool_choice defaults to auto. A useful extra is tool_stream, which streams the function-call object as it forms rather than delivering it whole — helpful for showing agent progress in a UI. Structured output is supported through response_format with json_object; a strict JSON-schema mode is not documented on this endpoint, so validate and retry rather than assuming schema conformance. Two integration traps are worth flagging. First, request_id and user_id have minimum lengths of 6 characters, so passing a short numeric identifier fails validation rather than being ignored. Second, if you consume GLM-5 through Amazon Bedrock, you are on a different API surface — the Converse and Invoke APIs, or the bedrock-mantle Chat Completions endpoint — and the Z.ai-specific fields above do not exist there.
CostPrice, caching, batching, what drives the bill
List price from Z.ai's own pricing page, USD per million tokens:
| Item | Price per 1M tokens |
|---|---|
| Input | $1.00 |
| Cached input | $0.20 |
| Output | $3.20 |
For comparison on the same page, GLM-5.1, GLM-5.2 and GLM-5.3 all cost $1.40 in, $0.26 cached and $4.40 out, and GLM-5-Turbo is $1.20 / $0.24 / $4.00 — so plain glm-5 is the cheapest member of the family, and the version you pick is itself a cost decision. GLM-5.3 arrived on 2026-08-19 at that same $1.40/$4.40 rate and did not displace this row: Z.ai still lists and sells plain glm-5 at $1.00/$3.20, so the newer version is an option beside it, not a migration you are forced into. Cached input storage is listed as "Limited-time Free", which means the discount structure may change; check the page before committing to a caching-heavy design. Thinking is on by default and its tokens are generated tokens, so output dominates unless you disable it. Prices through other surfaces are set by those vendors, not by Z.ai: Amazon Bedrock bills zai.glm-5 under its own pricing with Standard, Priority and Flex service tiers available and Reserved not offered. Z.ai also sells a GLM Coding Plan subscription starting at $18 per month for use inside coding tools rather than by the token, which is a different economic model entirely and worth evaluating if your usage is a single developer rather than a pipeline.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Z.ai API | Yes | https://api.z.ai/api/paas/v4/chat/completions, model id glm-5. OpenAI SDK compatible by base-URL swap. |
| Amazon Bedrock | Yes | Model id zai.glm-5, launched 2026-02-11. 200K context, 128K output. Converse, Invoke and Chat Completions APIs; bedrock-runtime and bedrock-mantle endpoints. Twelve regions across US, EU, APAC and South America. Standard, Priority and Flex tiers; Reserved not supported. |
| Google Vertex AI | Yes | Listed as GLM 5 in the Vertex managed-API partner model list, alongside GLM 4.7. |
| Alibaba Cloud Model Studio | Yes | Listed as glm-5 among the GLM models offered there. |
| Microsoft Foundry | Unverified | Not present in Foundry's models-sold-by-Azure or partners-and-community lists read here. Third-party reports of a Fireworks-hosted route could not be confirmed on a Microsoft page. |
| Hugging Face weights | Yes | zai-org/GLM-5 and zai-org/GLM-5-FP8, MIT licence. |
| Self-hosting | Yes | 744B total / 40B active. The FP8 variant is the realistic deployment target. |
StrengthsWhat it is good at
- MIT-licensed weights and first-party availability on Amazon Bedrock as
zai.glm-5across twelve regions — the only model in this group with both an open licence and US-cloud distribution. - Real architecture transparency:
config.jsonpublishes the full MoE and latent-attention geometry, and the card publishes pretraining token count and the RL infrastructure by name. - Per-task sampling recipes on the model card — separate settings for reasoning, code and agent work — rather than one generic recommendation.
- 128K max output against a 200K context, so more than half the window can be response, which suits long code generation.
- Cached input at $0.20 against $1.00, currently with free cache storage, and a $18/month coding-plan alternative for single-developer use.
LimitsWhere it falls down
- Text only in and out. No vision at all —
glm-5v-turbois a separate model, so any multimodal step needs a second call. - Reasoning control is binary.
reasoning_effortis documented as GLM-5.2 only, so onglm-5there is no intermediate level and no token budget. - Sampling range is half the usual:
temperaturetops out at 1.0, which is also the default, so there is no headroom above the tuned operating point. - Only four
stopsequences, against sixteen on DeepSeek — tight for agent framing. - The advertised "200K" context is really 202752 positions, and GLM-5.2 has moved on to a 1M window, so this model is already a generation behind its own family on long context.
Against its neighboursHow it compares
Against DeepSeek V4 Pro, the answer now depends on the clock: DeepSeek moved that model to a peak/off-peak schedule, so GLM-5 is slightly cheaper during DeepSeek's peak windows ($1.00/$3.20 against $1.32/$3.96) and about 1.5x more expensive off-peak (against $0.66/$1.98). GLM-5 has a 200K context against DeepSeek's 1M. What it buys is distribution: Bedrock and Vertex both carry GLM-5 and neither carries DeepSeek V4, so if your deployment has to sit behind an AWS or Google contract, GLM-5 is the one that exists. Both are MIT. Against its own successors GLM-5.1, GLM-5.2 and GLM-5.3 — all three at $1.40/$4.40 — plain glm-5 is cheaper ($1.00/$3.20) but gives up a 1M context window and the reasoning_effort control, a meaningful gap if you need graded reasoning. GLM-5.3, added on 2026-08-19, is the newest of them and carries no price premium over 5.1 or 5.2. Against Qwen3-Max, GLM-5 is the open-weight option with a portable deployment story where Qwen3-Max is closed and Model-Studio-only.
Getting startedThe smallest call that works
POST https://api.z.ai/api/paas/v4/chat/completions
Authorization: Bearer $ZAI_API_KEY
Content-Type: application/json
{
"model": "glm-5",
"messages": [
{"role": "user", "content": "Hello"}
]
}Raise max_tokens first — the card's own reasoning recipe uses 131072, and leaving a small default in place is the commonest way to truncate this model mid-plan. Then match the card's per-task sampling: temperature 0.7 with top_p 0.95 for code, or 1.0 and 1.0 for agent work. Send "thinking": {"type": "disabled"} for extraction.
SourcesWhere every claim above came from
- GLM-5 — Z.ai API documentation
- Chat Completion API reference — Z.ai
- Pricing — Z.ai API documentation
- Quick Start — Z.ai API documentation
- zai-org/GLM-5 — Hugging Face model card
- zai-org/GLM-5 config.json
- GLM 5 model card — Amazon Bedrock
- Vertex AI managed models for MaaS — Google Cloud
- GLM — Alibaba Cloud Model Studio
- Could not confirm from a primary source as of 2026-07-25: the knowledge cutoff, whether reasoning tokens on
glm-5are billed at the output rate, the field name carrying the reasoning trace on this model, the layer-level expert count breakdown beyond whatconfig.jsonstates, and Microsoft Foundry availability. The SWE-bench Verified figure of 77.8 is Z.ai's own claim, not an independent result. Corrections to the site's stored row: it listed a 128K context — the vendor and Bedrock both say 200K, andconfig.jsonsays 202752 — and it listed modalities as "Text, vision", but Z.ai and Bedrock both document GLM-5 as text-only.
What changedWhat changed here
Updated this page Kimi K3 is available on Amazon Bedrock with a 1M-token context window, making it callable through managed AWS endpoints.
- Zhipu AI's GLM-5.3-Flash Runs on 100,000 Domestic Chips, Tops OpenRouter Usage
Zhipu's GLM-5.3-Flash now tops OpenRouter usage and runs on 100,000 domestic chips — a signal that open-weight Chinese models are becoming a realistic option for cost-sensitive builds.
Three kinds of claim, strongest first. Signal runs every morning.