Home › Frontier Models › GLM-5
Model · Reference

GLM-5

Cost-sensitive agents

In one line

744B MoE with 40B active under MIT, 200K context and 128K output — the only model in this group you can get on both Bedrock and Vertex.

Why this oneWhat it is actually for

You reach for GLM-5 when you want an open-weight agentic model that your procurement team can actually buy through a US cloud. It is the unusual case in this group: MIT-licensed weights on Hugging Face and a first-party listing on Amazon Bedrock as zai.glm-5 across twelve regions, plus Google Vertex AI and Alibaba Cloud Model Studio. That combination means you can prototype against Z.ai's API, deploy through Bedrock for compliance, and fall back to self-hosting if either relationship sours — without changing models. At $1.00 in and $3.20 out on Z.ai it is not the cheapest model here, but it is the one with the fewest single points of failure. Its stated strength is long-horizon coding and engineering agents.

What it isArchitecture, lineage, training

GLM-5 is a sparse mixture-of-experts model from Z.ai, formerly Zhipu. The Hugging Face card for zai-org/GLM-5 gives 744B total parameters with 40B active, up from GLM-4.5's 355B total and 32B active, pretrained on 28.5T tokens against GLM-4.5's 23T, with post-training on an asynchronous RL infrastructure the team calls "slime". The card credits DeepSeek Sparse Attention for the deployment-cost reduction.

The architecture numbers come from config.json in that repository. The class is GlmMoeDsaForCausalLM — the "Dsa" is the sparse-attention indexer, visible as index_n_heads 32, index_head_dim 128 and index_topk 2048. hidden_size is 6144 across num_hidden_layers 78, with first_k_dense_replace 3, so the first three layers are dense feed-forward at intermediate_size 12288 and the remaining 75 are MoE. num_attention_heads is 64 with num_key_value_heads also 64, but the head geometry is multi-head latent attention: kv_lora_rank 512, q_lora_rank 2048, qk_nope_head_dim 192, qk_rope_head_dim 64 and v_head_dim 256. The MoE block has n_routed_experts 256 plus n_shared_experts 1, with num_experts_per_tok 8 and moe_intermediate_size 2048; routing uses a sigmoid scoring function with topk_method noaux_tc and routed_scaling_factor 2.5. vocab_size is 154880 and rope_theta is 1000000.

One number is worth reading carefully. max_position_embeddings is 202752 — about 198K, not a round 200K. The Z.ai docs and the Amazon Bedrock model card both say "200K", which is a rounding of the real ceiling. If you are building a chunker that assumes exactly 204800 positions, it will overflow. Licence is MIT, and a GLM-5-FP8 variant exists alongside the bf16 release.

At a glanceSee it

GLM-5 diagram

GLM-5 per token: sparse-attention indexing, three dense layers, then eight of 256 routed experts.

CapacityContext, output and what fits

FactValueSource
Model idglm-5 on Z.ai; zai.glm-5 on Amazon BedrockZ.ai GLM-5 guide; Bedrock model card
Context window200K stated; max_position_embeddings is exactly 202752Z.ai GLM-5 guide; config.json
Max output tokens128K. The API caps max_tokens at 131072Z.ai GLM-5 guide; chat-completion reference
Input modalitiesText onlyZ.ai GLM-5 guide; Bedrock card marks image, audio and video input as unsupported
Output modalitiesText onlyBedrock model card
Total / active parameters744B total, 40B activeHugging Face model card
Pretraining tokens28.5THugging Face model card
Knowledge cutoffNot disclosedNot stated on any Z.ai page read
Weights availableYes — zai-org/GLM-5, plus a GLM-5-FP8 variantHugging Face
LicenceMITHugging Face model card; Bedrock links to the GLM-5 repository LICENSE
Bedrock launch date2026-02-11, lifecycle Active, no EOL dateAmazon Bedrock model card

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
modelSelects the modelglm-5Z.ai also serves glm-5.1 and glm-5.2 on the same endpoint at higher prices; glm-5 is the cheapest of the three.
messagesConversation historyRequiredCached input is billed at $0.20 against $1.00, so a stable prefix is worth five times its length in savings.
max_tokensCaps generated tokensRange [1, 131072]The ceiling is the model's full 128K output. Long agent turns need this raised well above any default.
temperatureSampling randomnessRange [0.0, 1.0], default 1.0 for this seriesNote the ceiling is 1.0, not 2.0 — half the range most APIs give you. There is no way to go hotter than 1.0.
top_pNucleus sampling massRange [0.01, 1.0], default 0.95The model card's per-task recommendations move this: 0.95 for reasoning and code, 1.0 for agent work.
do_sampleSwitches sampling off entirelyBoolean, default trueSetting it false gives greedy decoding directly, which is cleaner than trying to reach determinism by driving temperature toward 0.
top_kTop-k truncationNot documented on the Z.ai chat-completion referenceDo not send it. Use top_p or do_sample.
thinkingTurns the reasoning pass on or offObject; {"type": "enabled"} is the default, {"type": "disabled"} turns it offUnlike Qwen3-Max, thinking is on by default here. Disabling it is the fast path for classification and extraction.
reasoning_effortReasoning levelGLM-5.2 only. Not accepted by glm-5On glm-5 your reasoning control is binary — on or off — with no intermediate levels and no numeric budget.
toolsFunction, retrieval and web-search toolsArray, maximum 128Z.ai's tool array covers more than functions; retrieval and web search are declared the same way.
tool_choiceForces or forbids tool useDefault autoauto is the sane default for an agentic model; forcing calls tends to produce fabricated arguments.
tool_streamStreams function-call deltasBoolean, default falseLets a UI show a tool call forming rather than waiting for the whole call object.
response_formatOutput shapetext default, or json_objectStructured output is supported; a strict JSON-schema mode is not documented on this endpoint.
stopHalt sequencesArray, maximum 4 itemsOnly four — much tighter than DeepSeek's sixteen. Budget them.
streamServer-sent eventsBoolean, default falseWith thinking on by default, streaming is close to required for interactive use.
request_id / user_idRequest and end-user identifiers6–64 and 6–128 charactersNote the minimum lengths — a short id is rejected, which trips up integrations passing numeric user keys.

Three knobs earn their keep. thinking is the first: on by default, and switching it off is the single largest latency and cost change available on this model. do_sample is the second and is under-used — it gives you clean greedy decoding without the awkwardness of a temperature range that stops at 1.0. Third is max_tokens: GLM-5 is sold as a long-horizon agentic model and the card's own reasoning recipe uses max_new_tokens of 131072, so leaving this at a small default quietly defeats the thing you bought it for.

SamplingShaping the output distribution

GLM-5's sampling range is narrower than most. temperature is documented over [0.0, 1.0] with a default of 1.0 for this series, so unlike the OpenAI-shaped APIs there is no headroom above 1.0 at all — 1.0 is the hot end. top_p runs [0.01, 1.0] with a default of 0.95. Because the temperature ceiling is the default, essentially all of your tuning is downward, and the cleaner way to reach determinism is do_sample: false for greedy decoding rather than pushing temperature to zero. The model card publishes per-task recipes, which is the best guidance available: reasoning at temperature=1.0, top_p=0.95, max_new_tokens=131072; code at temperature=0.7, top_p=0.95; agent tasks at temperature=1.0, top_p=1.0. Those are worth following literally — the agent recipe deliberately disables nucleus truncation, which suggests the model's tool-selection behaviour depends on tail mass that top_p 0.95 would cut. top_k is not documented on the Z.ai endpoint. Start with the card's recipe for your task shape and change nothing else until you have a baseline.

ReasoningThinking, effort and budgets

GLM-5 has a thinking mode that is enabled by default, controlled by a thinking object taking {"type": "enabled"} or {"type": "disabled"}. That default is the opposite of Qwen3-Max's and matters for anyone benchmarking the two side by side. The control is binary: reasoning_effort exists on the Z.ai endpoint but the reference marks it as GLM-5.2 only, so on glm-5 there is no intermediate level and no numeric token budget. If you need finer control over how much the model deliberates, your options are to disable thinking entirely, to move to GLM-5.2, or to self-host and cap generation yourself. Z.ai's documentation for glm-5 does not state whether reasoning tokens are billed at the output rate or whether the reasoning trace is returned in a separate field on this model — both are worth verifying empirically against usage before you build cost models on assumptions. The Bedrock model card describes the model as "chat-completion" and does not enumerate a reasoning-content field either.

ToolsFunction calling and server tools

Tool calling is a headline capability — Z.ai positions GLM-5 for "long-horizon planning" agentic engineering and the card claims 77.8 on SWE-bench Verified, which is a vendor-reported number. The tools array takes up to 128 entries and covers function, retrieval and web-search tool types; tool_choice defaults to auto. A useful extra is tool_stream, which streams the function-call object as it forms rather than delivering it whole — helpful for showing agent progress in a UI. Structured output is supported through response_format with json_object; a strict JSON-schema mode is not documented on this endpoint, so validate and retry rather than assuming schema conformance. Two integration traps are worth flagging. First, request_id and user_id have minimum lengths of 6 characters, so passing a short numeric identifier fails validation rather than being ignored. Second, if you consume GLM-5 through Amazon Bedrock, you are on a different API surface — the Converse and Invoke APIs, or the bedrock-mantle Chat Completions endpoint — and the Z.ai-specific fields above do not exist there.

CostPrice, caching, batching, what drives the bill

List price from Z.ai's own pricing page, USD per million tokens:

ItemPrice per 1M tokens
Input$1.00
Cached input$0.20
Output$3.20

For comparison on the same page, GLM-5.1, GLM-5.2 and GLM-5.3 all cost $1.40 in, $0.26 cached and $4.40 out, and GLM-5-Turbo is $1.20 / $0.24 / $4.00 — so plain glm-5 is the cheapest member of the family, and the version you pick is itself a cost decision. GLM-5.3 arrived on 2026-08-19 at that same $1.40/$4.40 rate and did not displace this row: Z.ai still lists and sells plain glm-5 at $1.00/$3.20, so the newer version is an option beside it, not a migration you are forced into. Cached input storage is listed as "Limited-time Free", which means the discount structure may change; check the page before committing to a caching-heavy design. Thinking is on by default and its tokens are generated tokens, so output dominates unless you disable it. Prices through other surfaces are set by those vendors, not by Z.ai: Amazon Bedrock bills zai.glm-5 under its own pricing with Standard, Priority and Flex service tiers available and Reserved not offered. Z.ai also sells a GLM Coding Plan subscription starting at $18 per month for use inside coding tools rather than by the token, which is a different economic model entirely and worth evaluating if your usage is a single developer rather than a pipeline.

Where it runsSurfaces and availability

SurfaceAvailableNotes
Z.ai APIYeshttps://api.z.ai/api/paas/v4/chat/completions, model id glm-5. OpenAI SDK compatible by base-URL swap.
Amazon BedrockYesModel id zai.glm-5, launched 2026-02-11. 200K context, 128K output. Converse, Invoke and Chat Completions APIs; bedrock-runtime and bedrock-mantle endpoints. Twelve regions across US, EU, APAC and South America. Standard, Priority and Flex tiers; Reserved not supported.
Google Vertex AIYesListed as GLM 5 in the Vertex managed-API partner model list, alongside GLM 4.7.
Alibaba Cloud Model StudioYesListed as glm-5 among the GLM models offered there.
Microsoft FoundryUnverifiedNot present in Foundry's models-sold-by-Azure or partners-and-community lists read here. Third-party reports of a Fireworks-hosted route could not be confirmed on a Microsoft page.
Hugging Face weightsYeszai-org/GLM-5 and zai-org/GLM-5-FP8, MIT licence.
Self-hostingYes744B total / 40B active. The FP8 variant is the realistic deployment target.

StrengthsWhat it is good at

  • MIT-licensed weights and first-party availability on Amazon Bedrock as zai.glm-5 across twelve regions — the only model in this group with both an open licence and US-cloud distribution.
  • Real architecture transparency: config.json publishes the full MoE and latent-attention geometry, and the card publishes pretraining token count and the RL infrastructure by name.
  • Per-task sampling recipes on the model card — separate settings for reasoning, code and agent work — rather than one generic recommendation.
  • 128K max output against a 200K context, so more than half the window can be response, which suits long code generation.
  • Cached input at $0.20 against $1.00, currently with free cache storage, and a $18/month coding-plan alternative for single-developer use.

LimitsWhere it falls down

  • Text only in and out. No vision at all — glm-5v-turbo is a separate model, so any multimodal step needs a second call.
  • Reasoning control is binary. reasoning_effort is documented as GLM-5.2 only, so on glm-5 there is no intermediate level and no token budget.
  • Sampling range is half the usual: temperature tops out at 1.0, which is also the default, so there is no headroom above the tuned operating point.
  • Only four stop sequences, against sixteen on DeepSeek — tight for agent framing.
  • The advertised "200K" context is really 202752 positions, and GLM-5.2 has moved on to a 1M window, so this model is already a generation behind its own family on long context.

Against its neighboursHow it compares

Against DeepSeek V4 Pro, the answer now depends on the clock: DeepSeek moved that model to a peak/off-peak schedule, so GLM-5 is slightly cheaper during DeepSeek's peak windows ($1.00/$3.20 against $1.32/$3.96) and about 1.5x more expensive off-peak (against $0.66/$1.98). GLM-5 has a 200K context against DeepSeek's 1M. What it buys is distribution: Bedrock and Vertex both carry GLM-5 and neither carries DeepSeek V4, so if your deployment has to sit behind an AWS or Google contract, GLM-5 is the one that exists. Both are MIT. Against its own successors GLM-5.1, GLM-5.2 and GLM-5.3 — all three at $1.40/$4.40 — plain glm-5 is cheaper ($1.00/$3.20) but gives up a 1M context window and the reasoning_effort control, a meaningful gap if you need graded reasoning. GLM-5.3, added on 2026-08-19, is the newest of them and carries no price premium over 5.1 or 5.2. Against Qwen3-Max, GLM-5 is the open-weight option with a portable deployment story where Qwen3-Max is closed and Model-Studio-only.

Getting startedThe smallest call that works

code
POST https://api.z.ai/api/paas/v4/chat/completions
Authorization: Bearer $ZAI_API_KEY
Content-Type: application/json

{
  "model": "glm-5",
  "messages": [
    {"role": "user", "content": "Hello"}
  ]
}

Raise max_tokens first — the card's own reasoning recipe uses 131072, and leaving a small default in place is the commonest way to truncate this model mid-plan. Then match the card's per-task sampling: temperature 0.7 with top_p 0.95 for code, or 1.0 and 1.0 for agent work. Send "thinking": {"type": "disabled"} for extraction.

SourcesWhere every claim above came from

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page Kimi K3 is available on Amazon Bedrock with a 1M-token context window, making it callable through managed AWS endpoints.

    China frontier labs · 18 Sep 2026 · source

RecentAuto-linked from the brief, not a rewrite of this page

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning