Home › Frontier Models › Qwen3 Max
Model · Reference

Qwen3 Max

Multilingual apps at API scale

In one line

Alibaba's closed-weight 256K flagship snapshot, priced in three tiers by input length, with thinking off by default — not an open-weight model.

Why this oneWhat it is actually for

You reach for Qwen3-Max when you are already inside Alibaba Cloud Model Studio, want the Qwen family's multilingual behaviour, and specifically want the January 2026 snapshot rather than the newer qwen3.7-max that has replaced it on the model overview page. Pinning to qwen3-max-2026-01-23 is a legitimate reason to be here: it is a stable, still-billable model with a fixed evaluation baseline, and moving to 3.7 means re-qualifying every prompt. It is also the cheapest Qwen Max tier if your prompts stay under 32K, at $1.20 per million input. You reach elsewhere if you need open weights — despite what a lot of secondary sources imply, the Max line is not open-weighted — or a context window past 256K.

What it isArchitecture, lineage, training

Qwen3-Max is Alibaba's proprietary flagship in the Qwen3 generation. The model id qwen3-max currently resolves to the snapshot qwen3-max-2026-01-23, and a qwen3-max-preview snapshot also exists. Alibaba's text-generation model listing gives it a 256K context window and marks thinking mode, function calling, built-in tools and structured output all as supported.

What Alibaba does not disclose is the architecture. No parameter count, no expert count, no active-parameter figure, no layer count, no attention description appears on any Alibaba Cloud Model Studio page read for this article. The site's stored row said "MoE" — that is a reasonable inference from the family, but it is not a vendor-stated fact and is not asserted here.

The open-weight question deserves to be settled plainly, because the roster row got it wrong. There is no Qwen3-Max repository under the official Qwen Hugging Face organisation; a search returns only community re-uploads and derivatives from unaffiliated accounts. Alibaba open-weights a large part of the Qwen family — the Qwen3, Qwen3.5 and Qwen3.6 dense and MoE releases under Apache 2.0 — but the Max line is the commercial, API-only tier. You cannot download Qwen3-Max, cannot fine-tune it on your own hardware, and cannot self-host it. Treat "API / open-weight" on this row as incorrect.

One more thing worth knowing before you build on it: qwen3-max no longer appears on Model Studio's Supported Models overview page, which now lists qwen3.7-max and qwen3.8-max-preview. It is still fully priced and served across Beijing, Hong Kong, Singapore, Tokyo, Frankfurt and US Virginia, and carries no deprecation notice — but it is a superseded generation.

At a glanceSee it

Qwen3 Max diagram

Qwen3-Max prices by input length: three tiers, and the whole bill doubles or triples as the prompt grows.

CapacityContext, output and what fits

FactValueSource
Model idqwen3-max, "currently equivalent to qwen3-max-2026-01-23"Model Studio model pricing
Deprecationqwen3-max-2026-01-23 is scheduled for deprecation on October 10, 2026, 00:00:00, with qwen3.7-max named as the replacement model. The unversioned qwen3-max alias is not itself listed as deprecatedModel Studio model deprecation
Context window256K. Supported indirectly: the model's top pricing tier is bounded at "128K < Token ≤ 256K". qwen3-max no longer appears in the Model Studio supported-models list, so no separately stated context figure could be locatedModel Studio model pricing
Max output tokensCould not be confirmed from a primary source as of 2026-07-26Not stated on any Model Studio page read
Input modalitiesText. The pricing page describes it in the Qwen-Max text series; vision input could not be confirmedModel Studio model pricing
Output modalitiesText, plus reasoning_content when thinking is enabled. The pricing page lists "Non-Thinking and Thinking modes"Model Studio model pricing; deep thinking guide
ArchitectureNot disclosed — no parameter count, expert count or layer count published—
Knowledge cutoffNot disclosed—
Weights availableNo. No Qwen3-Max repository under the official Qwen Hugging Face organisation as of 2026-07-26Hugging Face model search
LicenceNot applicable — commercial API only, no weights released—
RegionsChina Beijing, Hong Kong, Singapore, Tokyo, Frankfurt, US VirginiaModel Studio model pricing
Free quota1 million tokens, valid 90 days after activation; International deployment onlyModel Studio model pricing

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
messagesConversation historyRequiredTotal input length decides which of three price tiers the whole request bills at, so trimming history has an outsized effect here.
max_tokensCaps generated tokensInteger; ceiling not documentedBecause the model's output ceiling is undocumented, set this explicitly rather than relying on a default you cannot look up.
temperatureSampling randomnessRange [0, 2). Alibaba says it cannot be set to 0The floor is exclusive — for near-deterministic output use a very small positive value, not 0.
top_pNucleus sampling massRange (0, 1.0)Alibaba's guidance is to keep defaults and adjust only one of temperature or top_p, never both.
top_k, frequency_penaltyTop-k truncation and repetition penaltyNot documented on the OpenAI-compatibility parameter pageDo not assume they work. Use presence_penalty for repetition and top_p for truncation.
presence_penaltyDiscourages repeated tokensRange [-2.0, 2.0]The only repetition control documented on this interface. Small positive values loosen looping without derailing structure.
nNumber of responsesDefault 1, range 1–4Every candidate bills separately and at the tier your input length set.
seedRandom seedInteger. The OpenAI-compatibility reference states it "supports unsigned 64-bit integers"Improves reproducibility across runs. There is no documented 32-bit restriction on the OpenAI-compatible protocol.
stopHalt sequencesString or array; strings or token ids; default noneAccepting token ids as well as strings is unusual and handy for structured decoding.
stream / stream_optionsStreaming and usage reportingDefault false; {"include_usage": true}Usage is the only place you see cached-token counts, which decide your real input price.
toolsFunction tools the model can callArray, default noneFunction calling is documented on this interface. Note the compatibility page states tools cannot be combined with streaming.
enable_thinkingTurns the reasoning pass onBoolean. Disabled by default — Alibaba lists qwen3-max explicitly under "thinking disabled by default", unlike the Qwen3.7 Max/Plus series where it is onNot a standard OpenAI field: pass it via extra_body in the Python SDK, as a top-level field in Node, and top-level beside model in batch JSONL.
thinking_budgetCaps reasoning tokensInteger, via extra_body. Defaults to the model's maximum chain-of-thought length. Documented for Qwen3 models in thinking modeA true numeric budget, not a coarse level — the main advantage this model has over Kimi K3's three-step effort dial. The model stops reasoning once the limit is reached.
preserve_thinkingCarries prior reasoning into later turnsNot supported on qwen3-max. Alibaba's deep-thinking guide lists its supported models as Qwen3.7 Max/Plus, Qwen3.6 Max/Plus, Kimi-k2.7-code and Kimi-k2.6 — this model is not among themIf you need reasoning carried across turns, that is a reason to move to qwen3.7-max, not a knob you can set here.
cache_controlMarks an explicit cache breakpoint{"type": "ephemeral"}. qwen3-max is listed as supported for explicit cacheExplicit cache needs a 1,024-token minimum prefix; creation bills at 125% of the standard input price and hits at 10%. TTL is 5 minutes, resetting on each hit.
enable_searchLets the model use web search resultsBoolean, via extra_body. Could not be confirmed — it does not appear on the OpenAI-compatibility parameter reference, and no Model Studio page read lists qwen3-max as supporting itTreat as unverified. If you need grounded search on this interface, confirm against a current Alibaba page before designing around it.

The knobs that matter here are not the sampling ones. enable_thinking is first, because it is off by default on this model — if you assumed this was a reasoning model out of the box, it was not, and your quality numbers reflect that. thinking_budget is second: a numeric cap is a much better cost instrument than a three-level effort switch, and thinking content bills as output. Third is cache_control, because the explicit-cache economics are strong but asymmetric — a 125% write against a 10% read only pays back if you actually re-hit within five minutes. And note what is absent: no preserve_thinking, so every turn re-reasons from scratch and you pay for it at the tier your input length set.

SamplingShaping the output distribution

Qwen3-Max uses the standard Model Studio sampling surface with two details worth knowing. temperature is documented over [0, 2) and Alibaba explicitly says not to set it to 0 — the interval is open at the bottom, so "deterministic" here means a very small positive value rather than zero. top_p is documented over the open interval (0, 1.0). Alibaba's own advice is to leave both at their defaults and, if you must change something, change one of them and not the other; that is the correct instinct on any nucleus-plus-temperature stack, where the two interact multiplicatively and joint tuning makes results unattributable. top_k does not appear on the OpenAI-compatibility parameter page, so treat it as unavailable on this interface. seed is supported and is the better reproducibility lever than temperature-zeroing, but note that under the OpenAI-compatible protocol it must be within [0, 2^31 − 1], which is narrower than the unsigned 64-bit range the DashScope interface accepts — an easy source of silent 400s when porting. A sensible start is defaults everywhere, a fixed seed, and presence_penalty at 0.

ReasoningThinking, effort and budgets

Qwen3-Max is a hybrid-thinking model, and the important detail is that on the Qwen3 generation thinking is disabled by default — unlike Qwen3.5, 3.6 and 3.7, where it is on by default. You turn it on per request with enable_thinking: true. Because this is not a standard OpenAI parameter, it is passed through extra_body in the Python SDK, as a top-level parameter in the Node SDK, and as a top-level field beside model in batch JSONL request bodies — placing it inside extra_body in a batch file is a documented mistake. The budget is genuinely controllable via thinking_budget, an integer cap on reasoning tokens, and preserve_thinking passes reasoning forward into later turns. Reasoning is visible: it comes back in reasoning_content, separate from content, in both streaming and non-streaming responses. Billing is stated plainly, which is more than most vendors manage: "Thinking content is billed per output token", and if a hybrid model produces no reasoning, standard pricing applies.

ToolsFunction calling and server tools

Model Studio's capability table marks Qwen3-Max as supporting function calling, built-in tools and structured output. The parameter documentation read for this page covers tools but does not enumerate tool_choice, parallel_tool_calls or response_format on the OpenAI-compatible interface, so while the capabilities are asserted by the vendor, the exact field names and their defaults could not be confirmed here — do not copy them from another vendor's docs and assume they apply. One tool is documented directly: enable_search, a boolean passed via extra_body, which lets the model use web search results server-side. That is a real grounding capability and also a real evaluation hazard, since it makes responses non-reproducible and injects content you did not supply; keep it off in any pipeline where you need to attribute an answer to your own corpus. The practical failure mode across the Model Studio surface is parameter placement rather than tool semantics: non-standard fields live in extra_body in Python but top-level in Node and in batch files, and getting that wrong produces a request that is silently ignored rather than rejected.

CostPrice, caching, batching, what drives the bill

List price from Alibaba Cloud Model Studio's model pricing page, International (Singapore) deployment, USD per million tokens. Pricing is tiered by the number of input tokens in the request, and the tier applies to output as well:

Input lengthInput per 1MOutput per 1M
0 < tokens ≤ 32K$1.20$6.00
32K < tokens ≤ 128K$2.40$12.00
128K < tokens ≤ 256K$3.00$15.00

This tiering is the dominant cost fact and it is easy to miss: a prompt that creeps from 31K to 33K tokens doubles the price of the entire request, input and output alike. Prices vary by region — the table above is Singapore; Beijing, Hong Kong, Frankfurt, Tokyo and US Virginia are priced separately. A free quota of 1 million tokens is granted for 90 days after activation. Context caching cuts input further: implicit cache is automatic with roughly a 256-token minimum prefix and bills cached tokens at 20% of the input price for most models, while explicit cache — marked with cache_control: {"type": "ephemeral"} — needs a 1024-token prefix, bills creation at 125% of input and hits at 10%, with a 5-minute TTL that resets on each hit. Cached counts appear in usage.prompt_tokens_details. Thinking content bills as output tokens.

Where it runsSurfaces and availability

SurfaceAvailableNotes
Alibaba Cloud Model StudioYesModel id qwen3-max across six regions. OpenAI-compatible, Anthropic-compatible and DashScope interfaces.
Amazon BedrockNoBedrock's Qwen section lists Qwen3 235B A22B 2507, Qwen3 32B, Qwen3 Coder variants, Qwen3 Next 80B A3B and Qwen3 VL 235B — no Max model.
Google Vertex AINoVertex's managed-API list shows Qwen 3 Next Instruct 80B, Qwen 3 Next Thinking 80B, Qwen 3 Coder and Qwen 3 235B. Qwen3-Max is not listed.
Microsoft FoundryNoOnly Qwen-32B appears, as a fine-tuning target in public preview.
Hugging Face weightsNoNo official Qwen/Qwen3-Max repository. Search returns community re-uploads only.
Self-hostingNoCommercial API tier; weights not released.

StrengthsWhat it is good at

  • A numeric thinking_budget rather than a coarse effort level, which is a materially better cost instrument than the three-step dials on Kimi K3 and GLM-5.2.
  • Three interface flavours on one model — OpenAI-compatible, Anthropic-compatible and DashScope — so it drops into most existing clients without a rewrite.
  • Explicit context caching with published economics: 1024-token minimum, 125% write, 10% read, 5-minute TTL that resets on hit.
  • Six regions including Frankfurt and US Virginia, which gives non-China data residency options that DeepSeek's and Moonshot's own APIs do not.
  • Alibaba states plainly that thinking content bills per output token — a disclosure most vendors on this board omit.

LimitsWhere it falls down

  • Not open-weight, contrary to how it is frequently described. There is no official Qwen/Qwen3-Max repository; the Max line is Alibaba's commercial tier.
  • Architecture entirely undisclosed — no parameter count, expert count, active-parameter figure or attention description on any vendor page.
  • Max output tokens are not published anywhere this article could find, which makes capacity planning guesswork.
  • Input-length price tiering means a request crossing 32K or 128K tokens doubles or triples in cost, input and output together.
  • Superseded: it no longer appears on Model Studio's Supported Models overview, which now lists qwen3.7-max and qwen3.8-max-preview.

Against its neighboursHow it compares

Against Kimi K3, the other Frontier-tier Chinese model here, Qwen3-Max is cheaper at short prompts ($1.20 versus $3.00 input) and far cheaper on output ($6.00 versus $15.00), and it gives you a numeric thinking budget and real sampling controls where K3 gives you neither. K3 wins on context — 1M against 256K — and on native vision and video, which Qwen3-Max is not documented to accept. Against its own successor qwen3.7-max, the older snapshot is the stable, already-qualified option; 3.7 is what Alibaba now lists and prices as the current flagship, and staying on 3.6-era Max is a decision to accept a frozen baseline. Against GLM-5, Qwen3-Max is closed and Model-Studio-only where GLM-5 ships MIT weights and runs on Bedrock and Vertex.

Getting startedThe smallest call that works

code
POST https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions
Authorization: Bearer $DASHSCOPE_API_KEY
Content-Type: application/json

{
  "model": "qwen3-max",
  "messages": [
    {"role": "user", "content": "List three uses for a context cache."}
  ]
}

The base URL shown is the workspace-scoped form in the current Model Studio docs; substitute your workspace id and region. Change enable_thinking first — it is off by default on this generation, and turning it on is the single largest quality change available. Pass it via extra_body in Python, top-level in Node. Then set thinking_budget to cap what that costs you.

SourcesWhere every claim above came from

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page Alibaba launched Qwen3.8-Max, a 2.4T-parameter open-weight model that changes the self-host frontier calculus.

    Update sub/mdl-qwen3-max to account for Qwen3.8-Max, a 2.4-trillion-parameter open-weight model that challenges Qwen3 Max's closed-weight description.

    China frontier labs · 3 Aug 2026 · source

RecentAuto-linked from the brief, not a rewrite of this page

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning