Alibaba's closed-weight 256K flagship snapshot, priced in three tiers by input length, with thinking off by default — not an open-weight model.
Why this oneWhat it is actually for
You reach for Qwen3-Max when you are already inside Alibaba Cloud Model Studio, want the Qwen family's multilingual behaviour, and specifically want the January 2026 snapshot rather than the newer qwen3.7-max that has replaced it on the model overview page. Pinning to qwen3-max-2026-01-23 is a legitimate reason to be here: it is a stable, still-billable model with a fixed evaluation baseline, and moving to 3.7 means re-qualifying every prompt. It is also the cheapest Qwen Max tier if your prompts stay under 32K, at $1.20 per million input. You reach elsewhere if you need open weights — despite what a lot of secondary sources imply, the Max line is not open-weighted — or a context window past 256K.
What it isArchitecture, lineage, training
Qwen3-Max is Alibaba's proprietary flagship in the Qwen3 generation. The model id qwen3-max currently resolves to the snapshot qwen3-max-2026-01-23, and a qwen3-max-preview snapshot also exists. Alibaba's text-generation model listing gives it a 256K context window and marks thinking mode, function calling, built-in tools and structured output all as supported.
What Alibaba does not disclose is the architecture. No parameter count, no expert count, no active-parameter figure, no layer count, no attention description appears on any Alibaba Cloud Model Studio page read for this article. The site's stored row said "MoE" — that is a reasonable inference from the family, but it is not a vendor-stated fact and is not asserted here.
The open-weight question deserves to be settled plainly, because the roster row got it wrong. There is no Qwen3-Max repository under the official Qwen Hugging Face organisation; a search returns only community re-uploads and derivatives from unaffiliated accounts. Alibaba open-weights a large part of the Qwen family — the Qwen3, Qwen3.5 and Qwen3.6 dense and MoE releases under Apache 2.0 — but the Max line is the commercial, API-only tier. You cannot download Qwen3-Max, cannot fine-tune it on your own hardware, and cannot self-host it. Treat "API / open-weight" on this row as incorrect.
One more thing worth knowing before you build on it: qwen3-max no longer appears on Model Studio's Supported Models overview page, which now lists qwen3.7-max and qwen3.8-max-preview. It is still fully priced and served across Beijing, Hong Kong, Singapore, Tokyo, Frankfurt and US Virginia, and carries no deprecation notice — but it is a superseded generation.
At a glanceSee it
Qwen3-Max prices by input length: three tiers, and the whole bill doubles or triples as the prompt grows.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| Model id | qwen3-max, "currently equivalent to qwen3-max-2026-01-23" | Model Studio model pricing |
| Deprecation | qwen3-max-2026-01-23 is scheduled for deprecation on October 10, 2026, 00:00:00, with qwen3.7-max named as the replacement model. The unversioned qwen3-max alias is not itself listed as deprecated | Model Studio model deprecation |
| Context window | 256K. Supported indirectly: the model's top pricing tier is bounded at "128K < Token ≤ 256K". qwen3-max no longer appears in the Model Studio supported-models list, so no separately stated context figure could be located | Model Studio model pricing |
| Max output tokens | Could not be confirmed from a primary source as of 2026-07-26 | Not stated on any Model Studio page read |
| Input modalities | Text. The pricing page describes it in the Qwen-Max text series; vision input could not be confirmed | Model Studio model pricing |
| Output modalities | Text, plus reasoning_content when thinking is enabled. The pricing page lists "Non-Thinking and Thinking modes" | Model Studio model pricing; deep thinking guide |
| Architecture | Not disclosed — no parameter count, expert count or layer count published | — |
| Knowledge cutoff | Not disclosed | — |
| Weights available | No. No Qwen3-Max repository under the official Qwen Hugging Face organisation as of 2026-07-26 | Hugging Face model search |
| Licence | Not applicable — commercial API only, no weights released | — |
| Regions | China Beijing, Hong Kong, Singapore, Tokyo, Frankfurt, US Virginia | Model Studio model pricing |
| Free quota | 1 million tokens, valid 90 days after activation; International deployment only | Model Studio model pricing |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
messages | Conversation history | Required | Total input length decides which of three price tiers the whole request bills at, so trimming history has an outsized effect here. |
max_tokens | Caps generated tokens | Integer; ceiling not documented | Because the model's output ceiling is undocumented, set this explicitly rather than relying on a default you cannot look up. |
temperature | Sampling randomness | Range [0, 2). Alibaba says it cannot be set to 0 | The floor is exclusive — for near-deterministic output use a very small positive value, not 0. |
top_p | Nucleus sampling mass | Range (0, 1.0) | Alibaba's guidance is to keep defaults and adjust only one of temperature or top_p, never both. |
top_k, frequency_penalty | Top-k truncation and repetition penalty | Not documented on the OpenAI-compatibility parameter page | Do not assume they work. Use presence_penalty for repetition and top_p for truncation. |
presence_penalty | Discourages repeated tokens | Range [-2.0, 2.0] | The only repetition control documented on this interface. Small positive values loosen looping without derailing structure. |
n | Number of responses | Default 1, range 1–4 | Every candidate bills separately and at the tier your input length set. |
seed | Random seed | Integer. The OpenAI-compatibility reference states it "supports unsigned 64-bit integers" | Improves reproducibility across runs. There is no documented 32-bit restriction on the OpenAI-compatible protocol. |
stop | Halt sequences | String or array; strings or token ids; default none | Accepting token ids as well as strings is unusual and handy for structured decoding. |
stream / stream_options | Streaming and usage reporting | Default false; {"include_usage": true} | Usage is the only place you see cached-token counts, which decide your real input price. |
tools | Function tools the model can call | Array, default none | Function calling is documented on this interface. Note the compatibility page states tools cannot be combined with streaming. |
enable_thinking | Turns the reasoning pass on | Boolean. Disabled by default — Alibaba lists qwen3-max explicitly under "thinking disabled by default", unlike the Qwen3.7 Max/Plus series where it is on | Not a standard OpenAI field: pass it via extra_body in the Python SDK, as a top-level field in Node, and top-level beside model in batch JSONL. |
thinking_budget | Caps reasoning tokens | Integer, via extra_body. Defaults to the model's maximum chain-of-thought length. Documented for Qwen3 models in thinking mode | A true numeric budget, not a coarse level — the main advantage this model has over Kimi K3's three-step effort dial. The model stops reasoning once the limit is reached. |
preserve_thinking | Carries prior reasoning into later turns | Not supported on qwen3-max. Alibaba's deep-thinking guide lists its supported models as Qwen3.7 Max/Plus, Qwen3.6 Max/Plus, Kimi-k2.7-code and Kimi-k2.6 — this model is not among them | If you need reasoning carried across turns, that is a reason to move to qwen3.7-max, not a knob you can set here. |
cache_control | Marks an explicit cache breakpoint | {"type": "ephemeral"}. qwen3-max is listed as supported for explicit cache | Explicit cache needs a 1,024-token minimum prefix; creation bills at 125% of the standard input price and hits at 10%. TTL is 5 minutes, resetting on each hit. |
enable_search | Lets the model use web search results | Boolean, via extra_body. Could not be confirmed — it does not appear on the OpenAI-compatibility parameter reference, and no Model Studio page read lists qwen3-max as supporting it | Treat as unverified. If you need grounded search on this interface, confirm against a current Alibaba page before designing around it. |
The knobs that matter here are not the sampling ones. enable_thinking is first, because it is off by default on this model — if you assumed this was a reasoning model out of the box, it was not, and your quality numbers reflect that. thinking_budget is second: a numeric cap is a much better cost instrument than a three-level effort switch, and thinking content bills as output. Third is cache_control, because the explicit-cache economics are strong but asymmetric — a 125% write against a 10% read only pays back if you actually re-hit within five minutes. And note what is absent: no preserve_thinking, so every turn re-reasons from scratch and you pay for it at the tier your input length set.
SamplingShaping the output distribution
Qwen3-Max uses the standard Model Studio sampling surface with two details worth knowing. temperature is documented over [0, 2) and Alibaba explicitly says not to set it to 0 — the interval is open at the bottom, so "deterministic" here means a very small positive value rather than zero. top_p is documented over the open interval (0, 1.0). Alibaba's own advice is to leave both at their defaults and, if you must change something, change one of them and not the other; that is the correct instinct on any nucleus-plus-temperature stack, where the two interact multiplicatively and joint tuning makes results unattributable. top_k does not appear on the OpenAI-compatibility parameter page, so treat it as unavailable on this interface. seed is supported and is the better reproducibility lever than temperature-zeroing, but note that under the OpenAI-compatible protocol it must be within [0, 2^31 − 1], which is narrower than the unsigned 64-bit range the DashScope interface accepts — an easy source of silent 400s when porting. A sensible start is defaults everywhere, a fixed seed, and presence_penalty at 0.
ReasoningThinking, effort and budgets
Qwen3-Max is a hybrid-thinking model, and the important detail is that on the Qwen3 generation thinking is disabled by default — unlike Qwen3.5, 3.6 and 3.7, where it is on by default. You turn it on per request with enable_thinking: true. Because this is not a standard OpenAI parameter, it is passed through extra_body in the Python SDK, as a top-level parameter in the Node SDK, and as a top-level field beside model in batch JSONL request bodies — placing it inside extra_body in a batch file is a documented mistake. The budget is genuinely controllable via thinking_budget, an integer cap on reasoning tokens, and preserve_thinking passes reasoning forward into later turns. Reasoning is visible: it comes back in reasoning_content, separate from content, in both streaming and non-streaming responses. Billing is stated plainly, which is more than most vendors manage: "Thinking content is billed per output token", and if a hybrid model produces no reasoning, standard pricing applies.
ToolsFunction calling and server tools
Model Studio's capability table marks Qwen3-Max as supporting function calling, built-in tools and structured output. The parameter documentation read for this page covers tools but does not enumerate tool_choice, parallel_tool_calls or response_format on the OpenAI-compatible interface, so while the capabilities are asserted by the vendor, the exact field names and their defaults could not be confirmed here — do not copy them from another vendor's docs and assume they apply. One tool is documented directly: enable_search, a boolean passed via extra_body, which lets the model use web search results server-side. That is a real grounding capability and also a real evaluation hazard, since it makes responses non-reproducible and injects content you did not supply; keep it off in any pipeline where you need to attribute an answer to your own corpus. The practical failure mode across the Model Studio surface is parameter placement rather than tool semantics: non-standard fields live in extra_body in Python but top-level in Node and in batch files, and getting that wrong produces a request that is silently ignored rather than rejected.
CostPrice, caching, batching, what drives the bill
List price from Alibaba Cloud Model Studio's model pricing page, International (Singapore) deployment, USD per million tokens. Pricing is tiered by the number of input tokens in the request, and the tier applies to output as well:
| Input length | Input per 1M | Output per 1M |
|---|---|---|
| 0 < tokens ≤ 32K | $1.20 | $6.00 |
| 32K < tokens ≤ 128K | $2.40 | $12.00 |
| 128K < tokens ≤ 256K | $3.00 | $15.00 |
This tiering is the dominant cost fact and it is easy to miss: a prompt that creeps from 31K to 33K tokens doubles the price of the entire request, input and output alike. Prices vary by region — the table above is Singapore; Beijing, Hong Kong, Frankfurt, Tokyo and US Virginia are priced separately. A free quota of 1 million tokens is granted for 90 days after activation. Context caching cuts input further: implicit cache is automatic with roughly a 256-token minimum prefix and bills cached tokens at 20% of the input price for most models, while explicit cache — marked with cache_control: {"type": "ephemeral"} — needs a 1024-token prefix, bills creation at 125% of input and hits at 10%, with a 5-minute TTL that resets on each hit. Cached counts appear in usage.prompt_tokens_details. Thinking content bills as output tokens.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Alibaba Cloud Model Studio | Yes | Model id qwen3-max across six regions. OpenAI-compatible, Anthropic-compatible and DashScope interfaces. |
| Amazon Bedrock | No | Bedrock's Qwen section lists Qwen3 235B A22B 2507, Qwen3 32B, Qwen3 Coder variants, Qwen3 Next 80B A3B and Qwen3 VL 235B — no Max model. |
| Google Vertex AI | No | Vertex's managed-API list shows Qwen 3 Next Instruct 80B, Qwen 3 Next Thinking 80B, Qwen 3 Coder and Qwen 3 235B. Qwen3-Max is not listed. |
| Microsoft Foundry | No | Only Qwen-32B appears, as a fine-tuning target in public preview. |
| Hugging Face weights | No | No official Qwen/Qwen3-Max repository. Search returns community re-uploads only. |
| Self-hosting | No | Commercial API tier; weights not released. |
StrengthsWhat it is good at
- A numeric
thinking_budgetrather than a coarse effort level, which is a materially better cost instrument than the three-step dials on Kimi K3 and GLM-5.2. - Three interface flavours on one model — OpenAI-compatible, Anthropic-compatible and DashScope — so it drops into most existing clients without a rewrite.
- Explicit context caching with published economics: 1024-token minimum, 125% write, 10% read, 5-minute TTL that resets on hit.
- Six regions including Frankfurt and US Virginia, which gives non-China data residency options that DeepSeek's and Moonshot's own APIs do not.
- Alibaba states plainly that thinking content bills per output token — a disclosure most vendors on this board omit.
LimitsWhere it falls down
- Not open-weight, contrary to how it is frequently described. There is no official
Qwen/Qwen3-Maxrepository; the Max line is Alibaba's commercial tier. - Architecture entirely undisclosed — no parameter count, expert count, active-parameter figure or attention description on any vendor page.
- Max output tokens are not published anywhere this article could find, which makes capacity planning guesswork.
- Input-length price tiering means a request crossing 32K or 128K tokens doubles or triples in cost, input and output together.
- Superseded: it no longer appears on Model Studio's Supported Models overview, which now lists
qwen3.7-maxandqwen3.8-max-preview.
Against its neighboursHow it compares
Against Kimi K3, the other Frontier-tier Chinese model here, Qwen3-Max is cheaper at short prompts ($1.20 versus $3.00 input) and far cheaper on output ($6.00 versus $15.00), and it gives you a numeric thinking budget and real sampling controls where K3 gives you neither. K3 wins on context — 1M against 256K — and on native vision and video, which Qwen3-Max is not documented to accept. Against its own successor qwen3.7-max, the older snapshot is the stable, already-qualified option; 3.7 is what Alibaba now lists and prices as the current flagship, and staying on 3.6-era Max is a decision to accept a frozen baseline. Against GLM-5, Qwen3-Max is closed and Model-Studio-only where GLM-5 ships MIT weights and runs on Bedrock and Vertex.
Getting startedThe smallest call that works
POST https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions
Authorization: Bearer $DASHSCOPE_API_KEY
Content-Type: application/json
{
"model": "qwen3-max",
"messages": [
{"role": "user", "content": "List three uses for a context cache."}
]
}The base URL shown is the workspace-scoped form in the current Model Studio docs; substitute your workspace id and region. Change enable_thinking first — it is off by default on this generation, and turning it on is the single largest quality change available. Pass it via extra_body in Python, top-level in Node. Then set thinking_budget to cap what that costs you.
SourcesWhere every claim above came from
- Alibaba Cloud Model Studio model pricing
- Text generation models — Qwen, DeepSeek, GLM — Model Studio
- Supported Models and Capabilities Overview — Model Studio
- Use deep thinking models via API — Model Studio
- Qwen Context Cache feature — Model Studio
- OpenAI compatibility — Model Studio
- Hugging Face model search for Qwen3-Max
- Models at a glance — Amazon Bedrock
- Vertex AI managed models for MaaS — Google Cloud
- Could not confirm from a primary source as of 2026-07-25: max output tokens, any architecture detail including parameter and expert counts, knowledge cutoff, vision input support, and the exact field names for
tool_choice,parallel_tool_callsandresponse_formaton the OpenAI-compatible interface. Corrections to the site's stored row: it listed this model as "API / open-weight" with fine-tuning "Yes (open)" — no official Qwen3-Max weights exist, so both are wrong; it listed modalities as "Text, vision" — vision could not be confirmed and the model is listed under text generation; and it listed the architecture as "MoE", which Alibaba does not state anywhere.
What changedWhat changed here
Updated this page Alibaba launched Qwen3.8-Max, a 2.4T-parameter open-weight model that changes the self-host frontier calculus.
Update sub/mdl-qwen3-max to account for Qwen3.8-Max, a 2.4-trillion-parameter open-weight model that challenges Qwen3 Max's closed-weight description.
- Alibaba Unveils Qwen3.8-Max, Its Biggest AI Model Yet
Alibaba unveiled Qwen3.8-Max, its largest model to date, extending the Qwen line that many builders already use as a cheap open-weight alternative to US frontier models. If you're routing workloads by cost, a new top-end Qwen changes the price/quality frontier you should be re-testing against.
- Qwen3.8 Max Debuts With 2.4 Trillion Parameters, Trailing Only Fable 5
Alibaba's Qwen3.8 Max debuted with 2.4 trillion parameters, trailing only Fable 5 in scale. For open-weight deployments, this is a new high-end option to evaluate against your existing Qwen models, especially when you want frontier-adjacent capability you can self-host.
- Qwen3.8 Max Debuts With 2.4 Trillion Parameters, Trailing Only Fable 5
Alibaba's Qwen3.8 Max debuts with 2.4 trillion parameters, trailing only Fable 5 in the reported frontier ranking. That's another near-top-tier model to weigh when selecting vendors or planning multi-model strategies.
- Qwen3.8 Max Debuts With 2.4 Trillion Parameters, Trailing Only Fable 5
Alibaba's Qwen3.8-Max debuts as a 2.4-trillion-parameter open-weight model that the report ranks second only to Fable 5. The scale signals that the open-weight frontier now reaches trillion-parameter territory, which matters if you're planning large self-hosted deployments.
- Qwen3.8 Max Debuts With 2.4 Trillion Parameters, Trailing Only Fable 5
Alibaba's Qwen3.8 Max debuts with 2.4 trillion parameters, trailing only Fable 5 on the benchmarks. That scale pushes open-weight in-house deployments further, but you'll need serious infrastructure to serve it.
- Alibaba Launches 2.4 Trillion-Parameter Qwen3.8-Max AI Model, Challenging Moonshot and OpenAI
Alibaba launched Qwen3.8-Max, a 2.4-trillion-parameter open-weight model it says challenges Moonshot and OpenAI. That changes what you can self-host at frontier-ish quality, so re-run your eval suite against it before locking your model strategy.
- China’s Alibaba takes another swipe at America’s AI supremacy
Alibaba released its largest and "most capable AI model to date," Qwen3.8-Max, claiming parity with Anthropic, OpenAI, and Kimi K3, and made it widely available — if you're already building on Qwen, this raises the ceiling without switching stacks.
Three kinds of claim, strongest first. Signal runs every morning.