A 1.6T-parameter MoE with 49B active, MIT-licensed weights, a 1M context and a 384K output ceiling, at $1.32 in and $3.96 out at peak, and half that off-peak.
Why this oneWhat it is actually for
You reach for V4 Pro when you want DeepSeek's reasoning quality rather than its price floor, and you want it without giving up the ability to run the same weights yourself later. It is roughly three times the input price of V4 Flash and three times the output price, which buys 1.6T total parameters against Flash's 284B and 49B active against 13B — a real gap on world knowledge and multi-step agentic work. The 384K max output is the other reason: it is one of very few models that will emit a third of a million tokens in one response, which matters for whole-file rewrites and long plan generation. If your workload is high-volume extraction or summarisation, Flash is the correct call and Pro is waste.
What it isArchitecture, lineage, training
DeepSeek V4 Pro is a sparse mixture-of-experts model released 2026-04-24. DeepSeek's release note states 1.6T total parameters with 49B active per token, and the Hugging Face card describes a "hybrid attention mechanism combining Compressed Sparse Attention and Heavily Compressed Attention", plus Manifold-Constrained Hyper-Connections and the Muon optimizer. The card claims the design needs only 27% of the single-token inference FLOPs and 10% of the KV cache of DeepSeek-V3.2.
The real numbers come from config.json in deepseek-ai/DeepSeek-V4-Pro. hidden_size is 7168 across num_hidden_layers 61. num_attention_heads is 128 with num_key_value_heads of 1 and head_dim 512 — a latent-attention layout, with q_lora_rank 1536 and o_lora_rank 1024. The MoE block has n_routed_experts 384 plus n_shared_experts 1, with num_experts_per_tok 6 and moe_intermediate_size 3072; routing uses topk_method noaux_tc and a sqrtsoftplus scoring function. vocab_size is 129280 and max_position_embeddings is 1048576, reached by YaRN scaling with factor 16 over an original_max_position_embeddings of 65536; rope_theta is 10000 with a separate compress_rope_theta of 160000. The sparse-attention indexer is visible as index_n_heads 64, index_head_dim 128 and index_topk 1024, alongside a sliding_window of 128. Weights ship FP8 e4m3 with 128x128 blocks and expert_dtype fp4, and there is one num_nextn_predict_layers for multi-token prediction.
Training, per the model card: over 32T pretraining tokens, then two-stage post-training — domain-specific expert cultivation followed by unified consolidation via on-policy distillation. Licence is MIT.
At a glanceSee it
Per-token path through DeepSeek V4 Pro: sparse attention, then six of 384 routed experts plus one shared.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| Model id | deepseek-v4-pro | DeepSeek Models & Pricing |
| Context window | 1M tokens; max_position_embeddings is exactly 1048576 | Models & Pricing; config.json |
| Max output tokens | 384K | DeepSeek Models & Pricing |
| Input modalities | Text | Model card; Microsoft Foundry lists input as text, languages en and zh |
| Output modalities | Text, plus reasoning_content when thinking is on | DeepSeek API reference |
| Total / active parameters | 1.6T total, 49B active | DeepSeek V4 release note |
| Pretraining tokens | Over 32T | Hugging Face model card |
| Knowledge cutoff | Not disclosed | Not stated on any DeepSeek page read |
| Weights available | Yes — deepseek-ai/DeepSeek-V4-Pro, plus a -Base variant | Hugging Face |
| Licence | MIT | Hugging Face model card |
| Concurrency limit | 500 | DeepSeek Models & Pricing |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
messages | Conversation so far | Required | Prior turns and tool schemas all bill as input; caching only helps the unchanged prefix. |
max_tokens | Caps the completion | Integer; ceiling 384K per the Models & Pricing page. The API reference documents no explicit range | Set it deliberately. Too low is the documented cause of truncated JSON. |
temperature | Sampling randomness | Range 0 to 2, default 1 | DeepSeek's own guidance table: 0.0 for coding and maths, 1.0 for data cleaning and analysis, 1.3 for general conversation and translation, 1.5 for creative writing and poetry. |
top_p | Nucleus sampling mass | Range 0 to 1, default 1 | The V4 model card recommends leaving it at 1.0. Move temperature or top_p, not both. |
thinking | Turns the reasoning pass on or off | Object. {"type":"enabled"} / {"type":"disabled"}; default enabled | Thinking is on unless you turn it off. Disabling drops the deliberation and the reasoning tokens you pay for. Tool calls are supported with thinking enabled. |
reasoning_effort | How hard it thinks | Top-level field, sent alongside thinking, not nested inside it. Effective values high or max; default high. low and medium are accepted but remapped to high; xhigh is remapped to max | The remapping is the trap: passing low to save money buys you nothing — you get high and pay for it. The docs also note effort is automatically raised to max for some complex agent requests. The model card recommends a context window of at least 384K for Think Max. |
tools | Function tools | Array, maximum 128 | Beyond a couple of dozen tools, selection accuracy degrades before the cap bites. |
tool_choice | Forces or forbids tool use | none, auto, required, or a named function | required forces a call even when none is appropriate. |
response_format | Output shape | {"type":"text"} (default) or {"type":"json_object"} | JSON mode also requires the word "json" in the prompt plus a format example, or output is unreliable. |
stop | Halt sequences | String or array, up to 16 sequences | Sixteen is a generous cap; useful for agent frame delimiters. |
stream | Server-sent events | Boolean | Near-mandatory with thinking on, since first useful token can be far out. |
stream_options | Usage in the stream | {"include_usage": true} | The only way to read prompt_cache_hit_tokens and prompt_cache_miss_tokens per call. |
logprobs | Log probabilities of output tokens | Boolean | Enables confidence gating on extraction tasks. |
top_logprobs | Alternatives per position | Integer, 0 to 20 | Requires logprobs; grows response size, not token cost. |
user_id | End-user identifier for safety review | String, ≤ 512 chars, charset [a-zA-Z0-9\-_] | No effect on generation. Note the restricted charset — email addresses and raw UUIDs with other punctuation will be rejected. |
frequency_penalty, presence_penalty | Repetition penalties | Not supported — the API reference states "This parameter is no longer supported. It will not take effect if you pass it to the API" | They fail silently rather than erroring. Handle repetition with stop sequences and prompt instructions instead. |
Three knobs carry the weight. thinking is the biggest: it defaults to enabled, so turning it off is an active decision that collapses latency and output tokens, and for extraction or classification that costs you almost nothing in quality. reasoning_effort is the second, and the important fact is that it is a two-position switch wearing a five-position label — only high (the default) and max do anything distinct, and the cheaper-sounding values are remapped upward. Third is temperature, which unusually has vendor-published per-task values: 0.0 for code and maths is a real recommendation from DeepSeek, not folklore, even though the model card's generic deployment advice is 1.0.
SamplingShaping the output distribution
DeepSeek gives you a genuine temperature dial and, unusually, publishes what to set it to. The parameter-settings page lists 0.0 for coding and maths, 1.0 for data cleaning and analysis, 1.3 for general conversation and translation, and 1.5 for creative writing, against a stated default of 1.0. The V4 model card, which is aimed at people running the weights themselves, recommends temperature = 1.0 and top_p = 1.0 as the standard deployment setting. Those two pieces of advice are not in conflict so much as aimed at different callers: the card describes the operating point the model was tuned at, the API page describes how to bend it per task. Both temperature and top_p are accepted, and the usual discipline applies — move one, not both, or you will not know which change caused what. A sensible starting point for agentic and coding work is temperature 0.0 with top_p left at 1.0, then raise temperature only if the model is getting stuck in a rut. Note there is no top_k and no repetition penalty: both penalty parameters were removed.
ReasoningThinking, effort and budgets
Thinking mode is supported and, per DeepSeek's own pricing page, is the default for both V4 models. It is controlled by a thinking object taking {"type":"enabled"} or {"type":"disabled"}, with a separate top-level reasoning_effort string accepting high or max. The model card refers to the top setting as "Think Max" and states it needs a context window of at least 384K tokens — that is a hard operational constraint, not a suggestion, and it is the reason the max output figure and the Think Max requirement are the same number. Reasoning is surfaced as reasoning_content separate from content. Tool calling works while thinking is enabled; the docs note this has been true since V3.2. What the pricing page does not say is whether thinking tokens are billed at a different rate from ordinary output tokens — it gives one output price and no breakdown. Budget as if all reasoning tokens bill at the $3.96 peak output rate, or $1.98 off-peak.
ToolsFunction calling and server tools
Tool calling is supported, works alongside thinking mode, and takes up to 128 tools. tool_choice accepts none, auto, required or a specific function. The docs do not state whether multiple tool calls come back in one response, so do not architect around parallel dispatch without testing it. For structured arguments there is a strict mode: point the client at the beta base URL https://api.deepseek.com/beta and set "strict": true on the function. Strict mode supports object, string, number, integer, boolean, array, enum, anyOf, $ref and $def, but rejects minLength, maxLength, minItems and maxItems, and requires every property to be listed in required with additionalProperties: false. Two failure modes to plan for. JSON mode via response_format is documented as occasionally returning empty content, which DeepSeek says it is still optimising — retry rather than assume a parse bug. And separately, if you reach this model through Microsoft Foundry rather than DeepSeek's API, Foundry's own model table lists tool calling as not supported for DeepSeek-V4-Pro.
CostPrice, caching, batching, what drives the bill
List price from DeepSeek's Models & Pricing page, USD per million tokens. DeepSeek prices this model by time of day: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, and off-peak rates — every other hour, the whole weekend included — are exactly half the peak rates.
| Item | Off-peak | Peak |
|---|---|---|
| Input, cache miss | $0.66 | $1.32 |
| Input, cache hit | $0.022 | $0.044 |
| Output | $1.98 | $3.96 |
The cache-hit rate is 30x below the miss rate on either schedule, which is the single most important economic fact about this model. Caching is "enabled by default for all users, without needing to modify their code" and requires no parameter; the API reports prompt_cache_hit_tokens and prompt_cache_miss_tokens in the usage object. Caches are cleared automatically after "a few hours to a few days" of disuse. No cache-write fee is documented. Note what DeepSeek does not publish: there is no documented minimum cache unit size, prefix granularity or cache-construction rule, so how partial prefix overlap is treated could not be confirmed — measure it with the two usage counters rather than assuming a boundary. No batch discount is published. The time-of-day split is the discount, and it is the cheapest lever on this page: a job you can defer out of the two peak windows costs half as much for no code change at all. Plan the clock before you plan the prompt.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| DeepSeek API | Yes | https://api.deepseek.com, model id deepseek-v4-pro. OpenAI ChatCompletions and Anthropic-compatible interfaces both supported. |
| Microsoft Foundry | Yes | Sold directly by Azure as DeepSeek-V4-Pro: 1,000,000 input / 384,000 output tokens, languages en and zh. Foundry's table lists tool calling as not supported. |
| Alibaba Cloud Model Studio | Yes | Listed among the third-party models with hybrid thinking support. |
| Amazon Bedrock | No | Bedrock's DeepSeek section lists V3.2, V3.1 and R1 only. |
| Google Vertex AI | No | Vertex's managed-API list shows DeepSeek-V3.2, V3.1, R1-0528 and OCR; V4 does not appear. |
| Hugging Face weights | Yes | deepseek-ai/DeepSeek-V4-Pro and DeepSeek-V4-Pro-Base, MIT licence, FP8 with FP4 experts. |
| Self-hosting | Yes | Weights are public. Serving a 1.6T-parameter model is a serious hardware commitment even at FP8/FP4. |
StrengthsWhat it is good at
- MIT-licensed weights for a genuine frontier-scale model — 1.6T total, 49B active — with a
-Basevariant published alongside the instruct model. - 384K max output tokens, which is an order of magnitude beyond most models and makes whole-repository rewrites feasible in one response.
- Cache-hit input at $0.044 per million against a $1.32 miss at peak — a 30x gap — applied automatically, with per-request hit/miss counts exposed in
usage. - Vendor-published per-task temperature guidance — 0.0 for code and maths, 1.5 for creative — rather than leaving callers to guess.
- Strict JSON-schema tool arguments available on the beta base URL, with the supported and unsupported schema keywords enumerated.
LimitsWhere it falls down
- Text only. No image or video input at all, which rules it out for document-vision and UI-agent work that Kimi K3 and the Qwen Max line handle.
frequency_penaltyandpresence_penaltywere removed from the API, so repetition has to be managed by prompt andstopsequences.- JSON mode is documented as occasionally returning empty content — the vendor calls it a known issue under optimisation, so you need retry logic.
- Not on Bedrock or Vertex. Microsoft Foundry carries it, but with tool calling listed as unsupported on that surface.
- Reasoning-effort
maxrequires a 384K context window to operate, which constrains how you can deploy it self-hosted.
Against its neighboursHow it compares
Against DeepSeek V4 Flash, the sibling on the same endpoint: Pro costs exactly 3x more on input and 3x more on output on the same schedule, for 1.6T total against 284B and 49B active against 13B. Same 1M context, same 384K output ceiling, same parameters, same MIT licence. The upgrade buys world knowledge and agentic reliability; DeepSeek's own release note says Flash's reasoning "closely approaches" Pro when given a larger thinking budget, which is a strong hint that for reasoning-heavy but knowledge-light work, Flash plus more thinking is the better trade. Against Kimi K3, Pro is roughly 3.8x cheaper on output at peak and 7.6x off-peak and gives you sampling controls and downloadable weights, but has no vision and a smaller reported coding profile. Against Claude Sonnet 5, Pro is the self-hostable option; Sonnet is the one with broad cloud distribution.
Getting startedThe smallest call that works
POST https://api.deepseek.com/chat/completions
Authorization: Bearer $DEEPSEEK_API_KEY
Content-Type: application/json
{
"model": "deepseek-v4-pro",
"messages": [
{"role": "user", "content": "Explain what a KV cache is, in two sentences."}
]
}Change temperature first: set it to 0.0 for code and maths, per DeepSeek's own guidance table. Then decide on thinking — send {"type":"disabled"} for extraction and classification and watch the bill drop. Add "stream": true with "stream_options": {"include_usage": true} to see your cache hit rate. Leave top_p at 1.0.
SourcesWhere every claim above came from
- Models & Pricing — DeepSeek API Docs
- DeepSeek V4 Preview Release — DeepSeek API Docs
- Change Log — DeepSeek API Docs
- Create Chat Completion — DeepSeek API Docs
- Temperature Settings — DeepSeek API Docs
- Context Caching — DeepSeek API Docs
- Function Calling — DeepSeek API Docs
- JSON Output — DeepSeek API Docs
- deepseek-ai/DeepSeek-V4-Pro — Hugging Face model card
- deepseek-ai/DeepSeek-V4-Pro config.json
- Foundry Models sold by Azure — Microsoft Foundry
- Models at a glance — Amazon Bedrock
- Could not confirm from a primary source as of 2026-07-25: the knowledge cutoff, whether thinking tokens bill at a rate different from ordinary output tokens, and whether multiple tool calls are returned in a single response. Pricing re-read from the vendor page on 2026-08-19: this model is now on a documented peak/off-peak schedule and the figures above match the vendor table exactly. No batch discount is published.
Price and capacity verified 2026-09-12 against https://api-docs.deepseek.com/quick_start/pricing/. re-read 2026-09-12 (research pass, primary source read today: https://api-docs.deepseek.com/quick_start/pricing/ — figures confirmed: input_per_m 1.32, output_per_m 3.96, cache_hit_input_per_m 0.044, context 1M, max_output 384K.) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $1.32 in / $3.96 out per M, $0.044 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-09-04 (B2; the CITATION was corrected, not the price. The cited URL lacked a trailing slash and served "Your First API Call" with ZERO price-like numbers on it — a citation nothing could re-read, on the model 304 of 305 kits run on. With the slash it serves "Models & Pricing", and a live read that day found ALL TEN figures the two DeepSeek rows carry — peak, off-peak and cache-hit — present and unchanged: 0.007 0.014 0.022 0.044 0.22 0.44 0.66 1.32 1.98 3.96. No price moved. Prior: manual 2026-08-23 (currency proposal src-2026-08-23-8eb120; every peak and off-peak figure unchanged; the peak window carries a 'Monday through Friday' qualifier this row did not record).
What changedWhat changed here
Updated this page DeepSeek now serves all V4-Pro API traffic with V4.1-Flash and bills at Flash rates, so the V4-Pro page's pricing no longer describes what a caller pays.
Update the DeepSeek V4 Pro page to state that its API traffic is now routed to V4.1-Flash at Flash rates, so the listed $1.32/$3.96 peak pricing no longer applies to live calls.
Updated this page A 552B-parameter open-weight MoE model, DeepSeek-V4.1-Flash, is available on Alibaba Cloud at Flash pricing.
Add DeepSeek-V4.1-Flash to the DeepSeek V4 Flash page as a 552B-parameter MoE successor available on Alibaba Cloud, and note what it means for the self-host-versus-API decision.
- 'Better than DeepSeek': Xiaomi's MiMo-V2.6-Pro debuts as the top open weights model in the world alongside cheaper V2.6-Flash
Xiaomi's MiMo-V2.6-Pro debuts as the top open-weights model in the world, with a cheaper V2.6-Flash variant alongside it. A new open-weights leader changes what you can self-host or fine-tune without paying frontier API rates, so it's worth re-benchmarking your model choice.
- DeepSeek-V4.1-Flash Ships Causal Encoder–Decoder MoE With 1M Context and Extreme KV Compression
DeepSeek shipped V4.1 Flash, a causal encoder–decoder MoE with a 1M-token context window and aggressive KV-cache compression. If you're architecting long-document or long-session agents, this is a cheap way to stop chunking and retrieval gymnastics for context that now fits in one call.
- DeepSeek-V4.1-Flash Outpaces GPT-5.6 Sol on Some Agentic, Coding Tests
DeepSeek-V4.1-Flash reportedly beats GPT-5.6 Sol on some agentic and coding tests, which is the kind of result that should make you re-test your model routing rather than assume the US frontier model wins on your task. Run your own eval before locking in a provider.
- DeepSeek Launches V4.1-Flash With Lower Memory and API Costs
DeepSeek launched V4.1-Flash with lower memory footprint and lower API costs, and separate coverage has it outpacing GPT-5.6 Sol on some agentic and coding tests. For cost-sensitive agent or coding workloads, this is a cheaper option worth benchmarking against your current model before assuming the frontier US models are the only viable choice.
- DeepSeek Unifies Three Modes: V4.1 Flash Launches Today at Fastest Speed, V4 Pro Will Be Discontinued
DeepSeek unified its three serving modes into V4.1 Flash, launched today as its fastest tier, and is discontinuing V4 Pro. If you route traffic to DeepSeek, you need to re-point any V4 Pro calls before it goes away.
- DeepSeek V4.1 Flash Prices Cached Input To $0.003 Per Million Tokens
DeepSeek V4.1 Flash prices cached input at $0.003 per million tokens, a rate that makes prompt-caching architectures dramatically cheaper for high-volume, repetitive-context workloads. Worth re-running your cost model if you currently avoid caching because of price.
- DeepSeek-V4.1-Flash Packs 552B Parameters With Efficient MoE Inference
DeepSeek-V4.1-Flash is a 552B-parameter MoE model built for efficient inference, which is the technical reason the Flash tier can undercut the Pro tier. Worth studying as a case of MoE serving economics if you're choosing open-weight models to self-host.
- DeepSeek Routes All V4-Pro API Traffic to V4.1-Flash at Flash Rates
DeepSeek is routing all V4-Pro API traffic to V4.1-Flash and charging Flash rates, so existing V4-Pro integrations now get the cheaper model by default. If you built on V4-Pro, re-check output quality and your cost assumptions rather than assuming the endpoint is unchanged.
Showing the 10 most recent references. 8 older were dropped — a reference ages, so this list does not grow forever.
Three kinds of claim, strongest first. Signal runs every morning.