Home › Frontier Models › DeepSeek V4 Pro
Model · Reference

DeepSeek V4 Pro

Cost-sensitive coding & math

In one line

A 1.6T-parameter MoE with 49B active, MIT-licensed weights, a 1M context and a 384K output ceiling, at $1.32 in and $3.96 out at peak, and half that off-peak.

Why this oneWhat it is actually for

You reach for V4 Pro when you want DeepSeek's reasoning quality rather than its price floor, and you want it without giving up the ability to run the same weights yourself later. It is roughly three times the input price of V4 Flash and three times the output price, which buys 1.6T total parameters against Flash's 284B and 49B active against 13B — a real gap on world knowledge and multi-step agentic work. The 384K max output is the other reason: it is one of very few models that will emit a third of a million tokens in one response, which matters for whole-file rewrites and long plan generation. If your workload is high-volume extraction or summarisation, Flash is the correct call and Pro is waste.

What it isArchitecture, lineage, training

DeepSeek V4 Pro is a sparse mixture-of-experts model released 2026-04-24. DeepSeek's release note states 1.6T total parameters with 49B active per token, and the Hugging Face card describes a "hybrid attention mechanism combining Compressed Sparse Attention and Heavily Compressed Attention", plus Manifold-Constrained Hyper-Connections and the Muon optimizer. The card claims the design needs only 27% of the single-token inference FLOPs and 10% of the KV cache of DeepSeek-V3.2.

The real numbers come from config.json in deepseek-ai/DeepSeek-V4-Pro. hidden_size is 7168 across num_hidden_layers 61. num_attention_heads is 128 with num_key_value_heads of 1 and head_dim 512 — a latent-attention layout, with q_lora_rank 1536 and o_lora_rank 1024. The MoE block has n_routed_experts 384 plus n_shared_experts 1, with num_experts_per_tok 6 and moe_intermediate_size 3072; routing uses topk_method noaux_tc and a sqrtsoftplus scoring function. vocab_size is 129280 and max_position_embeddings is 1048576, reached by YaRN scaling with factor 16 over an original_max_position_embeddings of 65536; rope_theta is 10000 with a separate compress_rope_theta of 160000. The sparse-attention indexer is visible as index_n_heads 64, index_head_dim 128 and index_topk 1024, alongside a sliding_window of 128. Weights ship FP8 e4m3 with 128x128 blocks and expert_dtype fp4, and there is one num_nextn_predict_layers for multi-token prediction.

Training, per the model card: over 32T pretraining tokens, then two-stage post-training — domain-specific expert cultivation followed by unified consolidation via on-policy distillation. Licence is MIT.

At a glanceSee it

DeepSeek V4 Pro diagram

Per-token path through DeepSeek V4 Pro: sparse attention, then six of 384 routed experts plus one shared.

CapacityContext, output and what fits

FactValueSource
Model iddeepseek-v4-proDeepSeek Models & Pricing
Context window1M tokens; max_position_embeddings is exactly 1048576Models & Pricing; config.json
Max output tokens384KDeepSeek Models & Pricing
Input modalitiesTextModel card; Microsoft Foundry lists input as text, languages en and zh
Output modalitiesText, plus reasoning_content when thinking is onDeepSeek API reference
Total / active parameters1.6T total, 49B activeDeepSeek V4 release note
Pretraining tokensOver 32THugging Face model card
Knowledge cutoffNot disclosedNot stated on any DeepSeek page read
Weights availableYes — deepseek-ai/DeepSeek-V4-Pro, plus a -Base variantHugging Face
LicenceMITHugging Face model card
Concurrency limit500DeepSeek Models & Pricing

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
messagesConversation so farRequiredPrior turns and tool schemas all bill as input; caching only helps the unchanged prefix.
max_tokensCaps the completionInteger; ceiling 384K per the Models & Pricing page. The API reference documents no explicit rangeSet it deliberately. Too low is the documented cause of truncated JSON.
temperatureSampling randomnessRange 0 to 2, default 1DeepSeek's own guidance table: 0.0 for coding and maths, 1.0 for data cleaning and analysis, 1.3 for general conversation and translation, 1.5 for creative writing and poetry.
top_pNucleus sampling massRange 0 to 1, default 1The V4 model card recommends leaving it at 1.0. Move temperature or top_p, not both.
thinkingTurns the reasoning pass on or offObject. {"type":"enabled"} / {"type":"disabled"}; default enabledThinking is on unless you turn it off. Disabling drops the deliberation and the reasoning tokens you pay for. Tool calls are supported with thinking enabled.
reasoning_effortHow hard it thinksTop-level field, sent alongside thinking, not nested inside it. Effective values high or max; default high. low and medium are accepted but remapped to high; xhigh is remapped to maxThe remapping is the trap: passing low to save money buys you nothing — you get high and pay for it. The docs also note effort is automatically raised to max for some complex agent requests. The model card recommends a context window of at least 384K for Think Max.
toolsFunction toolsArray, maximum 128Beyond a couple of dozen tools, selection accuracy degrades before the cap bites.
tool_choiceForces or forbids tool usenone, auto, required, or a named functionrequired forces a call even when none is appropriate.
response_formatOutput shape{"type":"text"} (default) or {"type":"json_object"}JSON mode also requires the word "json" in the prompt plus a format example, or output is unreliable.
stopHalt sequencesString or array, up to 16 sequencesSixteen is a generous cap; useful for agent frame delimiters.
streamServer-sent eventsBooleanNear-mandatory with thinking on, since first useful token can be far out.
stream_optionsUsage in the stream{"include_usage": true}The only way to read prompt_cache_hit_tokens and prompt_cache_miss_tokens per call.
logprobsLog probabilities of output tokensBooleanEnables confidence gating on extraction tasks.
top_logprobsAlternatives per positionInteger, 0 to 20Requires logprobs; grows response size, not token cost.
user_idEnd-user identifier for safety reviewString, ≤ 512 chars, charset [a-zA-Z0-9\-_]No effect on generation. Note the restricted charset — email addresses and raw UUIDs with other punctuation will be rejected.
frequency_penalty, presence_penaltyRepetition penaltiesNot supported — the API reference states "This parameter is no longer supported. It will not take effect if you pass it to the API"They fail silently rather than erroring. Handle repetition with stop sequences and prompt instructions instead.

Three knobs carry the weight. thinking is the biggest: it defaults to enabled, so turning it off is an active decision that collapses latency and output tokens, and for extraction or classification that costs you almost nothing in quality. reasoning_effort is the second, and the important fact is that it is a two-position switch wearing a five-position label — only high (the default) and max do anything distinct, and the cheaper-sounding values are remapped upward. Third is temperature, which unusually has vendor-published per-task values: 0.0 for code and maths is a real recommendation from DeepSeek, not folklore, even though the model card's generic deployment advice is 1.0.

SamplingShaping the output distribution

DeepSeek gives you a genuine temperature dial and, unusually, publishes what to set it to. The parameter-settings page lists 0.0 for coding and maths, 1.0 for data cleaning and analysis, 1.3 for general conversation and translation, and 1.5 for creative writing, against a stated default of 1.0. The V4 model card, which is aimed at people running the weights themselves, recommends temperature = 1.0 and top_p = 1.0 as the standard deployment setting. Those two pieces of advice are not in conflict so much as aimed at different callers: the card describes the operating point the model was tuned at, the API page describes how to bend it per task. Both temperature and top_p are accepted, and the usual discipline applies — move one, not both, or you will not know which change caused what. A sensible starting point for agentic and coding work is temperature 0.0 with top_p left at 1.0, then raise temperature only if the model is getting stuck in a rut. Note there is no top_k and no repetition penalty: both penalty parameters were removed.

ReasoningThinking, effort and budgets

Thinking mode is supported and, per DeepSeek's own pricing page, is the default for both V4 models. It is controlled by a thinking object taking {"type":"enabled"} or {"type":"disabled"}, with a separate top-level reasoning_effort string accepting high or max. The model card refers to the top setting as "Think Max" and states it needs a context window of at least 384K tokens — that is a hard operational constraint, not a suggestion, and it is the reason the max output figure and the Think Max requirement are the same number. Reasoning is surfaced as reasoning_content separate from content. Tool calling works while thinking is enabled; the docs note this has been true since V3.2. What the pricing page does not say is whether thinking tokens are billed at a different rate from ordinary output tokens — it gives one output price and no breakdown. Budget as if all reasoning tokens bill at the $3.96 peak output rate, or $1.98 off-peak.

ToolsFunction calling and server tools

Tool calling is supported, works alongside thinking mode, and takes up to 128 tools. tool_choice accepts none, auto, required or a specific function. The docs do not state whether multiple tool calls come back in one response, so do not architect around parallel dispatch without testing it. For structured arguments there is a strict mode: point the client at the beta base URL https://api.deepseek.com/beta and set "strict": true on the function. Strict mode supports object, string, number, integer, boolean, array, enum, anyOf, $ref and $def, but rejects minLength, maxLength, minItems and maxItems, and requires every property to be listed in required with additionalProperties: false. Two failure modes to plan for. JSON mode via response_format is documented as occasionally returning empty content, which DeepSeek says it is still optimising — retry rather than assume a parse bug. And separately, if you reach this model through Microsoft Foundry rather than DeepSeek's API, Foundry's own model table lists tool calling as not supported for DeepSeek-V4-Pro.

CostPrice, caching, batching, what drives the bill

List price from DeepSeek's Models & Pricing page, USD per million tokens. DeepSeek prices this model by time of day: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, and off-peak rates — every other hour, the whole weekend included — are exactly half the peak rates.

ItemOff-peakPeak
Input, cache miss$0.66$1.32
Input, cache hit$0.022$0.044
Output$1.98$3.96

The cache-hit rate is 30x below the miss rate on either schedule, which is the single most important economic fact about this model. Caching is "enabled by default for all users, without needing to modify their code" and requires no parameter; the API reports prompt_cache_hit_tokens and prompt_cache_miss_tokens in the usage object. Caches are cleared automatically after "a few hours to a few days" of disuse. No cache-write fee is documented. Note what DeepSeek does not publish: there is no documented minimum cache unit size, prefix granularity or cache-construction rule, so how partial prefix overlap is treated could not be confirmed — measure it with the two usage counters rather than assuming a boundary. No batch discount is published. The time-of-day split is the discount, and it is the cheapest lever on this page: a job you can defer out of the two peak windows costs half as much for no code change at all. Plan the clock before you plan the prompt.

Where it runsSurfaces and availability

SurfaceAvailableNotes
DeepSeek APIYeshttps://api.deepseek.com, model id deepseek-v4-pro. OpenAI ChatCompletions and Anthropic-compatible interfaces both supported.
Microsoft FoundryYesSold directly by Azure as DeepSeek-V4-Pro: 1,000,000 input / 384,000 output tokens, languages en and zh. Foundry's table lists tool calling as not supported.
Alibaba Cloud Model StudioYesListed among the third-party models with hybrid thinking support.
Amazon BedrockNoBedrock's DeepSeek section lists V3.2, V3.1 and R1 only.
Google Vertex AINoVertex's managed-API list shows DeepSeek-V3.2, V3.1, R1-0528 and OCR; V4 does not appear.
Hugging Face weightsYesdeepseek-ai/DeepSeek-V4-Pro and DeepSeek-V4-Pro-Base, MIT licence, FP8 with FP4 experts.
Self-hostingYesWeights are public. Serving a 1.6T-parameter model is a serious hardware commitment even at FP8/FP4.

StrengthsWhat it is good at

  • MIT-licensed weights for a genuine frontier-scale model — 1.6T total, 49B active — with a -Base variant published alongside the instruct model.
  • 384K max output tokens, which is an order of magnitude beyond most models and makes whole-repository rewrites feasible in one response.
  • Cache-hit input at $0.044 per million against a $1.32 miss at peak — a 30x gap — applied automatically, with per-request hit/miss counts exposed in usage.
  • Vendor-published per-task temperature guidance — 0.0 for code and maths, 1.5 for creative — rather than leaving callers to guess.
  • Strict JSON-schema tool arguments available on the beta base URL, with the supported and unsupported schema keywords enumerated.

LimitsWhere it falls down

  • Text only. No image or video input at all, which rules it out for document-vision and UI-agent work that Kimi K3 and the Qwen Max line handle.
  • frequency_penalty and presence_penalty were removed from the API, so repetition has to be managed by prompt and stop sequences.
  • JSON mode is documented as occasionally returning empty content — the vendor calls it a known issue under optimisation, so you need retry logic.
  • Not on Bedrock or Vertex. Microsoft Foundry carries it, but with tool calling listed as unsupported on that surface.
  • Reasoning-effort max requires a 384K context window to operate, which constrains how you can deploy it self-hosted.

Against its neighboursHow it compares

Against DeepSeek V4 Flash, the sibling on the same endpoint: Pro costs exactly 3x more on input and 3x more on output on the same schedule, for 1.6T total against 284B and 49B active against 13B. Same 1M context, same 384K output ceiling, same parameters, same MIT licence. The upgrade buys world knowledge and agentic reliability; DeepSeek's own release note says Flash's reasoning "closely approaches" Pro when given a larger thinking budget, which is a strong hint that for reasoning-heavy but knowledge-light work, Flash plus more thinking is the better trade. Against Kimi K3, Pro is roughly 3.8x cheaper on output at peak and 7.6x off-peak and gives you sampling controls and downloadable weights, but has no vision and a smaller reported coding profile. Against Claude Sonnet 5, Pro is the self-hostable option; Sonnet is the one with broad cloud distribution.

Getting startedThe smallest call that works

code
POST https://api.deepseek.com/chat/completions
Authorization: Bearer $DEEPSEEK_API_KEY
Content-Type: application/json

{
  "model": "deepseek-v4-pro",
  "messages": [
    {"role": "user", "content": "Explain what a KV cache is, in two sentences."}
  ]
}

Change temperature first: set it to 0.0 for code and maths, per DeepSeek's own guidance table. Then decide on thinking — send {"type":"disabled"} for extraction and classification and watch the bill drop. Add "stream": true with "stream_options": {"include_usage": true} to see your cache hit rate. Leave top_p at 1.0.

SourcesWhere every claim above came from

Checked

Price and capacity verified 2026-09-12 against https://api-docs.deepseek.com/quick_start/pricing/. re-read 2026-09-12 (research pass, primary source read today: https://api-docs.deepseek.com/quick_start/pricing/ — figures confirmed: input_per_m 1.32, output_per_m 3.96, cache_hit_input_per_m 0.044, context 1M, max_output 384K.) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $1.32 in / $3.96 out per M, $0.044 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-09-04 (B2; the CITATION was corrected, not the price. The cited URL lacked a trailing slash and served "Your First API Call" with ZERO price-like numbers on it — a citation nothing could re-read, on the model 304 of 305 kits run on. With the slash it serves "Models & Pricing", and a live read that day found ALL TEN figures the two DeepSeek rows carry — peak, off-peak and cache-hit — present and unchanged: 0.007 0.014 0.022 0.044 0.22 0.44 0.66 1.32 1.98 3.96. No price moved. Prior: manual 2026-08-23 (currency proposal src-2026-08-23-8eb120; every peak and off-peak figure unchanged; the peak window carries a 'Monday through Friday' qualifier this row did not record).

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page DeepSeek now serves all V4-Pro API traffic with V4.1-Flash and bills at Flash rates, so the V4-Pro page's pricing no longer describes what a caller pays.

    Update the DeepSeek V4 Pro page to state that its API traffic is now routed to V4.1-Flash at Flash rates, so the listed $1.32/$3.96 peak pricing no longer applies to live calls.

    China frontier labs · 13 Sep 2026 · source

  • Updated this page A 552B-parameter open-weight MoE model, DeepSeek-V4.1-Flash, is available on Alibaba Cloud at Flash pricing.

    Add DeepSeek-V4.1-Flash to the DeepSeek V4 Flash page as a 552B-parameter MoE successor available on Alibaba Cloud, and note what it means for the self-host-versus-API decision.

    China frontier labs · 13 Sep 2026 · source

RecentAuto-linked from the brief, not a rewrite of this page

Showing the 10 most recent references. 8 older were dropped — a reference ages, so this list does not grow forever.

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning