Home › Frontier Models › DeepSeek V4 Flash
Model · Reference

DeepSeek V4 Flash

High-volume back-office

In one line

284B MoE with 13B active, MIT weights, 1M context and 384K output at $0.44 in and $1.32 out at peak, half that off-peak — the cheapest million-token model here.

Why this oneWhat it is actually for

You reach for V4 Flash when volume is the problem. At $0.44 per million input tokens at peak — $0.22 off-peak — with a cache-hit rate of $0.014, a pipeline that reads a large fixed corpus every day costs a rounding error, and the 1M window means you rarely have to build a chunking layer at all. It is the right default for extraction, classification, summarisation, log triage and any back-office job where you would otherwise be arguing about batch windows. It is the wrong choice when world knowledge matters — 13B active parameters shows up as thinner factual recall than V4 Pro. One hard operational note: if you are still calling deepseek-chat or deepseek-reasoner, those aliases were retired on 2026-07-24 and you must now name this model directly.

What it isArchitecture, lineage, training

DeepSeek V4 Flash is the smaller of the two V4 models released 2026-04-24, and it is not a distillation of Pro so much as the same architecture built narrower. The release note gives 284B total parameters with 13B active; the model card describes the same "hybrid attention mechanism combining Compressed Sparse Attention and Heavily Compressed Attention" and the same 32T-plus pretraining corpus and two-stage post-training as Pro.

From config.json in deepseek-ai/DeepSeek-V4-Flash: hidden_size 4096 across num_hidden_layers 43, against Pro's 7168 and 61. num_attention_heads is 64 with num_key_value_heads 1 and head_dim 512, and q_lora_rank is 1024. The MoE block runs n_routed_experts 256 plus n_shared_experts 1, with num_experts_per_tok 6 and moe_intermediate_size 2048 — note the same six-expert routing as Pro over a smaller pool, and a lower routed_scaling_factor of 1.5 against Pro's 2.5. vocab_size is 129280 and max_position_embeddings 1048576, again via YaRN factor 16 over an original 65536. The sparse-attention indexer is halved relative to Pro: index_topk 512 against 1024, with index_n_heads 64 and index_head_dim 128. Quantisation matches — FP8 e4m3 in 128x128 blocks with expert_dtype fp4 — as does the single num_nextn_predict_layers for multi-token prediction.

The model card publishes its own benchmark numbers, which are vendor claims rather than independent results: 86.2 on MMLU-Pro, 34.1 on SimpleQA-Verified in Max mode, and 78.7 on MRCR at 1M context. Licence is MIT, and the card ships vLLM and SGLang deployment instructions including Docker.

At a glanceSee it

DeepSeek V4 Flash diagram

Where a V4 Flash request's cost is decided: the cache split on the way in, the thinking toggle on the way out.

CapacityContext, output and what fits

FactValueSource
Model iddeepseek-v4-flashDeepSeek Models & Pricing
Deprecated aliasesdeepseek-chat and deepseek-reasoner retired 2026-07-24 15:59 UTC; they mapped to this model's non-thinking and thinking modesModels & Pricing; Change Log
Context window1M tokens; max_position_embeddings is exactly 1048576Models & Pricing; config.json
Max output tokens384KDeepSeek Models & Pricing
Input modalitiesTextModel card; Microsoft Foundry lists text input, languages en and zh
Output modalitiesText, plus reasoning_content when thinking is onDeepSeek API reference
Total / active parameters284B total, 13B activeDeepSeek V4 release note
Pretraining tokensOver 32THugging Face model card
Knowledge cutoffNot disclosedNot stated on any DeepSeek page read
Weights availableYes — deepseek-ai/DeepSeek-V4-Flash, plus a -Base variantHugging Face
LicenceMITHugging Face model card
Concurrency limit2500DeepSeek Models & Pricing

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
modelSelects the modeldeepseek-v4-flashName it explicitly. deepseek-chat and deepseek-reasoner were discontinued on 2026-07-24 15:59 UTC; the docs state they are no longer available but do not describe the failure mode, so treat the call as hard-failing rather than degrading.
messagesConversation so farRequiredA stable system prefix is what makes the $0.014 cache-hit rate reachable; reordering it destroys the hit.
max_tokensCaps the completionInteger; ceiling 384K per the Models & Pricing page. The API reference documents no explicit rangeThe documented cause of truncated JSON is setting this too low.
temperatureSampling randomnessRange 0 to 2, default 1DeepSeek's guidance: 0.0 for coding and maths, 1.0 for data cleaning and analysis, 1.3 for general conversation and translation, 1.5 for creative writing and poetry.
top_pNucleus sampling massRange 0 to 1, default 1The model card recommends 1.0. Adjust temperature instead.
thinkingTurns the reasoning pass on or offObject. {"type":"enabled"} / {"type":"disabled"}; default enabledThe main cost lever on this model, and it is on unless you turn it off. Any interaction between thinking mode and FIM completion is not documented — the FIM guide does not mention thinking mode, and its example uses deepseek-v4-pro.
reasoning_effortHow hard it thinksTop-level field, sent alongside thinking. Effective values high or max; default high. low and medium are remapped to high; xhigh to maxPassing low does not save money — it is remapped to high. Only disabled thinking actually cuts the reasoning bill.
toolsFunction toolsArray, maximum 128The API reference documents tool calling on the V4 models and the thinking-mode guide states thinking mode supports tool calls. Claims about which earlier version introduced this do not survive the change log — V3.2-Speciale is documented as having no tool calls.
tool_choiceForces or forbids tool usenone, auto, required, or a named functionrequired will force a call even when nothing fits.
response_formatOutput shape{"type":"text"} (default) or {"type":"json_object"}JSON mode needs the word "json" plus a worked example in the prompt, or output is unreliable.
stopHalt sequencesString or array, up to 16 sequencesCheap way to bound runaway generations on a high-volume pipeline.
streamServer-sent eventsBooleanLess necessary here than on Pro; Flash's first token arrives quickly with thinking off.
stream_optionsUsage in the stream{"include_usage": true}Surfaces prompt_cache_hit_tokens and prompt_cache_miss_tokens — the numbers your unit economics depend on.
logprobsLog probabilitiesBooleanThe cheapest confidence signal available for automated extraction gating.
top_logprobsAlternatives per positionInteger, 0 to 20Requires logprobs. No token cost.
user_idEnd-user identifier for safety reviewString, ≤ 512 chars, charset [a-zA-Z0-9\-_]No effect on generation. The restricted charset rejects most raw email addresses.
frequency_penalty, presence_penaltyRepetition penaltiesNot supported — the API reference states "This parameter is no longer supported. It will not take effect if you pass it to the API"They fail silently rather than erroring. Use stop sequences and explicit prompt instructions instead.

On Flash the decisive knob is thinking, and note it defaults to enabled — most high-volume back-office work does not need a reasoning pass, and turning it off cuts both latency and the output tokens you pay for, which is the whole reason this model is cheap. Reaching for reasoning_effort: "low" instead does nothing, because it is remapped to high. The second lever is prompt-prefix stability, which is not a parameter but behaves like one: the 31x gap between $0.44 and $0.014 is only realised if your system prompt is byte-identical across calls. Third is logprobs, because at this price point the sensible pattern is to run everything through Flash and escalate low-confidence items to Pro.

SamplingShaping the output distribution

Flash exposes the same sampling surface as Pro: temperature defaulting to 1 with a maximum of 2, top_p defaulting to 1, and no top_k or repetition penalties. DeepSeek publishes concrete per-task temperature values — 0.0 for coding and maths, 1.0 for data cleaning and analysis, 1.3 for general conversation and translation, 1.5 for creative writing — while the Hugging Face card, written for people running the weights, recommends temperature = 1.0 with top_p = 1.0 as the tuned operating point. For a high-volume extraction pipeline, temperature 0.0 with top_p left alone is the right starting posture: it makes outputs far more reproducible run-to-run, which matters more than creativity when a downstream parser is consuming the result. Because both penalty parameters were removed from the API, a smaller model like this one can loop on repetitive inputs with no dedicated defence — bound it with max_tokens and stop. If you are self-hosting via vLLM or SGLang, you have access to sampler options the hosted API does not expose, which is one underrated argument for running the weights.

ReasoningThinking, effort and budgets

Flash has a full thinking mode, supported by default, toggled by a thinking object taking {"type":"enabled"} or {"type":"disabled"}, with reasoning_effort accepting high or max. This is the pair that the retired deepseek-reasoner and deepseek-chat aliases used to select between. The model card makes a specific and useful claim: Flash's reasoning "closely approaches" V4 Pro when given a larger thinking budget — so on reasoning-bound tasks, Flash at reasoning_effort: max is a serious alternative to Pro at a third of the price, provided you can afford the extra output tokens and provided you give it a 384K-plus context window, which Think Max requires. Reasoning is returned in reasoning_content. One non-obvious interaction: FIM completion is only available in non-thinking mode, so code-infill workloads must explicitly disable thinking. As on Pro, the pricing page gives a single output rate with no separate line for reasoning tokens.

ToolsFunction calling and server tools

Tool calling is supported with a 128-tool cap, works in thinking mode, and takes tool_choice values of none, auto, required or a named function. Whether Flash returns several tool calls in one response is not stated in the docs, so verify before building a parallel dispatcher. Strict JSON-schema arguments are available on the beta base URL https://api.deepseek.com/beta with "strict": true on the function; the supported keywords are object, string, number, integer, boolean, array, enum, anyOf, $ref and $def, while minLength, maxLength, minItems and maxItems are rejected and every property must appear in required with additionalProperties: false. The failure mode to design around is documented plainly: JSON mode "may occasionally return empty content", which DeepSeek describes as under active optimisation. At Flash's volumes that is not a rare event, so treat empty content as a retryable condition rather than an exception. If you consume this model through Microsoft Foundry instead, note that Foundry's table lists tool calling as not supported for DeepSeek-V4-Flash.

CostPrice, caching, batching, what drives the bill

List price from DeepSeek's Models & Pricing page, USD per million tokens. DeepSeek prices this model by time of day: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, and off-peak rates — every other hour, the whole weekend included — are exactly half the peak rates.

ItemOff-peakPeak
Input, cache miss$0.22$0.44
Input, cache hit$0.007$0.014
Output$0.66$1.32

Three facts govern the bill. First, the cache-hit rate is roughly 1/31st of the miss rate on either schedule and caching is "enabled by default for all users, without needing to modify their code" — no parameter, no cache-control block, no write fee documented. Caches are dropped after "a few hours to a few days" of disuse. DeepSeek publishes no minimum cache unit size or prefix-matching rule, so how partial overlap is scored could not be confirmed; instrument it with the usage counters rather than designing against an assumed boundary. Second, output is 3x input here, still a much flatter ratio than the frontier models, so verbosity hurts less and thinking mode hurts more in relative terms. Third, there is no batch discount, but there is a time-of-day discount: shifting a queue out of the two peak windows halves the bill with no code change, which for a throughput pipeline is usually the largest single saving available. The concurrency ceiling of 2500 is five times Pro's 500, which matters as much as price for a throughput pipeline. Read prompt_cache_hit_tokens and prompt_cache_miss_tokens from usage before you model your unit economics on the headline number.

Where it runsSurfaces and availability

SurfaceAvailableNotes
DeepSeek APIYeshttps://api.deepseek.com, model id deepseek-v4-flash. OpenAI ChatCompletions and Anthropic-compatible interfaces.
Microsoft FoundryYesSold directly by Azure as DeepSeek-V4-Flash: 1,000,000 input / 384,000 output, languages en and zh. Tool calling listed as not supported on that surface.
Alibaba Cloud Model StudioYesListed among third-party models with hybrid thinking support.
Amazon BedrockNoBedrock's DeepSeek section lists V3.2, V3.1 and R1 only.
Google Vertex AINoVertex's managed-API list shows V3.2, V3.1, R1-0528 and OCR.
Hugging Face weightsYesdeepseek-ai/DeepSeek-V4-Flash and -Base, MIT licence.
Self-hostingYesThe model card ships vLLM and SGLang instructions including Docker. At 284B total / 13B active this is genuinely serveable, unlike Pro.

StrengthsWhat it is good at

  • $0.44 per million input tokens across a full 1M context at peak, $0.22 off-peak — the cheapest million-token window on this board, with a cache-hit rate of $0.014.
  • 384K max output tokens at Efficient-tier pricing, a ceiling most Frontier-tier models do not offer.
  • MIT-licensed weights at a size you can actually serve — 284B total, 13B active — with vLLM and SGLang instructions on the model card.
  • Concurrency limit of 2500 against Pro's 500, which is what makes it a throughput model and not just a cheap one.
  • Vendor-published benchmarks on the card give a concrete quality anchor: 86.2 MMLU-Pro, 78.7 MRCR at 1M, 34.1 SimpleQA-Verified in Max mode.

LimitsWhere it falls down

  • 13B active parameters means noticeably thinner world knowledge than V4 Pro — the card's own 34.1 on SimpleQA-Verified is a factual-recall number, not a reasoning one.
  • Text only, no vision, so document-image and screenshot pipelines need a separate model.
  • The deepseek-chat and deepseek-reasoner aliases were retired on 2026-07-24; any integration still using them is already broken, not about to break.
  • JSON mode occasionally returns empty content by the vendor's own admission, and both repetition-penalty parameters were removed, so small-model looping has no dedicated defence.
  • Not on Bedrock or Vertex; on Microsoft Foundry, tool calling is listed as unsupported.

Against its neighboursHow it compares

Against DeepSeek V4 Pro, Flash is one third the price on both input and output with identical context, output ceiling, API surface and licence. The card says Flash's reasoning closely approaches Pro's given a larger thinking budget, so the honest split is: use Flash and spend the savings on reasoning_effort when the task is reasoning-bound, and reach for Pro when it is knowledge-bound or agentically long. Against Gemini 3 Flash, the other obvious Efficient-tier long-context option, DeepSeek's differentiator is downloadable MIT weights and a 384K output ceiling; Google's is multimodal input and first-party cloud distribution. Against Claude Haiku 4.5, Flash wins on context and price and loses on ecosystem breadth and vision.

Getting startedThe smallest call that works

code
POST https://api.deepseek.com/chat/completions
Authorization: Bearer $DEEPSEEK_API_KEY
Content-Type: application/json

{
  "model": "deepseek-v4-flash",
  "messages": [
    {"role": "user", "content": "Summarise this in one sentence: ..."}
  ],
  "thinking": {"type": "disabled"}
}

Change temperature to 0.0 first if a parser consumes the output — DeepSeek's own table recommends it for structured work. Then add "stream_options": {"include_usage": true} and check prompt_cache_hit_tokens; if it is zero, your system prompt is not stable and you are paying 50x more than you need to on input.

SourcesWhere every claim above came from

Checked

Price and capacity verified 2026-09-05 against https://api-docs.deepseek.com/quick_start/pricing/. re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $0.44 in / $1.32 out per M, $0.014 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-09-04 (B2; the CITATION was corrected, not the price. The cited URL lacked a trailing slash and served "Your First API Call" with ZERO price-like numbers on it — a citation nothing could re-read, on the model 304 of 305 kits run on. With the slash it serves "Models & Pricing", and a live read that day found ALL TEN figures the two DeepSeek rows carry — peak, off-peak and cache-hit — present and unchanged: 0.007 0.014 0.022 0.044 0.22 0.44 0.66 1.32 1.98 3.96. No price moved. Prior: manual 2026-08-23 (currency proposal src-2026-08-23-8eb120; every peak and off-peak figure unchanged; the peak window carries a 'Monday through Friday' qualifier this row did not record).

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page A reported DeepSeek 4.1 Flash cuts memory use 4x versus DeepSeek 4.0 Flash, which would lower self-hosting cost and allow longer contexts on the same hardware.

    Update the DeepSeek V4 Flash page to note the reported DeepSeek 4.1 Flash release and its claimed 4x memory reduction, since the page currently presents V4 Flash as the current Flash-class entry.

    China frontier labs · 20 Sep 2026 · source

  • Updated this page A 552B-parameter open-weight MoE model, DeepSeek-V4.1-Flash, is available on Alibaba Cloud at Flash pricing.

    Add DeepSeek-V4.1-Flash to the DeepSeek V4 Flash page as a 552B-parameter MoE successor available on Alibaba Cloud, and note what it means for the self-host-versus-API decision.

    China frontier labs · 13 Sep 2026 · source

RecentAuto-linked from the brief, not a rewrite of this page

Showing the 10 most recent references. 10 older were dropped — a reference ages, so this list does not grow forever.

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning