284B MoE with 13B active, MIT weights, 1M context and 384K output at $0.44 in and $1.32 out at peak, half that off-peak — the cheapest million-token model here.
Why this oneWhat it is actually for
You reach for V4 Flash when volume is the problem. At $0.44 per million input tokens at peak — $0.22 off-peak — with a cache-hit rate of $0.014, a pipeline that reads a large fixed corpus every day costs a rounding error, and the 1M window means you rarely have to build a chunking layer at all. It is the right default for extraction, classification, summarisation, log triage and any back-office job where you would otherwise be arguing about batch windows. It is the wrong choice when world knowledge matters — 13B active parameters shows up as thinner factual recall than V4 Pro. One hard operational note: if you are still calling deepseek-chat or deepseek-reasoner, those aliases were retired on 2026-07-24 and you must now name this model directly.
What it isArchitecture, lineage, training
DeepSeek V4 Flash is the smaller of the two V4 models released 2026-04-24, and it is not a distillation of Pro so much as the same architecture built narrower. The release note gives 284B total parameters with 13B active; the model card describes the same "hybrid attention mechanism combining Compressed Sparse Attention and Heavily Compressed Attention" and the same 32T-plus pretraining corpus and two-stage post-training as Pro.
From config.json in deepseek-ai/DeepSeek-V4-Flash: hidden_size 4096 across num_hidden_layers 43, against Pro's 7168 and 61. num_attention_heads is 64 with num_key_value_heads 1 and head_dim 512, and q_lora_rank is 1024. The MoE block runs n_routed_experts 256 plus n_shared_experts 1, with num_experts_per_tok 6 and moe_intermediate_size 2048 — note the same six-expert routing as Pro over a smaller pool, and a lower routed_scaling_factor of 1.5 against Pro's 2.5. vocab_size is 129280 and max_position_embeddings 1048576, again via YaRN factor 16 over an original 65536. The sparse-attention indexer is halved relative to Pro: index_topk 512 against 1024, with index_n_heads 64 and index_head_dim 128. Quantisation matches — FP8 e4m3 in 128x128 blocks with expert_dtype fp4 — as does the single num_nextn_predict_layers for multi-token prediction.
The model card publishes its own benchmark numbers, which are vendor claims rather than independent results: 86.2 on MMLU-Pro, 34.1 on SimpleQA-Verified in Max mode, and 78.7 on MRCR at 1M context. Licence is MIT, and the card ships vLLM and SGLang deployment instructions including Docker.
At a glanceSee it
Where a V4 Flash request's cost is decided: the cache split on the way in, the thinking toggle on the way out.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| Model id | deepseek-v4-flash | DeepSeek Models & Pricing |
| Deprecated aliases | deepseek-chat and deepseek-reasoner retired 2026-07-24 15:59 UTC; they mapped to this model's non-thinking and thinking modes | Models & Pricing; Change Log |
| Context window | 1M tokens; max_position_embeddings is exactly 1048576 | Models & Pricing; config.json |
| Max output tokens | 384K | DeepSeek Models & Pricing |
| Input modalities | Text | Model card; Microsoft Foundry lists text input, languages en and zh |
| Output modalities | Text, plus reasoning_content when thinking is on | DeepSeek API reference |
| Total / active parameters | 284B total, 13B active | DeepSeek V4 release note |
| Pretraining tokens | Over 32T | Hugging Face model card |
| Knowledge cutoff | Not disclosed | Not stated on any DeepSeek page read |
| Weights available | Yes — deepseek-ai/DeepSeek-V4-Flash, plus a -Base variant | Hugging Face |
| Licence | MIT | Hugging Face model card |
| Concurrency limit | 2500 | DeepSeek Models & Pricing |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
model | Selects the model | deepseek-v4-flash | Name it explicitly. deepseek-chat and deepseek-reasoner were discontinued on 2026-07-24 15:59 UTC; the docs state they are no longer available but do not describe the failure mode, so treat the call as hard-failing rather than degrading. |
messages | Conversation so far | Required | A stable system prefix is what makes the $0.014 cache-hit rate reachable; reordering it destroys the hit. |
max_tokens | Caps the completion | Integer; ceiling 384K per the Models & Pricing page. The API reference documents no explicit range | The documented cause of truncated JSON is setting this too low. |
temperature | Sampling randomness | Range 0 to 2, default 1 | DeepSeek's guidance: 0.0 for coding and maths, 1.0 for data cleaning and analysis, 1.3 for general conversation and translation, 1.5 for creative writing and poetry. |
top_p | Nucleus sampling mass | Range 0 to 1, default 1 | The model card recommends 1.0. Adjust temperature instead. |
thinking | Turns the reasoning pass on or off | Object. {"type":"enabled"} / {"type":"disabled"}; default enabled | The main cost lever on this model, and it is on unless you turn it off. Any interaction between thinking mode and FIM completion is not documented — the FIM guide does not mention thinking mode, and its example uses deepseek-v4-pro. |
reasoning_effort | How hard it thinks | Top-level field, sent alongside thinking. Effective values high or max; default high. low and medium are remapped to high; xhigh to max | Passing low does not save money — it is remapped to high. Only disabled thinking actually cuts the reasoning bill. |
tools | Function tools | Array, maximum 128 | The API reference documents tool calling on the V4 models and the thinking-mode guide states thinking mode supports tool calls. Claims about which earlier version introduced this do not survive the change log — V3.2-Speciale is documented as having no tool calls. |
tool_choice | Forces or forbids tool use | none, auto, required, or a named function | required will force a call even when nothing fits. |
response_format | Output shape | {"type":"text"} (default) or {"type":"json_object"} | JSON mode needs the word "json" plus a worked example in the prompt, or output is unreliable. |
stop | Halt sequences | String or array, up to 16 sequences | Cheap way to bound runaway generations on a high-volume pipeline. |
stream | Server-sent events | Boolean | Less necessary here than on Pro; Flash's first token arrives quickly with thinking off. |
stream_options | Usage in the stream | {"include_usage": true} | Surfaces prompt_cache_hit_tokens and prompt_cache_miss_tokens — the numbers your unit economics depend on. |
logprobs | Log probabilities | Boolean | The cheapest confidence signal available for automated extraction gating. |
top_logprobs | Alternatives per position | Integer, 0 to 20 | Requires logprobs. No token cost. |
user_id | End-user identifier for safety review | String, ≤ 512 chars, charset [a-zA-Z0-9\-_] | No effect on generation. The restricted charset rejects most raw email addresses. |
frequency_penalty, presence_penalty | Repetition penalties | Not supported — the API reference states "This parameter is no longer supported. It will not take effect if you pass it to the API" | They fail silently rather than erroring. Use stop sequences and explicit prompt instructions instead. |
On Flash the decisive knob is thinking, and note it defaults to enabled — most high-volume back-office work does not need a reasoning pass, and turning it off cuts both latency and the output tokens you pay for, which is the whole reason this model is cheap. Reaching for reasoning_effort: "low" instead does nothing, because it is remapped to high. The second lever is prompt-prefix stability, which is not a parameter but behaves like one: the 31x gap between $0.44 and $0.014 is only realised if your system prompt is byte-identical across calls. Third is logprobs, because at this price point the sensible pattern is to run everything through Flash and escalate low-confidence items to Pro.
SamplingShaping the output distribution
Flash exposes the same sampling surface as Pro: temperature defaulting to 1 with a maximum of 2, top_p defaulting to 1, and no top_k or repetition penalties. DeepSeek publishes concrete per-task temperature values — 0.0 for coding and maths, 1.0 for data cleaning and analysis, 1.3 for general conversation and translation, 1.5 for creative writing — while the Hugging Face card, written for people running the weights, recommends temperature = 1.0 with top_p = 1.0 as the tuned operating point. For a high-volume extraction pipeline, temperature 0.0 with top_p left alone is the right starting posture: it makes outputs far more reproducible run-to-run, which matters more than creativity when a downstream parser is consuming the result. Because both penalty parameters were removed from the API, a smaller model like this one can loop on repetitive inputs with no dedicated defence — bound it with max_tokens and stop. If you are self-hosting via vLLM or SGLang, you have access to sampler options the hosted API does not expose, which is one underrated argument for running the weights.
ReasoningThinking, effort and budgets
Flash has a full thinking mode, supported by default, toggled by a thinking object taking {"type":"enabled"} or {"type":"disabled"}, with reasoning_effort accepting high or max. This is the pair that the retired deepseek-reasoner and deepseek-chat aliases used to select between. The model card makes a specific and useful claim: Flash's reasoning "closely approaches" V4 Pro when given a larger thinking budget — so on reasoning-bound tasks, Flash at reasoning_effort: max is a serious alternative to Pro at a third of the price, provided you can afford the extra output tokens and provided you give it a 384K-plus context window, which Think Max requires. Reasoning is returned in reasoning_content. One non-obvious interaction: FIM completion is only available in non-thinking mode, so code-infill workloads must explicitly disable thinking. As on Pro, the pricing page gives a single output rate with no separate line for reasoning tokens.
ToolsFunction calling and server tools
Tool calling is supported with a 128-tool cap, works in thinking mode, and takes tool_choice values of none, auto, required or a named function. Whether Flash returns several tool calls in one response is not stated in the docs, so verify before building a parallel dispatcher. Strict JSON-schema arguments are available on the beta base URL https://api.deepseek.com/beta with "strict": true on the function; the supported keywords are object, string, number, integer, boolean, array, enum, anyOf, $ref and $def, while minLength, maxLength, minItems and maxItems are rejected and every property must appear in required with additionalProperties: false. The failure mode to design around is documented plainly: JSON mode "may occasionally return empty content", which DeepSeek describes as under active optimisation. At Flash's volumes that is not a rare event, so treat empty content as a retryable condition rather than an exception. If you consume this model through Microsoft Foundry instead, note that Foundry's table lists tool calling as not supported for DeepSeek-V4-Flash.
CostPrice, caching, batching, what drives the bill
List price from DeepSeek's Models & Pricing page, USD per million tokens. DeepSeek prices this model by time of day: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, and off-peak rates — every other hour, the whole weekend included — are exactly half the peak rates.
| Item | Off-peak | Peak |
|---|---|---|
| Input, cache miss | $0.22 | $0.44 |
| Input, cache hit | $0.007 | $0.014 |
| Output | $0.66 | $1.32 |
Three facts govern the bill. First, the cache-hit rate is roughly 1/31st of the miss rate on either schedule and caching is "enabled by default for all users, without needing to modify their code" — no parameter, no cache-control block, no write fee documented. Caches are dropped after "a few hours to a few days" of disuse. DeepSeek publishes no minimum cache unit size or prefix-matching rule, so how partial overlap is scored could not be confirmed; instrument it with the usage counters rather than designing against an assumed boundary. Second, output is 3x input here, still a much flatter ratio than the frontier models, so verbosity hurts less and thinking mode hurts more in relative terms. Third, there is no batch discount, but there is a time-of-day discount: shifting a queue out of the two peak windows halves the bill with no code change, which for a throughput pipeline is usually the largest single saving available. The concurrency ceiling of 2500 is five times Pro's 500, which matters as much as price for a throughput pipeline. Read prompt_cache_hit_tokens and prompt_cache_miss_tokens from usage before you model your unit economics on the headline number.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| DeepSeek API | Yes | https://api.deepseek.com, model id deepseek-v4-flash. OpenAI ChatCompletions and Anthropic-compatible interfaces. |
| Microsoft Foundry | Yes | Sold directly by Azure as DeepSeek-V4-Flash: 1,000,000 input / 384,000 output, languages en and zh. Tool calling listed as not supported on that surface. |
| Alibaba Cloud Model Studio | Yes | Listed among third-party models with hybrid thinking support. |
| Amazon Bedrock | No | Bedrock's DeepSeek section lists V3.2, V3.1 and R1 only. |
| Google Vertex AI | No | Vertex's managed-API list shows V3.2, V3.1, R1-0528 and OCR. |
| Hugging Face weights | Yes | deepseek-ai/DeepSeek-V4-Flash and -Base, MIT licence. |
| Self-hosting | Yes | The model card ships vLLM and SGLang instructions including Docker. At 284B total / 13B active this is genuinely serveable, unlike Pro. |
StrengthsWhat it is good at
- $0.44 per million input tokens across a full 1M context at peak, $0.22 off-peak — the cheapest million-token window on this board, with a cache-hit rate of $0.014.
- 384K max output tokens at Efficient-tier pricing, a ceiling most Frontier-tier models do not offer.
- MIT-licensed weights at a size you can actually serve — 284B total, 13B active — with vLLM and SGLang instructions on the model card.
- Concurrency limit of 2500 against Pro's 500, which is what makes it a throughput model and not just a cheap one.
- Vendor-published benchmarks on the card give a concrete quality anchor: 86.2 MMLU-Pro, 78.7 MRCR at 1M, 34.1 SimpleQA-Verified in Max mode.
LimitsWhere it falls down
- 13B active parameters means noticeably thinner world knowledge than V4 Pro — the card's own 34.1 on SimpleQA-Verified is a factual-recall number, not a reasoning one.
- Text only, no vision, so document-image and screenshot pipelines need a separate model.
- The
deepseek-chatanddeepseek-reasoneraliases were retired on 2026-07-24; any integration still using them is already broken, not about to break. - JSON mode occasionally returns empty content by the vendor's own admission, and both repetition-penalty parameters were removed, so small-model looping has no dedicated defence.
- Not on Bedrock or Vertex; on Microsoft Foundry, tool calling is listed as unsupported.
Against its neighboursHow it compares
Against DeepSeek V4 Pro, Flash is one third the price on both input and output with identical context, output ceiling, API surface and licence. The card says Flash's reasoning closely approaches Pro's given a larger thinking budget, so the honest split is: use Flash and spend the savings on reasoning_effort when the task is reasoning-bound, and reach for Pro when it is knowledge-bound or agentically long. Against Gemini 3 Flash, the other obvious Efficient-tier long-context option, DeepSeek's differentiator is downloadable MIT weights and a 384K output ceiling; Google's is multimodal input and first-party cloud distribution. Against Claude Haiku 4.5, Flash wins on context and price and loses on ecosystem breadth and vision.
Getting startedThe smallest call that works
POST https://api.deepseek.com/chat/completions
Authorization: Bearer $DEEPSEEK_API_KEY
Content-Type: application/json
{
"model": "deepseek-v4-flash",
"messages": [
{"role": "user", "content": "Summarise this in one sentence: ..."}
],
"thinking": {"type": "disabled"}
}Change temperature to 0.0 first if a parser consumes the output — DeepSeek's own table recommends it for structured work. Then add "stream_options": {"include_usage": true} and check prompt_cache_hit_tokens; if it is zero, your system prompt is not stable and you are paying 50x more than you need to on input.
SourcesWhere every claim above came from
- Models & Pricing — DeepSeek API Docs
- DeepSeek V4 Preview Release — DeepSeek API Docs
- Change Log — DeepSeek API Docs
- Create Chat Completion — DeepSeek API Docs
- Temperature Settings — DeepSeek API Docs
- Context Caching — DeepSeek API Docs
- Function Calling — DeepSeek API Docs
- JSON Output — DeepSeek API Docs
- deepseek-ai/DeepSeek-V4-Flash — Hugging Face model card
- deepseek-ai/DeepSeek-V4-Flash config.json
- Foundry Models sold by Azure — Microsoft Foundry
- Models at a glance — Amazon Bedrock
- Could not confirm from a primary source as of 2026-07-25: the knowledge cutoff, whether thinking tokens bill differently from ordinary output tokens, whether several tool calls can be returned in one response, and the exact field names for the beta chat-prefix-completion feature. Pricing re-read from the vendor page on 2026-08-19: this model is now on a documented peak/off-peak schedule and the figures above match the vendor table exactly; no batch discount is published. The benchmark figures quoted here are DeepSeek's own claims from its model card, not independent evaluations.
Price and capacity verified 2026-09-05 against https://api-docs.deepseek.com/quick_start/pricing/. re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $0.44 in / $1.32 out per M, $0.014 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-09-04 (B2; the CITATION was corrected, not the price. The cited URL lacked a trailing slash and served "Your First API Call" with ZERO price-like numbers on it — a citation nothing could re-read, on the model 304 of 305 kits run on. With the slash it serves "Models & Pricing", and a live read that day found ALL TEN figures the two DeepSeek rows carry — peak, off-peak and cache-hit — present and unchanged: 0.007 0.014 0.022 0.044 0.22 0.44 0.66 1.32 1.98 3.96. No price moved. Prior: manual 2026-08-23 (currency proposal src-2026-08-23-8eb120; every peak and off-peak figure unchanged; the peak window carries a 'Monday through Friday' qualifier this row did not record).
What changedWhat changed here
Updated this page A reported DeepSeek 4.1 Flash cuts memory use 4x versus DeepSeek 4.0 Flash, which would lower self-hosting cost and allow longer contexts on the same hardware.
Update the DeepSeek V4 Flash page to note the reported DeepSeek 4.1 Flash release and its claimed 4x memory reduction, since the page currently presents V4 Flash as the current Flash-class entry.
Updated this page A 552B-parameter open-weight MoE model, DeepSeek-V4.1-Flash, is available on Alibaba Cloud at Flash pricing.
Add DeepSeek-V4.1-Flash to the DeepSeek V4 Flash page as a 552B-parameter MoE successor available on Alibaba Cloud, and note what it means for the self-host-versus-API decision.
- DeepSeek Doubles Annual Revenue Run Rate to $1 Billion Ahead of IPO
DeepSeek's annualized revenue run rate doubled to $1 billion ahead of a planned IPO, with a $7.5 billion funding round reportedly underway. That's a signal the low-cost Chinese model provider is becoming a durable commercial player rather than a price-disrupting flash — relevant if you're betting on DeepSeek as a long-term dependency.
- 'Better than DeepSeek': Xiaomi's MiMo-V2.6-Pro debuts as the top open weights model in the world alongside cheaper V2.6-Flash
Xiaomi's MiMo-V2.6-Pro debuts as the top open-weights model in the world, with a cheaper V2.6-Flash variant alongside it. A new open-weights leader changes what you can self-host or fine-tune without paying frontier API rates, so it's worth re-benchmarking your model choice.
- New DeepSeek 4.1 Flash Cuts Memory 4X vs DeepSeek 4.0 Flash
A new DeepSeek 4.1 Flash reportedly cuts memory use 4x versus DeepSeek 4.0 Flash. Lower memory per request means cheaper self-hosting and longer contexts on the same hardware — worth re-benchmarking if you currently serve a Flash-class model.
- DeepSeek-V4.1-Flash Ships Causal Encoder–Decoder MoE With 1M Context and Extreme KV Compression
DeepSeek shipped V4.1 Flash, a causal encoder–decoder MoE with a 1M-token context window and aggressive KV-cache compression. If you're architecting long-document or long-session agents, this is a cheap way to stop chunking and retrieval gymnastics for context that now fits in one call.
- DeepSeek-V4.1-Flash Outpaces GPT-5.6 Sol on Some Agentic, Coding Tests
DeepSeek-V4.1-Flash reportedly beats GPT-5.6 Sol on some agentic and coding tests, which is the kind of result that should make you re-test your model routing rather than assume the US frontier model wins on your task. Run your own eval before locking in a provider.
- DeepSeek Launches V4.1-Flash With Lower Memory and API Costs
DeepSeek launched V4.1-Flash with lower memory footprint and lower API costs, and separate coverage has it outpacing GPT-5.6 Sol on some agentic and coding tests. For cost-sensitive agent or coding workloads, this is a cheaper option worth benchmarking against your current model before assuming the frontier US models are the only viable choice.
- DeepSeek Unifies Three Modes: V4.1 Flash Launches Today at Fastest Speed, V4 Pro Will Be Discontinued
DeepSeek unified its three serving modes into V4.1 Flash, launched today as its fastest tier, and is discontinuing V4 Pro. If you route traffic to DeepSeek, you need to re-point any V4 Pro calls before it goes away.
- DeepSeek V4.1 Flash Prices Cached Input To $0.003 Per Million Tokens
DeepSeek V4.1 Flash prices cached input at $0.003 per million tokens, a rate that makes prompt-caching architectures dramatically cheaper for high-volume, repetitive-context workloads. Worth re-running your cost model if you currently avoid caching because of price.
- DeepSeek-V4.1-Flash Packs 552B Parameters With Efficient MoE Inference
DeepSeek-V4.1-Flash is a 552B-parameter MoE model built for efficient inference, which is the technical reason the Flash tier can undercut the Pro tier. Worth studying as a case of MoE serving economics if you're choosing open-weight models to self-host.
Showing the 10 most recent references. 10 older were dropped — a reference ages, so this list does not grow forever.
Three kinds of claim, strongest first. Signal runs every morning.