Home › Frontier Models › Gemini 3 Flash
Model · Reference

Gemini 3 Flash

High-volume long-context

In one line

Same 1M window, 64K output and four input modalities as Gemini 3.1 Pro at a quarter the input price, and with no long-context surcharge at all.

Why this oneWhat it is actually for

Gemini 3 Flash is the one to reach for when the job is high-volume and the context is still large. It carries the same 1M-token window and the same 64K output ceiling as Gemini 3.1 Pro, and it takes the same four input modalities, but text input lists at 0.50 dollars per million against Pro at 2.00 — and, the part people miss, it has no long-context surcharge. Pro doubles its rate on every token in a request once the prompt passes 200K; Flash charges the same 0.50 whether the prompt is 5K or 900K. If your workload is document extraction, transcript triage or repo-wide search over big inputs, Flash is often cheaper than Pro by considerably more than the headline 4x.

What it isArchitecture, lineage, training

Gemini 3 Flash is a closed-weight, natively multimodal reasoning model. Its model card, published December 2025, is explicit about lineage in a way that is unusually useful: Gemini 3 Flash is based on Gemini 3 Pro, and it is described as built off the Gemini 3 Pro reasoning foundation with thinking levels to control the mix of quality, cost and latency. Architecture, training dataset, training data processing and known limitations are all delegated to the Gemini 3 Pro model card rather than restated.

Following that pointer: Gemini 3 Pro is a sparse mixture-of-experts transformer with native multimodal support for text, vision and audio inputs, trained on Google TPUs over a large-scale pre-training set of publicly available web documents, text, code, images, audio and video. Sparse MoE routes each input token to a subset of experts, which is what lets a family decouple model capacity from serving cost per token — and is the most plausible reading of how a Flash tier sits under a Pro tier at a quarter the price. Note what the card does not say. It does not say Flash is a distillation of Pro, and it does not disclose any parameter count, expert count, active-parameter figure or layer count for either model. Flash was separately trained on TPUs per its own card.

The API model string is gemini-3-flash-preview. Context is 1M input and 64K output, and the knowledge cutoff is January 2025, both from the Gemini 3 developer guide's model table and the Flash model card. Google positions it for agentic workflows, everyday coding, reasoning and planning, and multimodal analysis, and claims it significantly outperforms Gemini 2.5 Pro on the December 2025 benchmark set — a comparison against the previous generation's flagship, not against Gemini 3 Pro.

At a glanceSee it

Gemini 3 Flash diagram

Gemini 3 Flash prices input by modality, not by prompt length, and batch halves it.

CapacityContext, output and what fits

FactValueSource
API model stringgemini-3-flash-previewGemini 3 Developer Guide model table
Launch stagePreviewGemini API models page
Context window1,000,000 input tokensGemini 3 Developer Guide; Gemini 3 Flash model card
Max output tokens64KGemini 3 Developer Guide; Gemini 3 Flash model card
Input modalitiesText, images, audio, video filesGemini 3 Flash model card
Output modalitiesText onlyGemini 3 Flash model card
Knowledge cutoffJanuary 2025Gemini 3 Developer Guide; inherited from the Gemini 3 Pro card
ArchitectureBased on Gemini 3 Pro, which is a sparse mixture-of-experts transformerGemini 3 Flash model card; Gemini 3 Pro model card
Parameter countNot disclosedNeither model card states one
Training hardwareGoogle TPUsGemini 3 Flash model card
Weights availableNo — closed API, no licence to self-hostModel card lists API distribution channels only
Default thinking levelhighGemini API thinking docs
Model card publishedDecember 2025Gemini 3 Flash model card
Image segmentationNot supportedGemini 3 Developer Guide

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
temperatureSampling randomness0.0 to 2.0, default 1.0Google strongly recommends leaving it at 1.0 for all Gemini 3 models; setting it below 1.0 may cause looping or degraded performance on complex maths and reasoning tasks.
topPNucleus sampling threshold0.0 to 1.0; no default published for this modelNot a recommended knob. Google states temperature, top_p and top_k are no longer recommended for all Gemini 3.x models and advises removing them from requests entirely.
topKKeeps only the k highest-probability tokensInteger; no range or default published for this modelSame guidance as topP — drop it rather than tune it. Use schema-constrained output when you need to narrow the response.
maxOutputTokensCaps the response lengthUp to 65,536 (documented as 64K)Truncates when hit. Thinking tokens are billed regardless.
stopSequencesStrings that halt generationArray of stringsThe stop string is not echoed back. Handy for batch extraction with delimiters.
candidateCountNumber of response variantsInteger; no default published for this model, and the reference flags the field as unsupported on some modelsWhere honoured, each extra candidate multiplies output cost. Rarely worth it on a cost-optimised tier, and worth confirming it is accepted at all.
seedFixes the sampling seedIntegerImproves repeatability across runs; not documented as bit-exact.
presencePenalty / frequencyPenaltyDiscourage repetitionFloats; present in GenerationConfig but flagged unsupported on some models, and support on Gemini 3 Flash could not be confirmedDo not build on them here. Prefer schema-constrained output.
thinkingConfig.thinkingLevelMaximum depth of internal reasoningminimal, low, medium, high; default highFlash is one of the few models that accepts minimal. This is where the cost savings actually are.
thinkingConfig.thinkingBudgetLegacy integer thinking budgetInteger; legacy, retained for backward compatibilityCannot appear in the same request as thinkingLevel — the guide states that returns a 400 error.
mediaResolutionToken budget per image, video frame or PDF pageGlobally in generationConfig: low, medium, high. ultra_high is available per content item only, not globally. Unspecified defaults to 1120 tokens per image, 560 per PDF page, 70 per video frame280 tokens at low, 1120 unspecified, 2240 at ultra_high — images only; the docs list ultra_high as N/A for video and PDF. On a cheap model this is often the largest single line in the bill.
responseMimeTypeForces the output MIME typee.g. application/jsonJSON without a schema gives you JSON shape but no field guarantees.
responseSchema / responseJsonSchemaConstrains output to a schemaSchema object; JSON Schema subset onlyLarge or deeply nested schemas may be rejected.
responseLogprobs / logprobsReturn token log-probabilitiesBoolean, plus integer count — support on this model could not be confirmedBoth fields exist in GenerationConfig, but logprobs is not listed among the supported capabilities for gemini-3-flash-preview, and it is reported as rejected on the Gemini 3 Pro line. Verify it works before using it to route low-confidence items up to Gemini 3.1 Pro.
toolConfig.functionCallingConfig.modeForces or forbids function callsAUTO, ANY, NONEANY forces a call — the fix for a model that answers in prose when you wanted structured tool use.
cachedContentPoints at explicitly cached contentCache resource namegenerateContent only. The caching docs state explicit caching — manually creating and managing cache objects — is not supported in the Interactions API, which does implicit caching only.

On Flash, thinkingLevel is the parameter that decides whether this model is cheap or not. It defaults to high, which is a frontier-model setting on an efficiency-tier price, and Flash accepts minimal, which Gemini 3.1 Pro does not. For classification, extraction and routing, minimal or low is usually correct and cuts both latency and billed tokens sharply. mediaResolution is the other one that matters, because at 0.50 per million input tokens a default-resolution image at 1120 tokens is a meaningful share of a small request. Leave the sampling knobs alone — Google now advises removing topP and topK from Gemini 3.x requests rather than tuning them.

SamplingShaping the output distribution

The sampling guidance is identical across the Gemini 3 family and it is a hard recommendation, not a hint: keep temperature at its default of 1.0. The developer guide warns that lowering it may cause looping or degraded performance, especially on complex mathematical or reasoning tasks. This surprises people migrating classification jobs from older models, where temperature 0 was standard practice for anything deterministic. The fix is to move determinism out of the sampler — set thinkingLevel to minimal or low, and constrain the output with responseSchema. Both topP and topK are accepted; the reference gives topP a 0.0 to 1.0 range and no default, and gives topK neither. A reasonable starting point for high-volume work is temperature 1.0, topP and topK unset, thinkingLevel at low, and a schema on the response. Then raise thinking only for the items that fail validation.

ReasoningThinking, effort and budgets

Yes, and unlike Gemini 3.1 Pro it can be effectively switched off. thinkingLevel accepts minimal, low, medium and high on Gemini 3 Flash, and the default is high — which means an untuned Flash call is doing frontier-depth reasoning on every request and billing you for it. The docs describe minimal as matching the no-thinking setting for most queries, with the caveat that the model may still think a little on complex coding tasks. The level is a ceiling on depth, not a fixed spend. Billing is unambiguous: response pricing is the sum of output tokens and thinking tokens, and you pay for the full thought tokens even though only a summary is ever returned to you. Summaries are opt-in. The legacy thinkingBudget integer still exists but cannot be combined with thinkingLevel in one request.

ToolsFunction calling and server tools

Function calling works the same way as on Gemini 3.1 Pro, including parallel calls for independent functions and compositional calls that chain on earlier results. Tools are declared with type: "function", name, description and parameters on the Interactions API; tool_choice takes auto, any, none and a preview validated mode. On legacy generateContent the equivalent is toolConfig.functionCallingConfig.mode with AUTO, ANY or NONE. Structured output uses response_format with mime_type and a schema drawn from a JSON Schema subset; very large or deeply nested schemas may be rejected. Server-side Google Search grounding is available and billed separately from tokens. The failure mode specific to Flash is worth naming: thinking is what makes tool selection reliable, so if you set thinkingLevel to minimal to save money on an agentic loop, expect tool-choice accuracy to drop. Image segmentation is explicitly not supported.

CostPrice, caching, batching, what drives the bill

Prices are from the Gemini Developer API pricing page, read on 2026-07-26, and agree with the model table in the Gemini 3 Developer Guide, which lists this model flat at 0.50 input and 3.00 output. The Gemini Enterprise Agent Platform (Vertex) pricing page could not be retrieved to cross-check. US dollars per million tokens. The structurally important fact is what is missing: Gemini 3 Flash has no long-context tier. Its pricing is quoted once, with no separate rate for prompts over 200K tokens, where Gemini 3.1 Pro doubles input to 4.00 and lifts output to 18.00. Audio is the modality that costs extra, at double the text rate on input.

LineText, image, videoAudio
Standard input0.501.00
Standard output3.00
Cached input0.050.10
Cache storage1.00 per 1M tokens per hour
Batch or flex input0.250.50
Batch or flex output1.50
Priority input0.901.80
Priority output5.40
Priority cached input0.090.18
Priority cache storage1.80 per 1M tokens per hour

Cached input is one tenth of standard input — a 90 percent saving computed from the published rates, not a figure Google states, since the caching page says only that cost savings are passed on automatically. Caching is implicit: on by default for Gemini 2.5 and newer, nothing to enable, with usage.total_cached_tokens confirming a hit. Batch and flex halve everything. The one that catches people is thinking: it defaults to high and thinking tokens bill at the 3.00 output rate, so an untuned Flash job can cost several times what the input price implied.

Where it runsSurfaces and availability

SurfaceAvailableNotes
Gemini Developer APIYesgemini-3-flash-preview, on both the Interactions API and legacy generateContent.
Google AI StudioYesListed on the Gemini API models page.
Google Vertex AI / Gemini Enterprise Agent PlatformYesAppears in the Agent Platform pricing table with identical per-token rates.
Batch APIYesBatch and flex tier at half the standard rate on both input and output.
Amazon BedrockNoNot among the distribution channels Google lists for Gemini models.
Microsoft Foundry / AzureNoNot among the distribution channels Google lists for Gemini models.
Hugging Face InferenceNoClosed weights; no Hugging Face model card exists.
Self-hostingNoWeights are not released under any licence.
Supervised fine-tuningUnverifiedCould not confirm gemini-3-flash-preview on the Vertex supervised-tuning supported-model list; the Vertex pricing page shows no SFT line for it.

StrengthsWhat it is good at

  • Flat input pricing across the whole 1M window — 0.50 per million whether the prompt is 5K or 900K tokens, per the Vertex pricing table. Gemini 3.1 Pro doubles above 200K.
  • Full frontier-family modality set: text, images, audio and video in, per its own model card, at efficiency-tier prices.
  • Accepts thinkingLevel: minimal, which Gemini 3.1 Pro does not — the cleanest way to buy Gemini 3 quality on a latency budget.
  • Batch and flex tier halves both input and output, taking text input to 0.25 and output to 1.50 per million.
  • Implicit caching at a 90 percent discount with no code change, tracked through usage.total_cached_tokens.

LimitsWhere it falls down

  • Thinking defaults to high. Leave it alone and you pay frontier-depth thinking tokens at 3.00 per million on a model you chose because it was cheap.
  • Audio input costs double text at 1.00 per million, so speech-heavy pipelines are not the bargain the headline rate implies.
  • Google's own benchmark claim on the model card is against Gemini 2.5 Pro, not against Gemini 3 Pro — there is no published primary-source margin between Flash and the current Pro tier.
  • Architecture is disclosed only as based on Gemini 3 Pro. No parameter count, no expert count, and no statement of whether it is a distillation.
  • Knowledge cutoff of January 2025 and a -preview model string, with image segmentation explicitly unsupported.

Against its neighboursHow it compares

Against Gemini 3.1 Pro the split is cleaner than the tier names suggest. Same window, same output cap, same modalities, same knowledge cutoff, same parameter surface — the differences are price, the absence of a long-context surcharge here, and Flash's extra minimal thinking level. Reach for Pro when the task is genuinely hard; reach for Flash when it is merely long. Against Claude Haiku 4.5 and GPT-5.6 Luna, Flash's argument is the input side: a million-token window and native audio and video at efficiency-tier pricing is not something the other efficiency models offer. Its argument against DeepSeek Flash is enterprise distribution through Vertex AI rather than raw price per token.

Getting startedThe smallest call that works

code
curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-3-flash-preview:generateContent" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H 'Content-Type: application/json' \
  -X POST \
  -d '{
    "contents": [{
      "parts": [{"text": "How does AI work?"}]
    }],
    "generationConfig": {
      "thinkingConfig": {
        "thinkingLevel": "minimal"
      }
    }
  }'

Set thinkingLevel deliberately on the very first call — the default is high and that is where an unexpected bill comes from. For extraction work add responseMimeType and responseSchema next, and set mediaResolution to low if you are sending images. Google's newer surface is POST /v1beta/interactions with model, input and snake_case generation_config.

SourcesWhere every claim above came from

Checked

Price and capacity verified 2026-09-18 against https://ai.google.dev/pricing. re-read 2026-09-18 (primary source read today: https://ai.google.dev/pricing — the page carries gemini-3-flash-preview only; no stable gemini-3-flash exists. Price UNCHANGED at $0.50 in / $3.00 out. The stable Flash tier a reader should compare against is now gemini-3.8-flash, which is its own row. Closes the 2026-09-03 id-drift reading.)

What changedWhat changed here

RecentAuto-linked from the brief, not a rewrite of this page
  • Gemini 3.8 TTS Playground 23 Sep · Simon Willison

    Google released gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts with a library of over 2,000 voices and custom voice cloning from a 30-second audio sample. Voice interfaces just got much cheaper to prototype — the open CORS policy means you can call it straight from a browser.

  • Gemini 3.8 text-to-speech says hello 23 Sep · Google DeepMind

    Google released two new Gemini text-to-speech models, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, with a library of over 2,000 voices and custom voice cloning from a 30-second audio sample. For anyone building voice agents, that's a drop-in TTS layer with a much wider voice palette and a cheap lite tier — worth benchmarking against your current provider.

  • DeepSeek V4.1 Flash Prices Cached Input To $0.003 Per Million Tokens 14 Sep · China frontier labs

    DeepSeek V4.1 Flash prices cached input at $0.003 per million tokens, a rate that makes prompt-caching architectures dramatically cheaper for high-volume, repetitive-context workloads. Worth re-running your cost model if you currently avoid caching because of price.

  • DeepSeek V4.1 Flash Debuts With 552B Parameters and 1M Token Context 13 Sep · China frontier labs

    DeepSeek shipped V4.1 Flash with 552B parameters and a 1M token context window — a very large open-weight-class model with long-context capability. If you're architecting retrieval or long-document pipelines, this is a candidate to benchmark against your current model before you commit to a paid API.

  • DeepSeek's cheaper V4.1 Flash Model fuels AI price war | Tap to know more | Inshorts 12 Sep · China frontier labs

    DeepSeek's cheaper V4.1 Flash is explicitly framed as fuelling an AI price war, following the earlier V4.1 Flash debut. Cheaper long-context inference changes the economics of agent loops and batch jobs — worth re-running your cost model rather than assuming last quarter's per-token rates.

  • DeepSeek V4.1 Flash: 552B parameters, 8B active, 60% cached-input cut — ties Opus 5 12 Sep · China frontier labs

    DeepSeek's V4.1 Flash is a 552B-parameter model with only 8B active and a 60% cut on cached input, reportedly tying Opus 5 — that combination of sparse activation and cache pricing is exactly the kind of cost curve you'd design a high-volume pipeline around.

  • Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context 11 Sep · China frontier labs

    Baseten added DeepSeek-V4.1-Flash to its model APIs with a 1M-token context window, so you can now call it through a managed endpoint rather than self-hosting — worth testing against your current provider on long-context workloads.

  • How to Use DeepSeek V4.1 Flash API: 12 Steps, $0.22/M [2026] 10 Sep · China frontier labs

    A practical walkthrough of the DeepSeek V4.1 Flash API at $0.22 per million tokens. At that rate, high-volume agent loops that were cost-prohibitive on US frontier models become viable — worth benchmarking against your current provider before your next architecture decision.

  • Google says its new Gemini 3.8 Flash model ‘works harder’ but might cost more 2 Sep · The Verge AI

    Google shipped Gemini 3.8 Flash weeks after its predecessor, pitched as working harder via more reasoning steps on complex tasks and iterative tool calls, at the same introductory pricing of $0.75 per million input tokens and $3.75 per million output tokens. A cheaper, tool-calling Flash tier arriving this fast changes the default for high-volume agent loops that previously had to justify a reason

  • llm-gemini 0.34 2 Sep · Simon Willison

    Gemini 3.8 Flash is out with low, medium, and high thinking levels, and Google also released a 3.8 Flash Cyber variant restricted to trusted defenders. That gives you a Flash-class model you can dial between quick answers and deeper reasoning — and signals Google is gating cyber-grade capabilities deliberately.

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning