Home › Frontier Models › Gemini 3.1 Pro
Model · Reference

Gemini 3.1 Pro

Long docs, multimodal, GCP

In one line

Google's frontier multimodal model with a 1M-token window — not 2M — where crossing 200K input tokens doubles the price of the entire request.

Why this oneWhat it is actually for

You reach for Gemini 3.1 Pro when the input is the hard part: a few hundred pages of PDF, an hour of video, a whole repository, or a mixed bundle of all three. It takes text, images, audio and video natively in one request and returns text, so you skip the transcribe-then-summarise pipeline entirely. The second reason is procurement — if your data already sits in Google Cloud, Vertex AI hands you the model without onboarding a new vendor. What it is not is the cheap default. At 2.00 dollars per million input tokens it costs four times Gemini 3 Flash, and pushing past 200K input tokens re-prices the whole request, output included, at 4.00 and 18.00.

What it isArchitecture, lineage, training

Gemini 3.1 Pro is a closed-weight, natively multimodal reasoning model. Its own model card, published February 2026, says only that it is based on Gemini 3 Pro and defers every architecture question to that card. The Gemini 3 Pro card is where the disclosure actually lives: Gemini 3 Pro is described as a sparse mixture-of-experts transformer with native multimodal support for text, vision and audio inputs. The card also explains why — sparse MoE models activate a subset of parameters per input token, which decouples total model capacity from compute and serving cost per token.

That is the entire architectural disclosure. Google publishes no total parameter count, no active parameter count, no expert count, no layer, head or vocabulary numbers for any Gemini 3 model. Training is described only in kind: a large-scale pre-training set of publicly available web documents, text, code, images, audio and video, followed by a post-training stage; the hardware is Google TPUs. The knowledge cutoff is January 2025, stated in the Gemini 3 Pro card and repeated in the Gemini 3 developer guide's model table.

Lineage matters here. Gemini 3 Pro shipped November 2025 and the card states it is not a modification or fine-tune of a prior model. Everything after it in the family — Gemini 3 Flash, Gemini 3.1 Pro, the Flash-Lite, Flash Live and Image variants — is built on that base. So Gemini 3.1 Pro reads as a refresh of the November base rather than a new pre-train, and Google's claim for it is comparative rather than absolute: it significantly outperforms Gemini 3 Pro across the February 2026 benchmark set. The API model string is gemini-3.1-pro-preview, and the preview suffix is the launch stage, not decoration.

At a glanceSee it

Gemini 3.1 Pro diagram

A Gemini 3.1 Pro request, from prompt through cache and thinking to the token bill.

CapacityContext, output and what fits

FactValueSource
API model stringgemini-3.1-pro-previewGemini 3 Developer Guide model table
Launch stagePreviewGemini API models page
Context window1,000,000 input tokensGemini 3 Developer Guide; Gemini 3.1 Pro model card
Max output tokens64KGemini 3 Developer Guide; Gemini 3.1 Pro model card
Input modalitiesText, images, audio, video filesGemini 3.1 Pro model card
Output modalitiesText onlyGemini 3.1 Pro model card
Knowledge cutoffJanuary 2025Gemini 3 Pro model card; Gemini 3 Developer Guide
ArchitectureSparse mixture-of-experts transformer, natively multimodalGemini 3 Pro model card
Parameter countNot disclosedNeither model card states one
Expert count / active paramsNot disclosedNeither model card states one
Training hardwareGoogle TPUsGemini 3 Pro model card
Weights availableNo — closed API, no licence to self-hostModel card distribution section lists API channels only
Minimum tokens for implicit cache hit4,096Gemini API context caching docs
Model card publishedFebruary 2026Gemini 3.1 Pro model card

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
temperatureSampling randomness0.0 to 2.0, default 1.0Google explicitly warns against changing it. The Gemini 3 guide strongly recommends keeping it at 1.0 and says setting it below 1.0 may cause looping or degraded performance on complex maths and reasoning tasks.
topPNucleus sampling threshold0.0 to 1.0; no default published for this modelNot a recommended knob. Google states that temperature, top_p and top_k are no longer recommended for all Gemini 3.x models and advises removing them from requests entirely rather than tuning them.
topKKeeps only the k highest-probability tokensInteger; no range or default published for this modelSame guidance as topP — drop it rather than tune it. If you need narrower output, constrain it with responseSchema.
maxOutputTokensCaps the response lengthUp to 65,536 (documented as 64K)Truncates mid-sentence when hit. Note thinking tokens count toward what you are billed even if they are not the visible output.
stopSequencesStrings that halt generationArray of stringsThe stop string itself is not returned. Useful for delimiter-framed output.
candidateCountNumber of response variantsInteger; no default published for this model, and the reference flags the field as unsupported on some modelsWhere it is honoured, each extra candidate multiplies output token cost. Confirm the model accepts it before designing around it.
seedFixes the sampling seedIntegerImproves run-to-run repeatability. The docs do not promise bit-exact determinism.
presencePenalty / frequencyPenaltyDiscourage repeated tokensFloats; present in GenerationConfig but flagged unsupported on some models, and support on Gemini 3.1 Pro could not be confirmedDo not build on them here. Prefer prompt-level instruction or a response schema.
thinkingConfig.thinkingLevelMaximum depth of internal reasoninglow, medium, high; default high. minimal is not accepted on Pro — it is a Flash and Flash-Lite valueThe single biggest lever on latency, cost and quality. high can take significantly longer to reach the first non-thinking output token.
thinkingConfig.thinkingBudgetLegacy token budget for thinkingInteger; legacy, retained for backward compatibilityCannot be used in the same request as thinkingLevel — the guide states that doing so returns a 400 error.
mediaResolutionHow much token budget each image, video frame or PDF page getsGlobally in generationConfig: low, medium, high. ultra_high is available per content item only, not globally. Unspecified defaults to 1120 tokens per image, 560 per PDF page, 70 per video frameDirectly changes input cost — an image is 280 tokens at low, 1120 unspecified, and 2240 at ultra_high. ultra_high is documented for images only; the table lists it as N/A for video and PDF.
responseMimeTypeForces the output MIME typee.g. application/jsonSet it to JSON without a schema and you get JSON-shaped but unconstrained output.
responseSchema / responseJsonSchemaConstrains output to a schemaSchema object; JSON Schema subset onlyVery large or deeply nested schemas may be rejected outright.
responseLogprobs / logprobsReturn token log-probabilitiesBoolean, plus integer count — support on this model could not be confirmedBoth fields exist in GenerationConfig, but Google does not list logprobs among the supported capabilities for gemini-3.1-pro-preview, and developers report the call rejected with "Logprobs is not supported for this model". Test it before building confidence gating or an eval harness on it.
toolConfig.functionCallingConfig.modeForces or forbids function callsAUTO, ANY, NONEANY guarantees a call and is how you stop the model answering in prose when you wanted a tool.
cachedContentPoints the request at explicitly cached contentCache resource namegenerateContent only. The caching docs state that explicit caching — manually creating and managing cache objects — is not supported in the Interactions API, which does implicit caching only.

One knob carries almost all the weight. thinkingLevel is the real dial — it moves latency, cost and answer quality together, and dropping from high to low is the first thing to try when a Pro call feels slow. mediaResolution is the second, because on multimodal workloads image and PDF tokens dominate the input bill and the unspecified default sits at the expensive end for images. The sampling knobs are the ones to deliberately not touch: the guide tells you to leave temperature at 1.0, and says topP and topK are no longer recommended on any Gemini 3.x model and should be removed from requests. In the Interactions API the same fields appear as generation_config.thinking_level and per-item resolution.

SamplingShaping the output distribution

Google's guidance for Gemini 3 is unusually blunt: leave temperature at its default of 1.0. The developer guide states that lowering it may lead to unexpected behaviour such as looping or degraded performance, particularly on complex mathematical or reasoning tasks. That is the opposite of the habit most people bring from earlier models, where dropping to 0.2 was the reflex for anything factual. The reason is structural — this model does its convergence in the thinking phase, not in the sampler, so squeezing the distribution afterwards fights the mechanism rather than helping it. topP and topK are both accepted; the reference gives topP a 0.0 to 1.0 range and states no default, and gives topK no range at all. A sensible starting point is therefore: temperature 1.0, topP and topK unset, and reach for thinkingLevel when you want more determinism in substance. If you genuinely need tighter output, constrain the shape with responseSchema rather than the sampler.

ReasoningThinking, effort and budgets

Yes, and it is on by default. Gemini 3.1 Pro's default thinkingLevel is high. The parameter takes low, medium or high on Pro — minimal exists but is documented for Gemini 3 Flash and Flash-Lite only. It is a maximum depth, not a fixed spend, so a simple prompt at high will not necessarily burn a full budget. The legacy thinkingBudget integer still exists but cannot appear in the same request as thinkingLevel. Billing is explicit in the docs: when thinking is on, response pricing is the sum of output tokens and thinking tokens, and you are charged for the full thought tokens even though only a summary is ever returned. Summaries are opt-in via thinking_summaries set to auto in the Interactions API. Across turns, Gemini 3 uses encrypted thought signatures to carry reasoning context; the SDKs handle those for you.

ToolsFunction calling and server tools

Function calling is supported, with both parallel calls — several independent functions in one turn — and compositional calls, where the model chains a second call using the first one's result. In the Interactions API a tool is declared with type: "function" plus name, description and parameters; selection is controlled by tool_choice with modes auto, any, none and a preview validated mode that enforces schema adherence. On the legacy generateContent surface the equivalent is toolConfig.functionCallingConfig.mode with AUTO, ANY or NONE. Structured output goes through response_format with a mime_type and a schema, and only a subset of JSON Schema is supported — large or deeply nested schemas may be rejected. Server-side tools include Google Search grounding, which is billed separately from tokens. Two failure modes bite people: dropping thought signatures when you hand-roll multi-turn tool loops instead of using an SDK, and schemas that are accepted but silently under-enforced.

CostPrice, caching, batching, what drives the bill

Prices below are from the Gemini Developer API pricing page, read on 2026-07-26, and agree with the model table in the Gemini 3 Developer Guide. The Gemini Enterprise Agent Platform (Vertex) pricing page could not be retrieved to cross-check, so treat the Developer API page as the source. All figures are US dollars per million tokens. The structure that actually drives your bill is the 200K threshold: every line is quoted twice, once for prompts up to 200K tokens and once for prompts over 200K, and the output rate is selected by the size of the prompt, not the size of the response. So a 210K-token prompt reprices the entire request. It is not a uniform doubling, though: input and cached input double, while output rises by half — 12.00 to 18.00.

LinePrompt up to 200KPrompt over 200K
Standard input2.004.00
Standard output12.0018.00
Cached input0.200.40
Cache storage4.50 per 1M tokens per hour
Batch input1.002.00
Batch output6.009.00
Flex input1.002.00
Flex output6.009.00
Priority input3.607.20
Priority output21.6032.40
Priority cached input0.360.72
Priority cache storage8.10 per 1M tokens per hour

Cached input is one tenth of standard input at both prompt sizes — a 90 percent saving on hit, computed from the published rates rather than quoted as a discount by Google, whose caching page says only that cost savings are passed on automatically. Caching is implicit: on by default for Gemini 2.5 and newer, no API change needed, with a 4,096-token minimum for this model and a usage.total_cached_tokens field to check it worked. Batch and flex both halve input and output at either prompt size. Thinking tokens are billed at the output rate. Grounding with Google Search is billed separately from tokens.

Where it runsSurfaces and availability

SurfaceAvailableNotes
Gemini Developer APIYesgemini-3.1-pro-preview on generativelanguage.googleapis.com, both the Interactions API and legacy generateContent.
Google AI StudioYesListed in the model card distribution section.
Google Vertex AI / Gemini Enterprise Agent PlatformYesSame per-token pricing as the Developer API, including the 200K tier.
Gemini app, NotebookLM, Antigravity, Gemini EnterpriseYesNamed as distribution channels on the model card.
Amazon BedrockNoNot among the distribution channels the model card lists.
Microsoft Foundry / AzureNoNot among the distribution channels the model card lists.
Hugging Face InferenceNoClosed weights; no Hugging Face model card exists.
Self-hostingNoWeights are not released under any licence.
Supervised fine-tuningUnverifiedCould not confirm gemini-3.1-pro-preview on the Vertex supervised-tuning supported-model list; the Vertex pricing page shows no SFT line for it.

StrengthsWhat it is good at

  • A genuine 1M-token input window with text, image, audio and video accepted natively in the same request — per the Gemini 3.1 Pro model card, which also caps output at 64K.
  • Thinking depth is a first-class, three-valued dial rather than an all-or-nothing switch, so you can trade latency against quality per call using thinkingLevel.
  • Implicit caching needs no code change and lands a 90 percent discount on cached input above a 4,096-token minimum.
  • Per-media-item mediaResolution lets you spend 280 tokens or 2240 tokens on an image deliberately, which is real cost control on multimodal pipelines.
  • Identical pricing on the Developer API and Vertex AI, so moving from prototype to a Google Cloud deployment does not change the unit economics.

LimitsWhere it falls down

  • The 200K threshold is a cliff, not a ramp: exceed it and the whole request, output included, bills at 4.00 and 18.00. Long-context work costs double what the headline suggests.
  • Architecture disclosure stops at the phrase sparse mixture-of-experts. No parameter count, expert count or active-parameter figure is published, so you cannot reason about it from first principles.
  • Knowledge cutoff of January 2025 is old for a model card published in February 2026 — anything recent needs grounding or retrieval.
  • Output is text only. There is no image or audio output from this model, and image segmentation is explicitly unsupported on Gemini 3 Pro and Gemini 3 Flash.
  • Still a -preview model string, and the newer Interactions API drops explicit caching entirely — only implicit caching works there.

Against its neighboursHow it compares

Against Gemini 3 Flash, the honest question is whether you need the reasoning at all: Flash has the same 1M window, the same 64K output and the same modalities at a quarter the input price, and — critically — no long-context surcharge. On big prompts Flash is more than 4x cheaper, not 4x. Take Pro when the task is hard, not merely long. Against GPT-5.6 Sol and Claude Opus 5, Gemini's differentiator is the input side: native audio and video in one call, and a window measured in millions rather than hundreds of thousands. Its weakness against both is disclosure and stability — a preview model string, a January 2025 cutoff, and a pricing model that punishes exactly the long-context workload it is sold on.

Getting startedThe smallest call that works

code
curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-pro-preview:generateContent" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H 'Content-Type: application/json' \
  -X POST \
  -d '{
    "contents": [{
      "parts": [{"text": "How does AI work?"}]
    }],
    "generationConfig": {
      "thinkingConfig": {
        "thinkingLevel": "low"
      }
    }
  }'

Change thinkingLevel first — this example uses low, but the model default is high, and that gap is most of your latency. Leave temperature alone. If you are sending images or PDFs, set mediaResolution before you tune anything else. Google's newer surface is POST /v1beta/interactions, which takes model, input and generation_config with snake_case field names instead.

SourcesWhere every claim above came from

Checked

Price and capacity verified 2026-09-18 against https://ai.google.dev/pricing. re-read 2026-09-18 (primary source read today: https://ai.google.dev/pricing — $2.00 per million input and $12.00 per million output at the <=200k tier, UNCHANGED. The page still labels it gemini-3.1-pro-preview, a preview channel and not a stable id, which this row's own note has recorded since 2026-08-18. Closes the 2026-09-03 id-drift reading.)

What changedWhat changed here

RecentAuto-linked from the brief, not a rewrite of this page
  • Anthropic merges Claude chat and Cowork in one interface 16 Sep · TechCrunch AI

    Anthropic merged Claude chat and Cowork into a single interface, with Docs and Slides tools added, rolling out to Pro and Max subscribers first. If you're building on Claude, the surface you integrate against is consolidating into one general agent rather than separate chat and work products — worth re-checking your assumptions about which endpoints and plan tiers expose what.

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning