Google's frontier multimodal model with a 1M-token window — not 2M — where crossing 200K input tokens doubles the price of the entire request.
Why this oneWhat it is actually for
You reach for Gemini 3.1 Pro when the input is the hard part: a few hundred pages of PDF, an hour of video, a whole repository, or a mixed bundle of all three. It takes text, images, audio and video natively in one request and returns text, so you skip the transcribe-then-summarise pipeline entirely. The second reason is procurement — if your data already sits in Google Cloud, Vertex AI hands you the model without onboarding a new vendor. What it is not is the cheap default. At 2.00 dollars per million input tokens it costs four times Gemini 3 Flash, and pushing past 200K input tokens re-prices the whole request, output included, at 4.00 and 18.00.
What it isArchitecture, lineage, training
Gemini 3.1 Pro is a closed-weight, natively multimodal reasoning model. Its own model card, published February 2026, says only that it is based on Gemini 3 Pro and defers every architecture question to that card. The Gemini 3 Pro card is where the disclosure actually lives: Gemini 3 Pro is described as a sparse mixture-of-experts transformer with native multimodal support for text, vision and audio inputs. The card also explains why — sparse MoE models activate a subset of parameters per input token, which decouples total model capacity from compute and serving cost per token.
That is the entire architectural disclosure. Google publishes no total parameter count, no active parameter count, no expert count, no layer, head or vocabulary numbers for any Gemini 3 model. Training is described only in kind: a large-scale pre-training set of publicly available web documents, text, code, images, audio and video, followed by a post-training stage; the hardware is Google TPUs. The knowledge cutoff is January 2025, stated in the Gemini 3 Pro card and repeated in the Gemini 3 developer guide's model table.
Lineage matters here. Gemini 3 Pro shipped November 2025 and the card states it is not a modification or fine-tune of a prior model. Everything after it in the family — Gemini 3 Flash, Gemini 3.1 Pro, the Flash-Lite, Flash Live and Image variants — is built on that base. So Gemini 3.1 Pro reads as a refresh of the November base rather than a new pre-train, and Google's claim for it is comparative rather than absolute: it significantly outperforms Gemini 3 Pro across the February 2026 benchmark set. The API model string is gemini-3.1-pro-preview, and the preview suffix is the launch stage, not decoration.
At a glanceSee it
A Gemini 3.1 Pro request, from prompt through cache and thinking to the token bill.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| API model string | gemini-3.1-pro-preview | Gemini 3 Developer Guide model table |
| Launch stage | Preview | Gemini API models page |
| Context window | 1,000,000 input tokens | Gemini 3 Developer Guide; Gemini 3.1 Pro model card |
| Max output tokens | 64K | Gemini 3 Developer Guide; Gemini 3.1 Pro model card |
| Input modalities | Text, images, audio, video files | Gemini 3.1 Pro model card |
| Output modalities | Text only | Gemini 3.1 Pro model card |
| Knowledge cutoff | January 2025 | Gemini 3 Pro model card; Gemini 3 Developer Guide |
| Architecture | Sparse mixture-of-experts transformer, natively multimodal | Gemini 3 Pro model card |
| Parameter count | Not disclosed | Neither model card states one |
| Expert count / active params | Not disclosed | Neither model card states one |
| Training hardware | Google TPUs | Gemini 3 Pro model card |
| Weights available | No — closed API, no licence to self-host | Model card distribution section lists API channels only |
| Minimum tokens for implicit cache hit | 4,096 | Gemini API context caching docs |
| Model card published | February 2026 | Gemini 3.1 Pro model card |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
temperature | Sampling randomness | 0.0 to 2.0, default 1.0 | Google explicitly warns against changing it. The Gemini 3 guide strongly recommends keeping it at 1.0 and says setting it below 1.0 may cause looping or degraded performance on complex maths and reasoning tasks. |
topP | Nucleus sampling threshold | 0.0 to 1.0; no default published for this model | Not a recommended knob. Google states that temperature, top_p and top_k are no longer recommended for all Gemini 3.x models and advises removing them from requests entirely rather than tuning them. |
topK | Keeps only the k highest-probability tokens | Integer; no range or default published for this model | Same guidance as topP — drop it rather than tune it. If you need narrower output, constrain it with responseSchema. |
maxOutputTokens | Caps the response length | Up to 65,536 (documented as 64K) | Truncates mid-sentence when hit. Note thinking tokens count toward what you are billed even if they are not the visible output. |
stopSequences | Strings that halt generation | Array of strings | The stop string itself is not returned. Useful for delimiter-framed output. |
candidateCount | Number of response variants | Integer; no default published for this model, and the reference flags the field as unsupported on some models | Where it is honoured, each extra candidate multiplies output token cost. Confirm the model accepts it before designing around it. |
seed | Fixes the sampling seed | Integer | Improves run-to-run repeatability. The docs do not promise bit-exact determinism. |
presencePenalty / frequencyPenalty | Discourage repeated tokens | Floats; present in GenerationConfig but flagged unsupported on some models, and support on Gemini 3.1 Pro could not be confirmed | Do not build on them here. Prefer prompt-level instruction or a response schema. |
thinkingConfig.thinkingLevel | Maximum depth of internal reasoning | low, medium, high; default high. minimal is not accepted on Pro — it is a Flash and Flash-Lite value | The single biggest lever on latency, cost and quality. high can take significantly longer to reach the first non-thinking output token. |
thinkingConfig.thinkingBudget | Legacy token budget for thinking | Integer; legacy, retained for backward compatibility | Cannot be used in the same request as thinkingLevel — the guide states that doing so returns a 400 error. |
mediaResolution | How much token budget each image, video frame or PDF page gets | Globally in generationConfig: low, medium, high. ultra_high is available per content item only, not globally. Unspecified defaults to 1120 tokens per image, 560 per PDF page, 70 per video frame | Directly changes input cost — an image is 280 tokens at low, 1120 unspecified, and 2240 at ultra_high. ultra_high is documented for images only; the table lists it as N/A for video and PDF. |
responseMimeType | Forces the output MIME type | e.g. application/json | Set it to JSON without a schema and you get JSON-shaped but unconstrained output. |
responseSchema / responseJsonSchema | Constrains output to a schema | Schema object; JSON Schema subset only | Very large or deeply nested schemas may be rejected outright. |
responseLogprobs / logprobs | Return token log-probabilities | Boolean, plus integer count — support on this model could not be confirmed | Both fields exist in GenerationConfig, but Google does not list logprobs among the supported capabilities for gemini-3.1-pro-preview, and developers report the call rejected with "Logprobs is not supported for this model". Test it before building confidence gating or an eval harness on it. |
toolConfig.functionCallingConfig.mode | Forces or forbids function calls | AUTO, ANY, NONE | ANY guarantees a call and is how you stop the model answering in prose when you wanted a tool. |
cachedContent | Points the request at explicitly cached content | Cache resource name | generateContent only. The caching docs state that explicit caching — manually creating and managing cache objects — is not supported in the Interactions API, which does implicit caching only. |
One knob carries almost all the weight. thinkingLevel is the real dial — it moves latency, cost and answer quality together, and dropping from high to low is the first thing to try when a Pro call feels slow. mediaResolution is the second, because on multimodal workloads image and PDF tokens dominate the input bill and the unspecified default sits at the expensive end for images. The sampling knobs are the ones to deliberately not touch: the guide tells you to leave temperature at 1.0, and says topP and topK are no longer recommended on any Gemini 3.x model and should be removed from requests. In the Interactions API the same fields appear as generation_config.thinking_level and per-item resolution.
SamplingShaping the output distribution
Google's guidance for Gemini 3 is unusually blunt: leave temperature at its default of 1.0. The developer guide states that lowering it may lead to unexpected behaviour such as looping or degraded performance, particularly on complex mathematical or reasoning tasks. That is the opposite of the habit most people bring from earlier models, where dropping to 0.2 was the reflex for anything factual. The reason is structural — this model does its convergence in the thinking phase, not in the sampler, so squeezing the distribution afterwards fights the mechanism rather than helping it. topP and topK are both accepted; the reference gives topP a 0.0 to 1.0 range and states no default, and gives topK no range at all. A sensible starting point is therefore: temperature 1.0, topP and topK unset, and reach for thinkingLevel when you want more determinism in substance. If you genuinely need tighter output, constrain the shape with responseSchema rather than the sampler.
ReasoningThinking, effort and budgets
Yes, and it is on by default. Gemini 3.1 Pro's default thinkingLevel is high. The parameter takes low, medium or high on Pro — minimal exists but is documented for Gemini 3 Flash and Flash-Lite only. It is a maximum depth, not a fixed spend, so a simple prompt at high will not necessarily burn a full budget. The legacy thinkingBudget integer still exists but cannot appear in the same request as thinkingLevel. Billing is explicit in the docs: when thinking is on, response pricing is the sum of output tokens and thinking tokens, and you are charged for the full thought tokens even though only a summary is ever returned. Summaries are opt-in via thinking_summaries set to auto in the Interactions API. Across turns, Gemini 3 uses encrypted thought signatures to carry reasoning context; the SDKs handle those for you.
ToolsFunction calling and server tools
Function calling is supported, with both parallel calls — several independent functions in one turn — and compositional calls, where the model chains a second call using the first one's result. In the Interactions API a tool is declared with type: "function" plus name, description and parameters; selection is controlled by tool_choice with modes auto, any, none and a preview validated mode that enforces schema adherence. On the legacy generateContent surface the equivalent is toolConfig.functionCallingConfig.mode with AUTO, ANY or NONE. Structured output goes through response_format with a mime_type and a schema, and only a subset of JSON Schema is supported — large or deeply nested schemas may be rejected. Server-side tools include Google Search grounding, which is billed separately from tokens. Two failure modes bite people: dropping thought signatures when you hand-roll multi-turn tool loops instead of using an SDK, and schemas that are accepted but silently under-enforced.
CostPrice, caching, batching, what drives the bill
Prices below are from the Gemini Developer API pricing page, read on 2026-07-26, and agree with the model table in the Gemini 3 Developer Guide. The Gemini Enterprise Agent Platform (Vertex) pricing page could not be retrieved to cross-check, so treat the Developer API page as the source. All figures are US dollars per million tokens. The structure that actually drives your bill is the 200K threshold: every line is quoted twice, once for prompts up to 200K tokens and once for prompts over 200K, and the output rate is selected by the size of the prompt, not the size of the response. So a 210K-token prompt reprices the entire request. It is not a uniform doubling, though: input and cached input double, while output rises by half — 12.00 to 18.00.
| Line | Prompt up to 200K | Prompt over 200K |
|---|---|---|
| Standard input | 2.00 | 4.00 |
| Standard output | 12.00 | 18.00 |
| Cached input | 0.20 | 0.40 |
| Cache storage | 4.50 per 1M tokens per hour | |
| Batch input | 1.00 | 2.00 |
| Batch output | 6.00 | 9.00 |
| Flex input | 1.00 | 2.00 |
| Flex output | 6.00 | 9.00 |
| Priority input | 3.60 | 7.20 |
| Priority output | 21.60 | 32.40 |
| Priority cached input | 0.36 | 0.72 |
| Priority cache storage | 8.10 per 1M tokens per hour | |
Cached input is one tenth of standard input at both prompt sizes — a 90 percent saving on hit, computed from the published rates rather than quoted as a discount by Google, whose caching page says only that cost savings are passed on automatically. Caching is implicit: on by default for Gemini 2.5 and newer, no API change needed, with a 4,096-token minimum for this model and a usage.total_cached_tokens field to check it worked. Batch and flex both halve input and output at either prompt size. Thinking tokens are billed at the output rate. Grounding with Google Search is billed separately from tokens.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Gemini Developer API | Yes | gemini-3.1-pro-preview on generativelanguage.googleapis.com, both the Interactions API and legacy generateContent. |
| Google AI Studio | Yes | Listed in the model card distribution section. |
| Google Vertex AI / Gemini Enterprise Agent Platform | Yes | Same per-token pricing as the Developer API, including the 200K tier. |
| Gemini app, NotebookLM, Antigravity, Gemini Enterprise | Yes | Named as distribution channels on the model card. |
| Amazon Bedrock | No | Not among the distribution channels the model card lists. |
| Microsoft Foundry / Azure | No | Not among the distribution channels the model card lists. |
| Hugging Face Inference | No | Closed weights; no Hugging Face model card exists. |
| Self-hosting | No | Weights are not released under any licence. |
| Supervised fine-tuning | Unverified | Could not confirm gemini-3.1-pro-preview on the Vertex supervised-tuning supported-model list; the Vertex pricing page shows no SFT line for it. |
StrengthsWhat it is good at
- A genuine 1M-token input window with text, image, audio and video accepted natively in the same request — per the Gemini 3.1 Pro model card, which also caps output at 64K.
- Thinking depth is a first-class, three-valued dial rather than an all-or-nothing switch, so you can trade latency against quality per call using
thinkingLevel. - Implicit caching needs no code change and lands a 90 percent discount on cached input above a 4,096-token minimum.
- Per-media-item
mediaResolutionlets you spend 280 tokens or 2240 tokens on an image deliberately, which is real cost control on multimodal pipelines. - Identical pricing on the Developer API and Vertex AI, so moving from prototype to a Google Cloud deployment does not change the unit economics.
LimitsWhere it falls down
- The 200K threshold is a cliff, not a ramp: exceed it and the whole request, output included, bills at 4.00 and 18.00. Long-context work costs double what the headline suggests.
- Architecture disclosure stops at the phrase sparse mixture-of-experts. No parameter count, expert count or active-parameter figure is published, so you cannot reason about it from first principles.
- Knowledge cutoff of January 2025 is old for a model card published in February 2026 — anything recent needs grounding or retrieval.
- Output is text only. There is no image or audio output from this model, and image segmentation is explicitly unsupported on Gemini 3 Pro and Gemini 3 Flash.
- Still a
-previewmodel string, and the newer Interactions API drops explicit caching entirely — only implicit caching works there.
Against its neighboursHow it compares
Against Gemini 3 Flash, the honest question is whether you need the reasoning at all: Flash has the same 1M window, the same 64K output and the same modalities at a quarter the input price, and — critically — no long-context surcharge. On big prompts Flash is more than 4x cheaper, not 4x. Take Pro when the task is hard, not merely long. Against GPT-5.6 Sol and Claude Opus 5, Gemini's differentiator is the input side: native audio and video in one call, and a window measured in millions rather than hundreds of thousands. Its weakness against both is disclosure and stability — a preview model string, a January 2025 cutoff, and a pricing model that punishes exactly the long-context workload it is sold on.
Getting startedThe smallest call that works
curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-pro-preview:generateContent" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H 'Content-Type: application/json' \
-X POST \
-d '{
"contents": [{
"parts": [{"text": "How does AI work?"}]
}],
"generationConfig": {
"thinkingConfig": {
"thinkingLevel": "low"
}
}
}'Change thinkingLevel first — this example uses low, but the model default is high, and that gap is most of your latency. Leave temperature alone. If you are sending images or PDFs, set mediaResolution before you tune anything else. Google's newer surface is POST /v1beta/interactions, which takes model, input and generation_config with snake_case field names instead.
SourcesWhere every claim above came from
- Gemini 3 Developer Guide — Gemini API
- Gemini 3.1 Pro — Model Card, Google DeepMind
- Gemini 3 Pro Model Card, PDF
- Gemini Developer API pricing
- Agent Platform pricing, Google Cloud
- Thinking — Gemini API
- Context caching — Gemini API
- Function calling — Gemini API
- Structured output — Gemini API
- Media resolution — Gemini API
- GenerateContent API reference — GenerationConfig
- Gemini 3 Developer Guide — legacy generateContent
- Gemini 3 Developer Guide — Interactions API
- Could not confirm from a primary source as of 2026-07-25: any parameter count, expert count or active-parameter figure; whether
gemini-3.1-pro-previewis eligible for supervised fine-tuning on Vertex AI; the exact per-query price of Google Search grounding for this specific model; and the ranges or defaults fortopP,topK,presencePenaltyandfrequencyPenalty, which the API reference lists without them. The roster row this page was built from said 2M context — the model card and developer guide both say 1M, so 1M is what is stated above.
Price and capacity verified 2026-09-18 against https://ai.google.dev/pricing. re-read 2026-09-18 (primary source read today: https://ai.google.dev/pricing — $2.00 per million input and $12.00 per million output at the <=200k tier, UNCHANGED. The page still labels it gemini-3.1-pro-preview, a preview channel and not a stable id, which this row's own note has recorded since 2026-08-18. Closes the 2026-09-03 id-drift reading.)
What changedWhat changed here
- Anthropic merges Claude chat and Cowork in one interface
Anthropic merged Claude chat and Cowork into a single interface, with Docs and Slides tools added, rolling out to Pro and Max subscribers first. If you're building on Claude, the surface you integrate against is consolidating into one general agent rather than separate chat and work products — worth re-checking your assumptions about which endpoints and plan tiers expose what.
Three kinds of claim, strongest first. Signal runs every morning.