Home › Frontier Models › GPT-5.6 Terra
Model · Reference

GPT-5.6 Terra

Everyday production scale

In one line

The middle GPT-5.6 tier: byte-identical API to Sol, same 1,050,000-token window and 128,000-token output ceiling, at half Sol's input price and three fifths of its output price.

Why this oneWhat it is actually for

Terra is the tier you run in production once your evals stop distinguishing it from Sol. It takes the same request body, the same context window, the same effort ladder up to max, the same pro mode and the same tool surface — only the model ID and the price change. That makes the decision empirical rather than architectural: send the same traffic to both, compare task success and cost per successful task, and keep Sol only for the request classes where it wins. OpenAI's own migration guidance says to preserve effective reasoning effort first and change the model second. Terra is where most everyday production work lands after that exercise.

What it isArchitecture, lineage, training

GPT-5.6 Terra is the balanced tier of the closed-weight GPT-5.6 family that OpenAI made generally available on 9 July 2026 alongside gpt-5.6-sol and gpt-5.6-luna. OpenAI's model page rates it "Reasoning: Higher, Speed: Fast" — one notch below Sol's "Highest" — and OpenAI's guidance describes it as "strong performance at a lower price".

No architecture is disclosed. There is no parameter count, no dense-versus-mixture-of-experts statement, no expert count, no layer or hidden-size numbers, and no training-corpus description beyond a knowledge cutoff of 16 February 2026. Critically, OpenAI never says how Terra relates to Sol — it does not claim Terra is a distillation, a pruned variant, or a separately trained model. Anyone telling you which it is, is guessing. What the vendor documents is that the three tiers share a documented interface, a 1,050,000-token context window, a 128,000-token output ceiling, and the same reasoning controls.

Behaviourally, Terra is a reasoning model in the same sense Sol is. It emits reasoning tokens billed as output tokens and never returned verbatim, it reasons adaptively across effort levels, and it supports the same GPT-5.6 additions: persisted reasoning across turns via reasoning.context, which the family defaults to all_turns; explicit prompt caching with request-placed breakpoints; max reasoning effort; and reasoning.mode set to pro for more model work per request at standard token rates.

The same real-time cyber and biology misuse classifiers run over Terra's generations as over Sol's. They can block a request or pause a stream for several seconds mid-generation, and OpenAI states they may occasionally intervene on legitimate dual-use security work.

At a glanceSee it

GPT-5.6 Terra diagram

Terra sits mid-ladder: one request body, three price points, identical context and controls.

CapacityContext, output and what fits

FactValueSource
Model IDgpt-5.6-terra; no shorter alias documentedOpenAI model page
Context window1,050,000 tokensOpenAI model page
Max output tokens128,000OpenAI model page
Max input tokens922,000 per Microsoft Foundry's table, within the same window; OpenAI publishes no separate input capMicrosoft Learn, Foundry Models sold by Azure
Input modalitiesText and image; audio and video not supportedOpenAI model page
Output modalitiesText onlyOpenAI model page
Knowledge cutoff16 February 2026. Microsoft Foundry's table says "Training Data up to June 2026" for the same ID — the vendors disagreeOpenAI model page; Microsoft Learn
Vendor classificationReasoning "Higher", Speed "Fast"OpenAI model page
Endpointsv1/responses, v1/chat/completions, v1/batch; every other endpoint listed on the model page is marked unsupportedOpenAI model page
Weights availableNo. Closed API; no licence or download offeredOpenAI model page lists API endpoints only
Fine-tuningNot supportedOpenAI model page
Predicted outputsNot supported. The model page does not mention the feature either way; OpenAI's Predicted Outputs guide limits it to the GPT-4o, GPT-4o-mini, GPT-4.1, GPT-4.1-mini and GPT-4.1-nano seriesOpenAI Predicted Outputs guide
Reasoning tokensSupported, billed as output tokens, not returned verbatimOpenAI model page; reasoning guide
Rate limits at usage tier 515,000 requests per minute, 40,000,000 tokens per minute, 15,000,000,000-token batch queueOpenAI model page

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
modelSelects the tiergpt-5.6-terraThe rest of the body is unchanged from Sol or Luna, which is what makes A/B testing across tiers cheap
input and instructionsTurn content and the system or developer messageString or item array; instructions optionalWith previous_response_id, prior instructions are not carried forward — you replace the system message per turn rather than accumulating it
reasoning.effortThinking budget before answeringnone, low, medium, high, xhigh, max. Default mediumThe dominant cost lever on this tier. OpenAI recommends low for extraction, routing and classification, medium or high for diagnosis and planning
reasoning.modeStandard or pro executionstandard default, proPro on Terra is a real option: it buys more model work at Terra's rates, which can beat standard-mode Sol on cost for some tasks. Measure, do not assume
reasoning.contextWhich prior reasoning items re-enter contextauto, current_turn, all_turns; GPT-5.6 defaults to all_turnsCheck the returned reasoning.context to see the effective mode. Switch to current_turn when a session changes task
reasoning.summaryRequests a summary of hidden reasoningauto, concise, detailedUseful for debugging agent loops; it is a summary, never the raw chain
text.verbosityDefault detail level of the visible answerlow, medium, high. Default mediumDirect control over billed output tokens. For code, OpenAI notes medium and high give longer, more structured output; low keeps it minimal
text.formatText or strict JSON Schema output{"type": "text"} default; json_schema with strict: trueSchema adherence is enforced, but a safety refusal returns a refusal item outside your schema — handle it explicitly
max_output_tokensCeiling covering visible output and reasoningMinimum 16; ceiling 128,000Hitting it returns status: incomplete with reason max_output_tokens, possibly with no visible text at all. OpenAI suggests reserving at least 25,000 tokens
max_tool_callsCaps total built-in tool calls in one responseNumber; applies across all built-in tools, not per toolA blunt but effective circuit breaker for runaway agent loops; further calls are ignored rather than erroring
tool_choiceWhether and which tool is calledauto default, none, required, allowed_tools, or a named functionallowed_tools narrows the callable set without editing the tools array, preserving the cached prefix
parallel_tool_callsPermits multiple tool calls per turnBoolean; parallel allowed by defaultfalse forces exactly zero or one call — the safe setting for non-idempotent side effects
service_tierProcessing laneauto default, default, flex, scale, priorityflex bills at batch rates for slower, occasionally unavailable service; priority costs 2x standard in both context bands. Check the returned service_tier, which can differ from what you asked for
temperature and top_pSampling controlsDocumented on the endpoint as 0 to 2 and 0 to 1Listed in the Responses reference but absent from every GPT-5.6 guide. Whether Terra honours them could not be confirmed from a primary source as of 2026-07-25
prompt_cache_key and prompt_cache_optionsCache routing and breakpoint policyFree-string key; mode is implicit default or explicit; ttl is 30m, the only supported valueTerra's cache write costs $2.50 per million against $0.20 reads, so misplaced breakpoints are the fastest way to erase the tier's savings
stop, seed, frequency_penalty, presence_penalty, logit_bias, nOlder sampling and control parametersNot supported on POST /v1/responsesNo seed means no reproducible sampling here. Constrain output with a strict schema and max_output_tokens instead

On Terra the two knobs that decide whether the tier is worth it are reasoning.effort and the caching pair. Effort is the quality dial you tune against Sol at the same setting; if Terra at high matches Sol at medium, you are still ahead on cost. The caching pair matters more here than on Sol because the absolute savings are smaller, so a badly placed breakpoint eats a larger share of the win. text.verbosity is the third: it trims billed output without touching how hard the model thinks.

SamplingShaping the output distribution

Terra behaves as a reasoning model, and OpenAI's GPT-5.6 documentation says nothing about its sampling distribution. temperature and top_p remain listed body parameters on POST /v1/responses, ranged 0 to 2 and 0 to 1, with the long-standing note to change one but not both — but no GPT-5.6 page mentions either, and whether Terra applies them could not be confirmed from a primary source as of 2026-07-25. The documented levers are different in kind: reasoning.effort changes how much the model explores before committing, and text.verbosity changes how much it writes once it has. OpenAI also notes GPT-5.6 is more concise by default than GPT-5.5, so brevity instructions inherited from an older prompt can now over-compress. A sensible start is effort medium, verbosity medium, no sampling parameters at all — then move effort down one level and see whether your evals notice.

ReasoningThinking, effort and budgets

Terra is a reasoning model with the full GPT-5.6 control set. reasoning.effort accepts none, low, medium, high, xhigh and max, defaulting to medium, and the budget is controlled only through that ladder plus max_output_tokens, which counts thinking and visible output together. Reasoning tokens are billed as output tokens — at Terra's $12 per million rather than Sol's $20, which is where much of the tier's saving comes from on reasoning-heavy work — and are not returned verbatim; you get the count in output_tokens_details.reasoning_tokens and an optional summary via reasoning.summary. reasoning.mode set to pro works on Terra too, at Terra's standard rates. Across turns, GPT-5.6 defaults reasoning.context to all_turns; when you run stateless with store: false, replay the encrypted reasoning items the API returns.

ToolsFunction calling and server tools

Terra's documented tool surface is the same as Sol's: function tools with JSON Schema, custom tools taking freeform text with optional Lark or Regex grammars, and hosted web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use and remote MCP through the Responses API. Parallel tool calls are on by default; parallel_tool_calls: false guarantees zero or one. Strict mode requires additionalProperties: false and all properties in required, with nullable type unions for optional fields. The practical failure modes: on Responses an omitted strict is normalised where possible and quietly falls back, returning strict: false, whereas Chat Completions stays non-strict by default — so the same tool definition behaves differently across the two APIs. Safety refusals arrive as a refusal item that ignores your schema. And reasoning items must be passed back alongside function outputs or the next turn loses the thread.

CostPrice, caching, batching, what drives the bill

List price from OpenAI's pricing page, per million tokens. GPT-5.6 bills in two context bands — OpenAI's model page states that prompts with more than 272,000 input tokens are priced at 2x input and 1.5x output for the full request:

LaneInputCached inputCache writeOutput
Standard, short context$2.00$0.20$2.50$12.00
Standard, long context (over 272K input)$4.00$0.40$5.00$18.00
Batch, short context$1.00$0.10$1.25$6.00
Batch, long context$2.00$0.20$2.50$9.00
Flex, short context$1.00$0.10$1.25$6.00
Flex, long context$2.00$0.20$2.50$9.00
Fast mode, short context$4.00$0.40$5.00$24.00
Fast mode, long context$8.00$0.80$10.00$36.00

Terra is half of Sol on input, cached input and cache write, and three fifths of Sol on output, in both context bands and in every lane, so what the tier saves depends on how output-heavy the workload is. What moves the bill within the tier: reasoning tokens bill as output at $12 per million short-context, so effort is a price setting; crossing 272,000 input tokens reprices the entire request at 2x input and 1.5x output, which on a 1,050,000-token window is easy to trip into unnoticed; cached reads are 90% off but cache writes cost 1.25 times the uncached input rate, a GPT-5.6-specific charge worth tracking through cache_write_tokens; batch is a flat 50% off with a 24-hour window; service_tier: flex bills at batch rates for slower service; Fast mode — renamed from Priority on 2026-07-30 — is 2x standard and is published for both context bands. Pro mode carries no rate premium — the reasoning guide says it bills at the selected model's standard token rates — but it produces more tokens. OpenAI says nothing either way about publishing a tokenizer for this family, so treat that as unknown; what it does document is the input-token endpoint, POST /v1/responses/input_tokens, recommended over a local tiktoken estimate because local tokenizers cannot account for images, files, tools and schemas.

Where it runsSurfaces and availability

SurfaceAvailableNotes
OpenAI APIYesResponses, Chat Completions and Batch endpoints listed on the model page. Responses is required for pro mode, persisted reasoning and hosted tools
Amazon BedrockUnverifiedOpenAI's Bedrock guide covers the GPT-5.6 family and uses openai.gpt-5.6-sol as its example, but does not enumerate Terra. Check the AWS model-support-by-Region page before committing. Where GPT-5.6 is available on Bedrock the context is capped at 272,000 tokens and hosted tools, pro mode and Programmatic Tool Calling are absent
Microsoft Foundry / AzureYesListed as gpt-5.6-terra, snapshot 2026-07-09, 1,050,000 window, 128,000 max output, Standard Global and Data Zone deployments
Google Vertex AINoNot present in Vertex AI Model Garden as of 2026-07-25; only OpenAI's open-weight gpt-oss models appear there
Hugging Face InferenceNoClosed weights, nothing to host
Self-hostingNoWeights are not distributed

StrengthsWhat it is good at

  • Half of Sol's price on input, cached input and cache write, and three fifths of Sol's price on output — with no reduction in context, output ceiling or effort range, per OpenAI's pricing page.
  • Full 1,050,000-token context and 128,000-token output, so long-context work does not force you up to the flagship tier.
  • Supports reasoning.mode: "pro" and max effort, which means you can buy extra model work at Terra's rates instead of moving to Sol.
  • OpenAI rates it "Speed: Fast", a classification Sol's page does not carry — relevant when latency is part of the product.
  • Generally available on Microsoft Foundry under the same model ID, with tier 5 and 6 subscriptions holding quota by default.

LimitsWhere it falls down

  • OpenAI publishes no benchmark separating Terra from Sol or Luna — only the words "Highest", "Higher" and "High". You cannot pick a tier from the documentation; you have to run the eval.
  • Nothing about the architecture is disclosed, including whether Terra is derived from Sol at all.
  • Fine-tuning and predicted outputs are both listed as not supported, so prompt, tools and retrieval are your only adaptation levers.
  • Cache writes cost 1.25 times uncached input. At Terra's smaller absolute savings, a churning prefix erases the tier advantage faster than it would on Sol.
  • Terra's presence on Amazon Bedrock is not confirmed by OpenAI's own Bedrock guide, so a multi-cloud plan built on it needs checking against AWS's Region and model tables first.

Against its neighboursHow it compares

Against gpt-5.6-sol, Terra is the same interface at half the input price and three fifths of the output price, with one vendor word of difference — "Higher" reasoning versus "Highest". The migration guide's procedure is the right one: hold reasoning.effort fixed, change only the model ID, and compare task success and cost per successful task. Terra at high beating Sol at medium is a perfectly ordinary outcome and still saves money. Against gpt-5.6-luna, Terra costs 10x on input and 10x on output, and Luna's advantage is not only price: at usage tier 5 Luna is documented at 30,000 requests and 180,000,000 tokens per minute against Terra's 15,000 and 40,000,000. If your bottleneck is throughput rather than depth, that gap matters more than the rate card.

Getting startedThe smallest call that works

code
curl https://api.openai.com/v1/responses \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-terra",
    "input": "Summarize this incident for the next on-call engineer.",
    "reasoning": { "effort": "medium" },
    "text": { "verbosity": "low" },
    "prompt_cache_key": "oncall-summary-v1"
  }'

First change the model ID to gpt-5.6-sol and back, holding everything else fixed, to measure what the tier is costing you in quality. Then tune reasoning.effort one level in each direction. Add max_output_tokens once you know your reasoning-token profile, and split prompt_cache_key across more values if a single key exceeds roughly 15 requests per minute.

SourcesWhere every claim above came from

Checked

Price and capacity verified 2026-09-12 against https://developers.openai.com/api/docs/pricing. re-read 2026-09-12 (research pass, primary source read today: https://developers.openai.com/api/docs/pricing; https://developers.openai.com/api/docs/models/gpt-5.6-terra — figures confirmed: input_per_m 2.0, output_per_m 12.0, cache_hit_input_per_m 0.2, context 1,050,000 (~1M), max_output 128K. Observed on the page, not added: Standard table verbatim: '| gpt-5.6-terra | $2.00 | $0.20 | $2.50 | $12.00 | $4.00 | $0.40 | $5.00 | $18.00'. Model page: '1,050,000 context window', '128,000 max output tokens'. Cache-write column ($2.50) not carried by the row. gpt-6-astra now above it on the page (see sol row).) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $2 in / $12 out per M, $0.2 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-08-27 (currency proposal src-2026-08-27-00eaf9; read of developers.openai.com/api/docs/pricing, Standard table — $2.00/$12.00 short, $4.00/$18.00 long, $0.20 cached, unchanged at both bands. gpt-5.6-cyber is still on that page ($12.50/$75.00 short context, long context unpriced) and still earns no row here — an operator call, unchanged. Cause of the hash move not determined; no row this file carries moved.)

What changedWhat changed here

RecentAuto-linked from the brief, not a rewrite of this page

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning