The middle GPT-5.6 tier: byte-identical API to Sol, same 1,050,000-token window and 128,000-token output ceiling, at half Sol's input price and three fifths of its output price.
Why this oneWhat it is actually for
Terra is the tier you run in production once your evals stop distinguishing it from Sol. It takes the same request body, the same context window, the same effort ladder up to max, the same pro mode and the same tool surface — only the model ID and the price change. That makes the decision empirical rather than architectural: send the same traffic to both, compare task success and cost per successful task, and keep Sol only for the request classes where it wins. OpenAI's own migration guidance says to preserve effective reasoning effort first and change the model second. Terra is where most everyday production work lands after that exercise.
What it isArchitecture, lineage, training
GPT-5.6 Terra is the balanced tier of the closed-weight GPT-5.6 family that OpenAI made generally available on 9 July 2026 alongside gpt-5.6-sol and gpt-5.6-luna. OpenAI's model page rates it "Reasoning: Higher, Speed: Fast" — one notch below Sol's "Highest" — and OpenAI's guidance describes it as "strong performance at a lower price".
No architecture is disclosed. There is no parameter count, no dense-versus-mixture-of-experts statement, no expert count, no layer or hidden-size numbers, and no training-corpus description beyond a knowledge cutoff of 16 February 2026. Critically, OpenAI never says how Terra relates to Sol — it does not claim Terra is a distillation, a pruned variant, or a separately trained model. Anyone telling you which it is, is guessing. What the vendor documents is that the three tiers share a documented interface, a 1,050,000-token context window, a 128,000-token output ceiling, and the same reasoning controls.
Behaviourally, Terra is a reasoning model in the same sense Sol is. It emits reasoning tokens billed as output tokens and never returned verbatim, it reasons adaptively across effort levels, and it supports the same GPT-5.6 additions: persisted reasoning across turns via reasoning.context, which the family defaults to all_turns; explicit prompt caching with request-placed breakpoints; max reasoning effort; and reasoning.mode set to pro for more model work per request at standard token rates.
The same real-time cyber and biology misuse classifiers run over Terra's generations as over Sol's. They can block a request or pause a stream for several seconds mid-generation, and OpenAI states they may occasionally intervene on legitimate dual-use security work.
At a glanceSee it
Terra sits mid-ladder: one request body, three price points, identical context and controls.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| Model ID | gpt-5.6-terra; no shorter alias documented | OpenAI model page |
| Context window | 1,050,000 tokens | OpenAI model page |
| Max output tokens | 128,000 | OpenAI model page |
| Max input tokens | 922,000 per Microsoft Foundry's table, within the same window; OpenAI publishes no separate input cap | Microsoft Learn, Foundry Models sold by Azure |
| Input modalities | Text and image; audio and video not supported | OpenAI model page |
| Output modalities | Text only | OpenAI model page |
| Knowledge cutoff | 16 February 2026. Microsoft Foundry's table says "Training Data up to June 2026" for the same ID — the vendors disagree | OpenAI model page; Microsoft Learn |
| Vendor classification | Reasoning "Higher", Speed "Fast" | OpenAI model page |
| Endpoints | v1/responses, v1/chat/completions, v1/batch; every other endpoint listed on the model page is marked unsupported | OpenAI model page |
| Weights available | No. Closed API; no licence or download offered | OpenAI model page lists API endpoints only |
| Fine-tuning | Not supported | OpenAI model page |
| Predicted outputs | Not supported. The model page does not mention the feature either way; OpenAI's Predicted Outputs guide limits it to the GPT-4o, GPT-4o-mini, GPT-4.1, GPT-4.1-mini and GPT-4.1-nano series | OpenAI Predicted Outputs guide |
| Reasoning tokens | Supported, billed as output tokens, not returned verbatim | OpenAI model page; reasoning guide |
| Rate limits at usage tier 5 | 15,000 requests per minute, 40,000,000 tokens per minute, 15,000,000,000-token batch queue | OpenAI model page |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
model | Selects the tier | gpt-5.6-terra | The rest of the body is unchanged from Sol or Luna, which is what makes A/B testing across tiers cheap |
input and instructions | Turn content and the system or developer message | String or item array; instructions optional | With previous_response_id, prior instructions are not carried forward — you replace the system message per turn rather than accumulating it |
reasoning.effort | Thinking budget before answering | none, low, medium, high, xhigh, max. Default medium | The dominant cost lever on this tier. OpenAI recommends low for extraction, routing and classification, medium or high for diagnosis and planning |
reasoning.mode | Standard or pro execution | standard default, pro | Pro on Terra is a real option: it buys more model work at Terra's rates, which can beat standard-mode Sol on cost for some tasks. Measure, do not assume |
reasoning.context | Which prior reasoning items re-enter context | auto, current_turn, all_turns; GPT-5.6 defaults to all_turns | Check the returned reasoning.context to see the effective mode. Switch to current_turn when a session changes task |
reasoning.summary | Requests a summary of hidden reasoning | auto, concise, detailed | Useful for debugging agent loops; it is a summary, never the raw chain |
text.verbosity | Default detail level of the visible answer | low, medium, high. Default medium | Direct control over billed output tokens. For code, OpenAI notes medium and high give longer, more structured output; low keeps it minimal |
text.format | Text or strict JSON Schema output | {"type": "text"} default; json_schema with strict: true | Schema adherence is enforced, but a safety refusal returns a refusal item outside your schema — handle it explicitly |
max_output_tokens | Ceiling covering visible output and reasoning | Minimum 16; ceiling 128,000 | Hitting it returns status: incomplete with reason max_output_tokens, possibly with no visible text at all. OpenAI suggests reserving at least 25,000 tokens |
max_tool_calls | Caps total built-in tool calls in one response | Number; applies across all built-in tools, not per tool | A blunt but effective circuit breaker for runaway agent loops; further calls are ignored rather than erroring |
tool_choice | Whether and which tool is called | auto default, none, required, allowed_tools, or a named function | allowed_tools narrows the callable set without editing the tools array, preserving the cached prefix |
parallel_tool_calls | Permits multiple tool calls per turn | Boolean; parallel allowed by default | false forces exactly zero or one call — the safe setting for non-idempotent side effects |
service_tier | Processing lane | auto default, default, flex, scale, priority | flex bills at batch rates for slower, occasionally unavailable service; priority costs 2x standard in both context bands. Check the returned service_tier, which can differ from what you asked for |
temperature and top_p | Sampling controls | Documented on the endpoint as 0 to 2 and 0 to 1 | Listed in the Responses reference but absent from every GPT-5.6 guide. Whether Terra honours them could not be confirmed from a primary source as of 2026-07-25 |
prompt_cache_key and prompt_cache_options | Cache routing and breakpoint policy | Free-string key; mode is implicit default or explicit; ttl is 30m, the only supported value | Terra's cache write costs $2.50 per million against $0.20 reads, so misplaced breakpoints are the fastest way to erase the tier's savings |
stop, seed, frequency_penalty, presence_penalty, logit_bias, n | Older sampling and control parameters | Not supported on POST /v1/responses | No seed means no reproducible sampling here. Constrain output with a strict schema and max_output_tokens instead |
On Terra the two knobs that decide whether the tier is worth it are reasoning.effort and the caching pair. Effort is the quality dial you tune against Sol at the same setting; if Terra at high matches Sol at medium, you are still ahead on cost. The caching pair matters more here than on Sol because the absolute savings are smaller, so a badly placed breakpoint eats a larger share of the win. text.verbosity is the third: it trims billed output without touching how hard the model thinks.
SamplingShaping the output distribution
Terra behaves as a reasoning model, and OpenAI's GPT-5.6 documentation says nothing about its sampling distribution. temperature and top_p remain listed body parameters on POST /v1/responses, ranged 0 to 2 and 0 to 1, with the long-standing note to change one but not both — but no GPT-5.6 page mentions either, and whether Terra applies them could not be confirmed from a primary source as of 2026-07-25. The documented levers are different in kind: reasoning.effort changes how much the model explores before committing, and text.verbosity changes how much it writes once it has. OpenAI also notes GPT-5.6 is more concise by default than GPT-5.5, so brevity instructions inherited from an older prompt can now over-compress. A sensible start is effort medium, verbosity medium, no sampling parameters at all — then move effort down one level and see whether your evals notice.
ReasoningThinking, effort and budgets
Terra is a reasoning model with the full GPT-5.6 control set. reasoning.effort accepts none, low, medium, high, xhigh and max, defaulting to medium, and the budget is controlled only through that ladder plus max_output_tokens, which counts thinking and visible output together. Reasoning tokens are billed as output tokens — at Terra's $12 per million rather than Sol's $20, which is where much of the tier's saving comes from on reasoning-heavy work — and are not returned verbatim; you get the count in output_tokens_details.reasoning_tokens and an optional summary via reasoning.summary. reasoning.mode set to pro works on Terra too, at Terra's standard rates. Across turns, GPT-5.6 defaults reasoning.context to all_turns; when you run stateless with store: false, replay the encrypted reasoning items the API returns.
ToolsFunction calling and server tools
Terra's documented tool surface is the same as Sol's: function tools with JSON Schema, custom tools taking freeform text with optional Lark or Regex grammars, and hosted web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use and remote MCP through the Responses API. Parallel tool calls are on by default; parallel_tool_calls: false guarantees zero or one. Strict mode requires additionalProperties: false and all properties in required, with nullable type unions for optional fields. The practical failure modes: on Responses an omitted strict is normalised where possible and quietly falls back, returning strict: false, whereas Chat Completions stays non-strict by default — so the same tool definition behaves differently across the two APIs. Safety refusals arrive as a refusal item that ignores your schema. And reasoning items must be passed back alongside function outputs or the next turn loses the thread.
CostPrice, caching, batching, what drives the bill
List price from OpenAI's pricing page, per million tokens. GPT-5.6 bills in two context bands — OpenAI's model page states that prompts with more than 272,000 input tokens are priced at 2x input and 1.5x output for the full request:
| Lane | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
| Standard, short context | $2.00 | $0.20 | $2.50 | $12.00 |
| Standard, long context (over 272K input) | $4.00 | $0.40 | $5.00 | $18.00 |
| Batch, short context | $1.00 | $0.10 | $1.25 | $6.00 |
| Batch, long context | $2.00 | $0.20 | $2.50 | $9.00 |
| Flex, short context | $1.00 | $0.10 | $1.25 | $6.00 |
| Flex, long context | $2.00 | $0.20 | $2.50 | $9.00 |
| Fast mode, short context | $4.00 | $0.40 | $5.00 | $24.00 |
| Fast mode, long context | $8.00 | $0.80 | $10.00 | $36.00 |
Terra is half of Sol on input, cached input and cache write, and three fifths of Sol on output, in both context bands and in every lane, so what the tier saves depends on how output-heavy the workload is. What moves the bill within the tier: reasoning tokens bill as output at $12 per million short-context, so effort is a price setting; crossing 272,000 input tokens reprices the entire request at 2x input and 1.5x output, which on a 1,050,000-token window is easy to trip into unnoticed; cached reads are 90% off but cache writes cost 1.25 times the uncached input rate, a GPT-5.6-specific charge worth tracking through cache_write_tokens; batch is a flat 50% off with a 24-hour window; service_tier: flex bills at batch rates for slower service; Fast mode — renamed from Priority on 2026-07-30 — is 2x standard and is published for both context bands. Pro mode carries no rate premium — the reasoning guide says it bills at the selected model's standard token rates — but it produces more tokens. OpenAI says nothing either way about publishing a tokenizer for this family, so treat that as unknown; what it does document is the input-token endpoint, POST /v1/responses/input_tokens, recommended over a local tiktoken estimate because local tokenizers cannot account for images, files, tools and schemas.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| OpenAI API | Yes | Responses, Chat Completions and Batch endpoints listed on the model page. Responses is required for pro mode, persisted reasoning and hosted tools |
| Amazon Bedrock | Unverified | OpenAI's Bedrock guide covers the GPT-5.6 family and uses openai.gpt-5.6-sol as its example, but does not enumerate Terra. Check the AWS model-support-by-Region page before committing. Where GPT-5.6 is available on Bedrock the context is capped at 272,000 tokens and hosted tools, pro mode and Programmatic Tool Calling are absent |
| Microsoft Foundry / Azure | Yes | Listed as gpt-5.6-terra, snapshot 2026-07-09, 1,050,000 window, 128,000 max output, Standard Global and Data Zone deployments |
| Google Vertex AI | No | Not present in Vertex AI Model Garden as of 2026-07-25; only OpenAI's open-weight gpt-oss models appear there |
| Hugging Face Inference | No | Closed weights, nothing to host |
| Self-hosting | No | Weights are not distributed |
StrengthsWhat it is good at
- Half of Sol's price on input, cached input and cache write, and three fifths of Sol's price on output — with no reduction in context, output ceiling or effort range, per OpenAI's pricing page.
- Full 1,050,000-token context and 128,000-token output, so long-context work does not force you up to the flagship tier.
- Supports
reasoning.mode: "pro"andmaxeffort, which means you can buy extra model work at Terra's rates instead of moving to Sol. - OpenAI rates it "Speed: Fast", a classification Sol's page does not carry — relevant when latency is part of the product.
- Generally available on Microsoft Foundry under the same model ID, with tier 5 and 6 subscriptions holding quota by default.
LimitsWhere it falls down
- OpenAI publishes no benchmark separating Terra from Sol or Luna — only the words "Highest", "Higher" and "High". You cannot pick a tier from the documentation; you have to run the eval.
- Nothing about the architecture is disclosed, including whether Terra is derived from Sol at all.
- Fine-tuning and predicted outputs are both listed as not supported, so prompt, tools and retrieval are your only adaptation levers.
- Cache writes cost 1.25 times uncached input. At Terra's smaller absolute savings, a churning prefix erases the tier advantage faster than it would on Sol.
- Terra's presence on Amazon Bedrock is not confirmed by OpenAI's own Bedrock guide, so a multi-cloud plan built on it needs checking against AWS's Region and model tables first.
Against its neighboursHow it compares
Against gpt-5.6-sol, Terra is the same interface at half the input price and three fifths of the output price, with one vendor word of difference — "Higher" reasoning versus "Highest". The migration guide's procedure is the right one: hold reasoning.effort fixed, change only the model ID, and compare task success and cost per successful task. Terra at high beating Sol at medium is a perfectly ordinary outcome and still saves money. Against gpt-5.6-luna, Terra costs 10x on input and 10x on output, and Luna's advantage is not only price: at usage tier 5 Luna is documented at 30,000 requests and 180,000,000 tokens per minute against Terra's 15,000 and 40,000,000. If your bottleneck is throughput rather than depth, that gap matters more than the rate card.
Getting startedThe smallest call that works
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-terra",
"input": "Summarize this incident for the next on-call engineer.",
"reasoning": { "effort": "medium" },
"text": { "verbosity": "low" },
"prompt_cache_key": "oncall-summary-v1"
}'First change the model ID to gpt-5.6-sol and back, holding everything else fixed, to measure what the tier is costing you in quality. Then tune reasoning.effort one level in each direction. Add max_output_tokens once you know your reasoning-token profile, and split prompt_cache_key across more values if a single key exceeds roughly 15 requests per minute.
SourcesWhere every claim above came from
- GPT-5.6 Terra Model | OpenAI API
- Models | OpenAI API
- Pricing | OpenAI API
- Create a model response | OpenAI API Reference
- Using GPT-5.6 | OpenAI API
- Reasoning models | OpenAI API
- API deployment checklist | OpenAI API
- Prompt caching | OpenAI API
- Function calling | OpenAI API
- Structured model outputs | OpenAI API
- Batch API | OpenAI API
- Flex processing | OpenAI API
- OpenAI models in Amazon Bedrock | OpenAI API
- Upgrading to GPT-5.6 Sol | OpenAI API
- Foundry Models sold by Azure | Microsoft Learn
- Could not confirm: openai.com's own launch post and pricing page returned HTTP 403 and were not read, so all prices here come from developers.openai.com/api/docs/pricing. No architecture, parameter count or relationship to Sol is published. No vendor benchmark separates Terra from its tier neighbours. Whether
temperatureandtop_ptake effect is not stated anywhere in the GPT-5.6 documentation. Terra's availability on Amazon Bedrock is not enumerated by OpenAI's Bedrock guide. OpenAI states a 16 February 2026 knowledge cutoff while Microsoft Foundry's table says training data to June 2026. The site row this page was built from said fine-tuning was available "via API"; the model page says it is not supported, and the vendor page wins.
Price and capacity verified 2026-09-12 against https://developers.openai.com/api/docs/pricing. re-read 2026-09-12 (research pass, primary source read today: https://developers.openai.com/api/docs/pricing; https://developers.openai.com/api/docs/models/gpt-5.6-terra — figures confirmed: input_per_m 2.0, output_per_m 12.0, cache_hit_input_per_m 0.2, context 1,050,000 (~1M), max_output 128K. Observed on the page, not added: Standard table verbatim: '| gpt-5.6-terra | $2.00 | $0.20 | $2.50 | $12.00 | $4.00 | $0.40 | $5.00 | $18.00'. Model page: '1,050,000 context window', '128,000 max output tokens'. Cache-write column ($2.50) not carried by the row. gpt-6-astra now above it on the page (see sol row).) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $2 in / $12 out per M, $0.2 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-08-27 (currency proposal src-2026-08-27-00eaf9; read of developers.openai.com/api/docs/pricing, Standard table — $2.00/$12.00 short, $4.00/$18.00 long, $0.20 cached, unchanged at both bands. gpt-5.6-cyber is still on that page ($12.50/$75.00 short context, long context unpriced) and still earns no row here — an operator call, unchanged. Cause of the hash move not determined; no row this file carries moved.)
What changedWhat changed here
- Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users
OpenAI shipped an improved GPT-5.6 Sol into ChatGPT and expanded free-tier access to GPT-5.6 Luna with unlimited everyday text chats — the entry point for experimenting with frontier reasoning just got cheaper and the default model better.
- Advancing the price-performance frontier with GPT‑5.6
OpenAI cut GPT-5.6 Terra prices by 20% and GPT-5.6 Luna by 80%, using GPT-5.6 Sol to optimize load balancing and inference kernels. If you pay per token for agent workloads, re-cost your pipelines now — the price-performance frontier just moved.
Three kinds of claim, strongest first. Signal runs every morning.