Home › Frontier Models › GPT-5.6 Luna
Model · Reference

GPT-5.6 Luna

Cheap high volume

In one line

The cheapest GPT-5.6 tier at $0.20/$1.20 per million, with the family's full 1,050,000-token context and, at usage tier 5, roughly 4.5x Sol's documented token throughput.

Why this oneWhat it is actually for

Luna is for volume: classification, extraction, routing, summarisation, first-pass triage — work where you send a lot of requests and each one is individually cheap to get wrong. It keeps the full 1,050,000-token window and the full effort ladder, so a long document does not force you onto a pricier tier. The throughput headroom is the part people miss: at usage tier 5 OpenAI documents 30,000 requests and 180,000,000 tokens per minute for Luna against 15,000 and 40,000,000 for Sol and Terra. If your constraint is requests per second rather than depth per request, that difference decides the tier before the rate card does.

What it isArchitecture, lineage, training

GPT-5.6 Luna is the efficient tier of the closed-weight GPT-5.6 family, generally available from 9 July 2026 alongside gpt-5.6-sol and gpt-5.6-terra. OpenAI's model page rates it "Reasoning: High, Speed: Fast" and describes it as optimised for cost-sensitive workloads and, in the migration guide, for "efficient, high-volume workloads".

OpenAI discloses no architecture for it. No parameter count, no dense-versus-mixture-of-experts statement, no expert or layer counts, no training-corpus description beyond a knowledge cutoff of 16 February 2026. The vendor also never states how Luna is produced — whether it is distilled from Sol, trained separately, or a smaller model in the same run. Since the weights are not published there is no config.json to check either, so architecture claims about this model from any source are unverifiable.

What is documented is that Luna is a full member of the family rather than a cut-down sibling. It carries the same 1,050,000-token context, the same 128,000-token output ceiling, the same text and image input, the same reasoning-token behaviour, and the same GPT-5.6 features: persisted reasoning via reasoning.context defaulting to all_turns, explicit prompt caching with request-placed breakpoints and a 30-minute minimum TTL, max reasoning effort, and reasoning.mode set to pro. The tool list on its model page is identical to Sol's, down to computer use and remote MCP.

The same real-time cyber and biology misuse classifiers run over Luna's output. They can refuse a request or pause generation mid-stream for several seconds — which matters more on a latency-sensitive tier than it does on a flagship one.

At a glanceSee it

GPT-5.6 Luna diagram

On Luna the bill is set by cache routing, a low effort setting, and whether work can go to batch.

CapacityContext, output and what fits

FactValueSource
Model IDgpt-5.6-luna; no shorter alias documentedOpenAI model page
Context window1,050,000 tokensOpenAI model page
Max output tokens128,000OpenAI model page
Max input tokens922,000 per Microsoft Foundry's table, within the same window; OpenAI states no separate input capMicrosoft Learn, Foundry Models sold by Azure
Input modalitiesText and image; audio and video not supportedOpenAI model page
Output modalitiesText onlyOpenAI model page
Knowledge cutoff16 February 2026. Microsoft Foundry lists "Training Data up to June 2026" for the same ID — the two vendor pages disagreeOpenAI model page; Microsoft Learn
Vendor classificationReasoning "High", Speed "Fast"OpenAI model page
Weights availableNo. Closed API; no licence, no download, no config.json to inspectOpenAI model page lists API endpoints only
Fine-tuningNot supportedOpenAI model page
Endpointsv1/responses, v1/chat/completions, v1/batchOpenAI model page
Rate limits at usage tier 530,000 requests per minute, 180,000,000 tokens per minute, 15,000,000,000-token batch queueOpenAI model page

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
modelSelects the tiergpt-5.6-lunaIdentical body to Sol and Terra, so promoting a failing request class to a higher tier is a one-field change
input and instructionsTurn content and the system or developer messageString or item arrayKeep instructions stable and first in the prompt — it is the part that caches
reasoning.effortHow much the model thinks firstnone, low, medium, high, xhigh, max. Default mediumThe single biggest lever on this tier. OpenAI recommends none as a latency baseline for classification and retrieval, and low when the workflow benefits from tool use. Leaving the default at medium is how high-volume bills quietly triple
reasoning.modeStandard or pro executionstandard default, proSupported, but rarely the right call here — pro raises latency and token count on the tier chosen for neither
reasoning.contextWhich prior reasoning items re-enter contextauto, current_turn, all_turns; GPT-5.6 defaults to all_turnsOn stateless single-shot classification this is moot; on chat sessions all_turns costs context you are paying to carry
text.verbosityDefault detail level of the answerlow, medium, high. Default mediumlow is the obvious default for high-volume extraction. Output is 6x the input rate in the short-context band (4.5x once a request crosses 272K input tokens), so trimming it is worth more per token than trimming the prompt
text.formatText or strict JSON Schema output{"type": "text"} default; json_schema with strict: trueThe right way to constrain a cheap model. Note a safety refusal returns a refusal item that does not match the schema, so parsers must branch on it
max_output_tokensCeiling on visible plus reasoning tokensMinimum 16; ceiling 128,000A tight ceiling on a high-effort request can burn reasoning tokens and return status: incomplete with no text — pair a low ceiling with low effort, not high
top_logprobsRequests per-token log probabilitiesInteger 0 to 20; also add message.output_text.logprobs to includeBoth the parameter and the include value are documented body fields of POST /v1/responses, but no OpenAI page states that GPT-5.6 Luna actually returns logprobs — the model page lists no such feature. Confirm the field is populated on a live response before designing confidence-based escalation around it
tool_choiceWhether and which tool is calledauto default, none, required, allowed_tools, or a named functionnone removes a whole class of latency and token overhead when the task never needs a tool
parallel_tool_callsAllows multiple calls per turnBoolean; parallel by defaultfalse gives exactly zero or one call, which keeps per-request cost predictable
backgroundRuns the response asynchronouslyBooleanLets long jobs survive client timeouts; you poll instead of holding a connection
service_tierProcessing laneauto default, default, flex, scale, priorityflex bills at batch rates for slower, occasionally unavailable service — a natural fit for Luna-shaped work. priority is 2x standard, and is now published for the long-context band as well as the short
prompt_cache_key and prompt_cache_optionsCache routing and breakpoint policyFree-string key; mode implicit default or explicit; ttl 30m, the only valueCaching only starts at 1,024 tokens of prefix. Writes cost $0.25 per million against $0.02 reads, so a churning prefix on high volume is the classic way to overpay
temperature and top_pSampling controlsDocumented on the endpoint as 0 to 2 and 0 to 1Present in the Responses reference, absent from every GPT-5.6 guide. Whether Luna honours them could not be confirmed from a primary source as of 2026-07-25
stop, seed, frequency_penalty, presence_penalty, logit_bias, nOlder control parametersNot supported on POST /v1/responsesNo stop sequences and no seed. stop does exist on v1/chat/completions, which this model also serves, but its reference entry is flagged "Not supported with latest reasoning models o3 and o4-mini" and OpenAI says nothing about GPT-5.6 either way — so treat that route as unconfirmed, not as a guaranteed fallback

On Luna, cost discipline is the whole game and three settings decide it. reasoning.effort left at the medium default is the most common source of a surprise bill on high-volume traffic — drop to low or none and measure. text.verbosity: "low" attacks the expensive side of the rate card, since output costs 6x input in the short-context band. And prompt_cache_key with a stable prefix turns repeat traffic into $0.02-per-million reads.

SamplingShaping the output distribution

Luna is a reasoning model and OpenAI documents nothing about its output distribution. temperature and top_p are listed as body parameters of POST /v1/responses with ranges 0 to 2 and 0 to 1, and the reference repeats the standing advice to change one but not both — but neither appears anywhere in the GPT-5.6 model pages or guides, and whether Luna applies them could not be confirmed from a primary source as of 2026-07-25. The controls OpenAI actually documents work on a different axis: reasoning.effort governs exploration before the answer, text.verbosity governs length after it, and the model is described as reasoning adaptively, spending fewer tokens on simple inputs. For high-volume work the sensible starting point is effort low, verbosity low, a strict JSON Schema, and no sampling parameters — then raise effort only for the request classes your evals show failing.

ReasoningThinking, effort and budgets

Yes. Luna takes the same reasoning.effort ladder as its tier neighbours — none, low, medium, high, xhigh, max — defaulting to medium, and OpenAI's model page rates its reasoning capability "High". That default is the thing to change first on a cheap tier: reasoning tokens bill as output at $1.20 per million, so a workload left at medium can cost several times the same workload at none. The budget is controllable only through the ladder plus max_output_tokens, which counts thinking and visible text together; overrun returns status: incomplete with reason max_output_tokens, sometimes before any visible output. Thinking is never returned verbatim — you get output_tokens_details.reasoning_tokens and an optional reasoning.summary. reasoning.mode: "pro" is supported but pulls against everything this tier is for.

ToolsFunction calling and server tools

Luna's model page lists the same tool surface as the flagship: function calling, structured outputs, streaming, plus hosted web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use and remote MCP. Parallel calls are on by default and parallel_tool_calls: false pins it to zero or one. Strict schemas need additionalProperties: false, every property in required, and nullable unions for optional fields; on Responses an omitted strict is normalised where possible and falls back silently, reported as strict: false on the returned tool. The failure modes that bite on a high-volume tier specifically: a safety refusal comes back as a refusal item that will not parse against your schema, so a naive pipeline throws; leaving a large tool catalogue in every request inflates the cached prefix and the bill, which is what tool_search with defer_loading: true exists to fix; and hosted tools do not exist on Amazon Bedrock at all.

CostPrice, caching, batching, what drives the bill

List price from OpenAI's pricing page, per million tokens. GPT-5.6 bills in two context bands — OpenAI's model page states that prompts with more than 272,000 input tokens are priced at 2x input and 1.5x output for the full request:

LaneInputCached inputCache writeOutput
Standard, short context$0.20$0.02$0.25$1.20
Standard, long context (over 272K input)$0.40$0.04$0.50$1.80
Batch, short context$0.10$0.01$0.125$0.60
Batch, long context$0.20$0.02$0.25$0.90
Flex, short context$0.10$0.01$0.125$0.60
Flex, long context$0.20$0.02$0.25$0.90
Fast mode, short context$0.40$0.04$0.50$2.40
Fast mode, long context$0.80$0.08$1.00$3.60

Luna is one tenth of Terra on every line, in both bands and every lane; against Sol there is no single ratio any more — one twentieth on input, cached input and cache write, and 6% on output. Within the tier the bill is driven by four things. Reasoning tokens bill as output at $1.20 per million short-context, so reasoning.effort is a price control — the gap between none and medium can dominate everything else. Context band is a step change: over 272,000 input tokens the whole request reprices at 2x input and 1.5x output, which also flattens the output-to-input ratio from 6x to 4.5x. Caching is 90% off on reads but GPT-5.6 charges 1.25x uncached input for writes, and caching only engages at 1,024 tokens of prefix, so short high-volume prompts may never cache at all; track cached_tokens and cache_write_tokens. Batch is a flat 50% off with a 24-hour completion window, and service_tier: flex gives batch rates synchronously in exchange for slower and occasionally unavailable service. OpenAI states nothing either way about publishing a tokenizer for this family, so treat that as unknown; it does document the input-token endpoint, POST /v1/responses/input_tokens, and recommends it over a local tiktoken estimate because local tokenizers cannot account for images, files, tools and schemas.

Where it runsSurfaces and availability

SurfaceAvailableNotes
OpenAI APIYesModel page lists v1/responses, v1/chat/completions and v1/batch. Chat Completions is the fallback if you need stop sequences
Amazon BedrockUnverifiedOpenAI's Bedrock guide describes GPT-5.6 family support but names only openai.gpt-5.6-sol; Luna is not enumerated. Check the AWS model-support-by-Region page. Where GPT-5.6 runs on Bedrock the context is capped at 272,000 tokens and hosted tools are unavailable
Microsoft Foundry / AzureYesListed as gpt-5.6-luna, snapshot 2026-07-09, 1,050,000 window, 128,000 max output
Google Vertex AINoNot listed in Vertex AI Model Garden as of 2026-07-25; only OpenAI's open-weight gpt-oss models appear there
Hugging Face InferenceNoWeights are not published
Self-hostingNoClosed weights; no config.json or model card exists to inspect

StrengthsWhat it is good at

  • The full 1,050,000-token context and 128,000-token output ceiling at $0.20 in and $1.20 out per million — long-document work does not force an upgrade to a pricier tier.
  • Documented tier-5 rate limits of 30,000 requests and 180,000,000 tokens per minute, double the requests and 4.5x the tokens of Sol and Terra.
  • Cached input at $0.02 per million, the cheapest read rate in the family, which makes a stable system prompt plus prompt_cache_key unusually effective here.
  • Same tool list as the flagship on its model page, including code interpreter, computer use and remote MCP — the cheap tier is not feature-stripped.
  • Combines with batch or service_tier: flex for another 50% off, taking effective input to $0.10 per million.

LimitsWhere it falls down

  • OpenAI rates it "High" reasoning against Sol's "Highest" but publishes no benchmark, so the actual quality gap on your task is unknown until you measure it.
  • The reasoning.effort default is medium, not low — a cheap tier with an expensive default, and the most common cause of a surprise bill on high-volume traffic.
  • Prompt caching only engages at 1,024 tokens or more of prefix, so short high-frequency prompts get no discount at all while cache writes still cost 1.25x input.
  • Fine-tuning is not supported, so you cannot buy back quality by specialising the model on your data.
  • Real-time safety classifiers can pause generation mid-stream for several seconds, which is a larger proportional latency hit on a tier chosen for speed.

Against its neighboursHow it compares

Against gpt-5.6-terra, Luna costs one tenth as much per token and, at usage tier 5, is documented at 30,000 requests and 180,000,000 tokens per minute versus Terra's 15,000 and 40,000,000 — so the choice is often about throughput ceilings rather than unit price. The vendor gives you one word of quality difference, "High" versus "Higher", and no benchmark, so the tier decision has to come from your own evals. A common shape that works: run everything on Luna at reasoning.effort: "low" with top_logprobs enabled, and escalate only low-confidence or high-stakes requests to Terra or Sol with the identical body. Against gpt-5.6-sol, Luna is one twentieth the input price and 6% of the output price, with the same context and the same tools — the gap you are paying for is judgment on hard problems, not capacity.

Getting startedThe smallest call that works

code
curl https://api.openai.com/v1/responses \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-luna",
    "input": "Classify this ticket as billing, bug, or feature request.",
    "reasoning": { "effort": "low" },
    "text": { "verbosity": "low" }
  }'

Change reasoning.effort first — try none and see whether accuracy moves, since the medium default is the expensive setting on this tier. Then add a strict text.format JSON Schema so downstream parsing is safe, and a stable prompt_cache_key once the prefix exceeds 1,024 tokens. Move offline work to the Batch endpoint for another 50%.

SourcesWhere every claim above came from

Checked

Price and capacity verified 2026-09-12 against https://developers.openai.com/api/docs/pricing. re-read 2026-09-12 (research pass, primary source read today: https://developers.openai.com/api/docs/pricing; https://developers.openai.com/api/docs/models/gpt-5.6-luna — figures confirmed: input_per_m 0.2, output_per_m 1.2, cache_hit_input_per_m 0.02, context 1,050,000 (~1M), max_output 128K. Observed on the page, not added: Standard table verbatim: '| gpt-5.6-luna | $0.20 | $0.02 | $0.25 | $1.20 | $0.40 | $0.04 | $0.50 | $1.80'. Model page: '1,050,000 context window', '128,000 max output tokens'. Cache-write column ($0.25) not carried by the row.) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $0.2 in / $1.2 out per M, $0.02 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-08-27 (currency proposal src-2026-08-27-00eaf9; read of developers.openai.com/api/docs/pricing, Standard table — $0.20/$1.20 short, $0.40/$1.80 long, $0.02 cached, unchanged at both bands. gpt-5.6-cyber is still on that page ($12.50/$75.00 short context, long context unpriced) and still earns no row here — an operator call, unchanged. Cause of the hash move not determined; no row this file carries moved.)

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page OpenAI launched GPT-6 Sol and Luna, two tiers in the GPT-6 family with different capability and cost balances.

    Add GPT-6 Sol and Luna as two tiers in the GPT-6 family alongside Astra, with their capability and cost positioning.

    OpenAI · 22 Sep 2026 · source

  • Updated this page GPT-6 Sol and Luna are reported at half the price of their GPT-5.6 equivalents, with Luna halved again.

    Record that GPT-6 Sol and Luna are priced at half their GPT-5.6 equivalents, with Luna halved again, and update the price figures on both pages.

    Simon Willison · 22 Sep 2026 · source

  • Updated this page GPT-5.6 Luna's price is reported to drop 80% to $0.45 per million tokens, which supersedes the page's current per-million figures.

    Change the GPT-5.6 Luna page's $0.20/$1.20 per-million pricing to the reported $0.45 per million, and flag the figure as reported rather than confirmed.

    China frontier labs · 12 Sep 2026 · source

RecentAuto-linked from the brief, not a rewrite of this page
  • Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war 22 Sep · Simon Willison

    Simon Willison's hands-on notes put numbers on the shift: GPT-6 Sol and Luna are half the price of their GPT-5.6 equivalents, with GPT-6 Luna halving the cost of the already-cheap GPT-5.6 Luna. That directly changes the cost math for anything you're running at volume.

  • Introducing GPT-6 Sol and Luna 22 Sep · OpenAI

    OpenAI launched GPT-6 Sol and Luna, two models cut from the same cloth as Astra but with different capability/cost balances. Two tiers means you now have a real choice between a cheap workhorse and a stronger model inside one family — worth testing which of your tasks actually needs Sol.

  • Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index 17 Aug · Simon Willison

    Qwen 3.8 27B — a 27B-parameter model — scored 52 on the Artificial Analysis Intelligence Index, the same as GPT-5.6 Luna at max and one point behind GLM-5.2 (753B) and DeepSeek V4 Pro (1.7T). A model this small matching much larger ones changes what you should consider for cost- and latency-sensitive deployments.

  • Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users 6 Aug · OpenAI

    OpenAI shipped an improved GPT-5.6 Sol into ChatGPT and expanded free-tier access to GPT-5.6 Luna with unlimited everyday text chats — the entry point for experimenting with frontier reasoning just got cheaper and the default model better.

  • Advancing the price-performance frontier with GPT‑5.6 30 Jul · Simon Willison

    OpenAI cut GPT-5.6 Terra prices by 20% and GPT-5.6 Luna by 80%, using GPT-5.6 Sol to optimize load balancing and inference kernels. If you pay per token for agent workloads, re-cost your pipelines now — the price-performance frontier just moved.

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning