The cheapest GPT-5.6 tier at $0.20/$1.20 per million, with the family's full 1,050,000-token context and, at usage tier 5, roughly 4.5x Sol's documented token throughput.
Why this oneWhat it is actually for
Luna is for volume: classification, extraction, routing, summarisation, first-pass triage — work where you send a lot of requests and each one is individually cheap to get wrong. It keeps the full 1,050,000-token window and the full effort ladder, so a long document does not force you onto a pricier tier. The throughput headroom is the part people miss: at usage tier 5 OpenAI documents 30,000 requests and 180,000,000 tokens per minute for Luna against 15,000 and 40,000,000 for Sol and Terra. If your constraint is requests per second rather than depth per request, that difference decides the tier before the rate card does.
What it isArchitecture, lineage, training
GPT-5.6 Luna is the efficient tier of the closed-weight GPT-5.6 family, generally available from 9 July 2026 alongside gpt-5.6-sol and gpt-5.6-terra. OpenAI's model page rates it "Reasoning: High, Speed: Fast" and describes it as optimised for cost-sensitive workloads and, in the migration guide, for "efficient, high-volume workloads".
OpenAI discloses no architecture for it. No parameter count, no dense-versus-mixture-of-experts statement, no expert or layer counts, no training-corpus description beyond a knowledge cutoff of 16 February 2026. The vendor also never states how Luna is produced — whether it is distilled from Sol, trained separately, or a smaller model in the same run. Since the weights are not published there is no config.json to check either, so architecture claims about this model from any source are unverifiable.
What is documented is that Luna is a full member of the family rather than a cut-down sibling. It carries the same 1,050,000-token context, the same 128,000-token output ceiling, the same text and image input, the same reasoning-token behaviour, and the same GPT-5.6 features: persisted reasoning via reasoning.context defaulting to all_turns, explicit prompt caching with request-placed breakpoints and a 30-minute minimum TTL, max reasoning effort, and reasoning.mode set to pro. The tool list on its model page is identical to Sol's, down to computer use and remote MCP.
The same real-time cyber and biology misuse classifiers run over Luna's output. They can refuse a request or pause generation mid-stream for several seconds — which matters more on a latency-sensitive tier than it does on a flagship one.
At a glanceSee it
On Luna the bill is set by cache routing, a low effort setting, and whether work can go to batch.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| Model ID | gpt-5.6-luna; no shorter alias documented | OpenAI model page |
| Context window | 1,050,000 tokens | OpenAI model page |
| Max output tokens | 128,000 | OpenAI model page |
| Max input tokens | 922,000 per Microsoft Foundry's table, within the same window; OpenAI states no separate input cap | Microsoft Learn, Foundry Models sold by Azure |
| Input modalities | Text and image; audio and video not supported | OpenAI model page |
| Output modalities | Text only | OpenAI model page |
| Knowledge cutoff | 16 February 2026. Microsoft Foundry lists "Training Data up to June 2026" for the same ID — the two vendor pages disagree | OpenAI model page; Microsoft Learn |
| Vendor classification | Reasoning "High", Speed "Fast" | OpenAI model page |
| Weights available | No. Closed API; no licence, no download, no config.json to inspect | OpenAI model page lists API endpoints only |
| Fine-tuning | Not supported | OpenAI model page |
| Endpoints | v1/responses, v1/chat/completions, v1/batch | OpenAI model page |
| Rate limits at usage tier 5 | 30,000 requests per minute, 180,000,000 tokens per minute, 15,000,000,000-token batch queue | OpenAI model page |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
model | Selects the tier | gpt-5.6-luna | Identical body to Sol and Terra, so promoting a failing request class to a higher tier is a one-field change |
input and instructions | Turn content and the system or developer message | String or item array | Keep instructions stable and first in the prompt — it is the part that caches |
reasoning.effort | How much the model thinks first | none, low, medium, high, xhigh, max. Default medium | The single biggest lever on this tier. OpenAI recommends none as a latency baseline for classification and retrieval, and low when the workflow benefits from tool use. Leaving the default at medium is how high-volume bills quietly triple |
reasoning.mode | Standard or pro execution | standard default, pro | Supported, but rarely the right call here — pro raises latency and token count on the tier chosen for neither |
reasoning.context | Which prior reasoning items re-enter context | auto, current_turn, all_turns; GPT-5.6 defaults to all_turns | On stateless single-shot classification this is moot; on chat sessions all_turns costs context you are paying to carry |
text.verbosity | Default detail level of the answer | low, medium, high. Default medium | low is the obvious default for high-volume extraction. Output is 6x the input rate in the short-context band (4.5x once a request crosses 272K input tokens), so trimming it is worth more per token than trimming the prompt |
text.format | Text or strict JSON Schema output | {"type": "text"} default; json_schema with strict: true | The right way to constrain a cheap model. Note a safety refusal returns a refusal item that does not match the schema, so parsers must branch on it |
max_output_tokens | Ceiling on visible plus reasoning tokens | Minimum 16; ceiling 128,000 | A tight ceiling on a high-effort request can burn reasoning tokens and return status: incomplete with no text — pair a low ceiling with low effort, not high |
top_logprobs | Requests per-token log probabilities | Integer 0 to 20; also add message.output_text.logprobs to include | Both the parameter and the include value are documented body fields of POST /v1/responses, but no OpenAI page states that GPT-5.6 Luna actually returns logprobs — the model page lists no such feature. Confirm the field is populated on a live response before designing confidence-based escalation around it |
tool_choice | Whether and which tool is called | auto default, none, required, allowed_tools, or a named function | none removes a whole class of latency and token overhead when the task never needs a tool |
parallel_tool_calls | Allows multiple calls per turn | Boolean; parallel by default | false gives exactly zero or one call, which keeps per-request cost predictable |
background | Runs the response asynchronously | Boolean | Lets long jobs survive client timeouts; you poll instead of holding a connection |
service_tier | Processing lane | auto default, default, flex, scale, priority | flex bills at batch rates for slower, occasionally unavailable service — a natural fit for Luna-shaped work. priority is 2x standard, and is now published for the long-context band as well as the short |
prompt_cache_key and prompt_cache_options | Cache routing and breakpoint policy | Free-string key; mode implicit default or explicit; ttl 30m, the only value | Caching only starts at 1,024 tokens of prefix. Writes cost $0.25 per million against $0.02 reads, so a churning prefix on high volume is the classic way to overpay |
temperature and top_p | Sampling controls | Documented on the endpoint as 0 to 2 and 0 to 1 | Present in the Responses reference, absent from every GPT-5.6 guide. Whether Luna honours them could not be confirmed from a primary source as of 2026-07-25 |
stop, seed, frequency_penalty, presence_penalty, logit_bias, n | Older control parameters | Not supported on POST /v1/responses | No stop sequences and no seed. stop does exist on v1/chat/completions, which this model also serves, but its reference entry is flagged "Not supported with latest reasoning models o3 and o4-mini" and OpenAI says nothing about GPT-5.6 either way — so treat that route as unconfirmed, not as a guaranteed fallback |
On Luna, cost discipline is the whole game and three settings decide it. reasoning.effort left at the medium default is the most common source of a surprise bill on high-volume traffic — drop to low or none and measure. text.verbosity: "low" attacks the expensive side of the rate card, since output costs 6x input in the short-context band. And prompt_cache_key with a stable prefix turns repeat traffic into $0.02-per-million reads.
SamplingShaping the output distribution
Luna is a reasoning model and OpenAI documents nothing about its output distribution. temperature and top_p are listed as body parameters of POST /v1/responses with ranges 0 to 2 and 0 to 1, and the reference repeats the standing advice to change one but not both — but neither appears anywhere in the GPT-5.6 model pages or guides, and whether Luna applies them could not be confirmed from a primary source as of 2026-07-25. The controls OpenAI actually documents work on a different axis: reasoning.effort governs exploration before the answer, text.verbosity governs length after it, and the model is described as reasoning adaptively, spending fewer tokens on simple inputs. For high-volume work the sensible starting point is effort low, verbosity low, a strict JSON Schema, and no sampling parameters — then raise effort only for the request classes your evals show failing.
ReasoningThinking, effort and budgets
Yes. Luna takes the same reasoning.effort ladder as its tier neighbours — none, low, medium, high, xhigh, max — defaulting to medium, and OpenAI's model page rates its reasoning capability "High". That default is the thing to change first on a cheap tier: reasoning tokens bill as output at $1.20 per million, so a workload left at medium can cost several times the same workload at none. The budget is controllable only through the ladder plus max_output_tokens, which counts thinking and visible text together; overrun returns status: incomplete with reason max_output_tokens, sometimes before any visible output. Thinking is never returned verbatim — you get output_tokens_details.reasoning_tokens and an optional reasoning.summary. reasoning.mode: "pro" is supported but pulls against everything this tier is for.
ToolsFunction calling and server tools
Luna's model page lists the same tool surface as the flagship: function calling, structured outputs, streaming, plus hosted web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use and remote MCP. Parallel calls are on by default and parallel_tool_calls: false pins it to zero or one. Strict schemas need additionalProperties: false, every property in required, and nullable unions for optional fields; on Responses an omitted strict is normalised where possible and falls back silently, reported as strict: false on the returned tool. The failure modes that bite on a high-volume tier specifically: a safety refusal comes back as a refusal item that will not parse against your schema, so a naive pipeline throws; leaving a large tool catalogue in every request inflates the cached prefix and the bill, which is what tool_search with defer_loading: true exists to fix; and hosted tools do not exist on Amazon Bedrock at all.
CostPrice, caching, batching, what drives the bill
List price from OpenAI's pricing page, per million tokens. GPT-5.6 bills in two context bands — OpenAI's model page states that prompts with more than 272,000 input tokens are priced at 2x input and 1.5x output for the full request:
| Lane | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
| Standard, short context | $0.20 | $0.02 | $0.25 | $1.20 |
| Standard, long context (over 272K input) | $0.40 | $0.04 | $0.50 | $1.80 |
| Batch, short context | $0.10 | $0.01 | $0.125 | $0.60 |
| Batch, long context | $0.20 | $0.02 | $0.25 | $0.90 |
| Flex, short context | $0.10 | $0.01 | $0.125 | $0.60 |
| Flex, long context | $0.20 | $0.02 | $0.25 | $0.90 |
| Fast mode, short context | $0.40 | $0.04 | $0.50 | $2.40 |
| Fast mode, long context | $0.80 | $0.08 | $1.00 | $3.60 |
Luna is one tenth of Terra on every line, in both bands and every lane; against Sol there is no single ratio any more — one twentieth on input, cached input and cache write, and 6% on output. Within the tier the bill is driven by four things. Reasoning tokens bill as output at $1.20 per million short-context, so reasoning.effort is a price control — the gap between none and medium can dominate everything else. Context band is a step change: over 272,000 input tokens the whole request reprices at 2x input and 1.5x output, which also flattens the output-to-input ratio from 6x to 4.5x. Caching is 90% off on reads but GPT-5.6 charges 1.25x uncached input for writes, and caching only engages at 1,024 tokens of prefix, so short high-volume prompts may never cache at all; track cached_tokens and cache_write_tokens. Batch is a flat 50% off with a 24-hour completion window, and service_tier: flex gives batch rates synchronously in exchange for slower and occasionally unavailable service. OpenAI states nothing either way about publishing a tokenizer for this family, so treat that as unknown; it does document the input-token endpoint, POST /v1/responses/input_tokens, and recommends it over a local tiktoken estimate because local tokenizers cannot account for images, files, tools and schemas.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| OpenAI API | Yes | Model page lists v1/responses, v1/chat/completions and v1/batch. Chat Completions is the fallback if you need stop sequences |
| Amazon Bedrock | Unverified | OpenAI's Bedrock guide describes GPT-5.6 family support but names only openai.gpt-5.6-sol; Luna is not enumerated. Check the AWS model-support-by-Region page. Where GPT-5.6 runs on Bedrock the context is capped at 272,000 tokens and hosted tools are unavailable |
| Microsoft Foundry / Azure | Yes | Listed as gpt-5.6-luna, snapshot 2026-07-09, 1,050,000 window, 128,000 max output |
| Google Vertex AI | No | Not listed in Vertex AI Model Garden as of 2026-07-25; only OpenAI's open-weight gpt-oss models appear there |
| Hugging Face Inference | No | Weights are not published |
| Self-hosting | No | Closed weights; no config.json or model card exists to inspect |
StrengthsWhat it is good at
- The full 1,050,000-token context and 128,000-token output ceiling at $0.20 in and $1.20 out per million — long-document work does not force an upgrade to a pricier tier.
- Documented tier-5 rate limits of 30,000 requests and 180,000,000 tokens per minute, double the requests and 4.5x the tokens of Sol and Terra.
- Cached input at $0.02 per million, the cheapest read rate in the family, which makes a stable system prompt plus
prompt_cache_keyunusually effective here. - Same tool list as the flagship on its model page, including code interpreter, computer use and remote MCP — the cheap tier is not feature-stripped.
- Combines with batch or
service_tier: flexfor another 50% off, taking effective input to $0.10 per million.
LimitsWhere it falls down
- OpenAI rates it "High" reasoning against Sol's "Highest" but publishes no benchmark, so the actual quality gap on your task is unknown until you measure it.
- The
reasoning.effortdefault ismedium, notlow— a cheap tier with an expensive default, and the most common cause of a surprise bill on high-volume traffic. - Prompt caching only engages at 1,024 tokens or more of prefix, so short high-frequency prompts get no discount at all while cache writes still cost 1.25x input.
- Fine-tuning is not supported, so you cannot buy back quality by specialising the model on your data.
- Real-time safety classifiers can pause generation mid-stream for several seconds, which is a larger proportional latency hit on a tier chosen for speed.
Against its neighboursHow it compares
Against gpt-5.6-terra, Luna costs one tenth as much per token and, at usage tier 5, is documented at 30,000 requests and 180,000,000 tokens per minute versus Terra's 15,000 and 40,000,000 — so the choice is often about throughput ceilings rather than unit price. The vendor gives you one word of quality difference, "High" versus "Higher", and no benchmark, so the tier decision has to come from your own evals. A common shape that works: run everything on Luna at reasoning.effort: "low" with top_logprobs enabled, and escalate only low-confidence or high-stakes requests to Terra or Sol with the identical body. Against gpt-5.6-sol, Luna is one twentieth the input price and 6% of the output price, with the same context and the same tools — the gap you are paying for is judgment on hard problems, not capacity.
Getting startedThe smallest call that works
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-luna",
"input": "Classify this ticket as billing, bug, or feature request.",
"reasoning": { "effort": "low" },
"text": { "verbosity": "low" }
}'Change reasoning.effort first — try none and see whether accuracy moves, since the medium default is the expensive setting on this tier. Then add a strict text.format JSON Schema so downstream parsing is safe, and a stable prompt_cache_key once the prefix exceeds 1,024 tokens. Move offline work to the Batch endpoint for another 50%.
SourcesWhere every claim above came from
- GPT-5.6 Luna Model | OpenAI API
- Models | OpenAI API
- Pricing | OpenAI API
- Create a model response | OpenAI API Reference
- Using GPT-5.6 | OpenAI API
- Reasoning models | OpenAI API
- API deployment checklist | OpenAI API
- Prompt caching | OpenAI API
- Batch API | OpenAI API
- Flex processing | OpenAI API
- Priority processing | OpenAI API
- Function calling | OpenAI API
- Structured model outputs | OpenAI API
- OpenAI models in Amazon Bedrock | OpenAI API
- Foundry Models sold by Azure | Microsoft Learn
- Could not confirm: openai.com's launch post and consumer pricing page returned HTTP 403 and were not read; all prices come from developers.openai.com/api/docs/pricing. No architecture is disclosed and no weights exist to check, so parameter count, dense-versus-MoE and any distillation relationship to Sol are unknown. No vendor benchmark separates Luna from Terra or Sol. Whether
temperatureandtop_ptake effect on this model is not stated. Luna's availability on Amazon Bedrock is not enumerated in OpenAI's Bedrock guide. OpenAI states a 16 February 2026 knowledge cutoff while Microsoft Foundry's table says training data to June 2026. The site row this page was built from listed fine-tuning as available "via API" and the context as "~1M"; the model page says fine-tuning is not supported and the window is 1,050,000, and the vendor page wins.
Price and capacity verified 2026-09-12 against https://developers.openai.com/api/docs/pricing. re-read 2026-09-12 (research pass, primary source read today: https://developers.openai.com/api/docs/pricing; https://developers.openai.com/api/docs/models/gpt-5.6-luna — figures confirmed: input_per_m 0.2, output_per_m 1.2, cache_hit_input_per_m 0.02, context 1,050,000 (~1M), max_output 128K. Observed on the page, not added: Standard table verbatim: '| gpt-5.6-luna | $0.20 | $0.02 | $0.25 | $1.20 | $0.40 | $0.04 | $0.50 | $1.80'. Model page: '1,050,000 context window', '128,000 max output tokens'. Cache-write column ($0.25) not carried by the row.) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $0.2 in / $1.2 out per M, $0.02 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-08-27 (currency proposal src-2026-08-27-00eaf9; read of developers.openai.com/api/docs/pricing, Standard table — $0.20/$1.20 short, $0.40/$1.80 long, $0.02 cached, unchanged at both bands. gpt-5.6-cyber is still on that page ($12.50/$75.00 short context, long context unpriced) and still earns no row here — an operator call, unchanged. Cause of the hash move not determined; no row this file carries moved.)
What changedWhat changed here
Updated this page OpenAI launched GPT-6 Sol and Luna, two tiers in the GPT-6 family with different capability and cost balances.
Add GPT-6 Sol and Luna as two tiers in the GPT-6 family alongside Astra, with their capability and cost positioning.
Updated this page GPT-6 Sol and Luna are reported at half the price of their GPT-5.6 equivalents, with Luna halved again.
Record that GPT-6 Sol and Luna are priced at half their GPT-5.6 equivalents, with Luna halved again, and update the price figures on both pages.
Updated this page GPT-5.6 Luna's price is reported to drop 80% to $0.45 per million tokens, which supersedes the page's current per-million figures.
Change the GPT-5.6 Luna page's $0.20/$1.20 per-million pricing to the reported $0.45 per million, and flag the figure as reported rather than confirmed.
- Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war
Simon Willison's hands-on notes put numbers on the shift: GPT-6 Sol and Luna are half the price of their GPT-5.6 equivalents, with GPT-6 Luna halving the cost of the already-cheap GPT-5.6 Luna. That directly changes the cost math for anything you're running at volume.
- Introducing GPT-6 Sol and Luna
OpenAI launched GPT-6 Sol and Luna, two models cut from the same cloth as Astra but with different capability/cost balances. Two tiers means you now have a real choice between a cheap workhorse and a stronger model inside one family — worth testing which of your tasks actually needs Sol.
- Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index
Qwen 3.8 27B — a 27B-parameter model — scored 52 on the Artificial Analysis Intelligence Index, the same as GPT-5.6 Luna at max and one point behind GLM-5.2 (753B) and DeepSeek V4 Pro (1.7T). A model this small matching much larger ones changes what you should consider for cost- and latency-sensitive deployments.
- Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users
OpenAI shipped an improved GPT-5.6 Sol into ChatGPT and expanded free-tier access to GPT-5.6 Luna with unlimited everyday text chats — the entry point for experimenting with frontier reasoning just got cheaper and the default model better.
- Advancing the price-performance frontier with GPT‑5.6
OpenAI cut GPT-5.6 Terra prices by 20% and GPT-5.6 Luna by 80%, using GPT-5.6 Sol to optimize load balancing and inference kernels. If you pay per token for agent workloads, re-cost your pipelines now — the price-performance frontier just moved.
Three kinds of claim, strongest first. Signal runs every morning.