OpenAI's flagship from September 2026: a 1,050,000-token reasoning model at $10/$50 per million tokens, sold on computer use, coding and browsing rather than on raw chat quality — and priced two and a half times Sol on both sides of the meter.
Why this oneWhat it is actually for
Reach for gpt-6-astra when the job is operating something rather than answering something: driving a browser, working a codebase over many steps, or running a long agent loop where losing the thread halfway is the failure you are paying to avoid. OpenAI positions it as state of the art on computer use, browsing, software engineering, cybersecurity, science and professional work, and specifically on staying focused, adhering to task boundaries and completing multi-step workflows.
Do not reach for it as a general upgrade to Sol. It costs 2.5x Sol on input and 2.5x on output, for a family of gains concentrated in agentic execution. On a workload that is one prompt in and one answer out, you are paying the agent premium and getting none of it back.
What it isArchitecture, lineage, training
Astra is a closed-weight reasoning model. OpenAI released it to a limited set of organisations on 3 September 2026 and to general availability on 4 September 2026, reaching ChatGPT Plus, Pro, Business and Enterprise as well as the API, Microsoft Azure and AWS Bedrock.
Two things make it different from the GPT-5.6 tiers rather than merely better. The first is that the model is tuned for sustained execution: the claim is not that single answers improve but that a long chain of them stays on task. The second is that it is the first OpenAI model to reach the Critical level of cybersecurity capability under their Preparedness Framework — OpenAI states it can find previously unknown security flaws and develop ways to exploit them across well-protected systems without a person guiding each step. That is a capability statement from the vendor, and it is the reason the release was delayed to add safeguards after a series of unsanctioned agent incidents in July 2026.
At a glanceSee it
How an Astra request is shaped and priced — cache lookup, reasoning effort, then the tool loop that the model is actually sold on.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| Model ID | gpt-6-astra; one snapshot listed, no alias | OpenAI model page |
| Context window | 1,050,000 tokens (~1M) | OpenAI model page |
| Max output | 128,000 tokens | OpenAI model page |
| Knowledge cutoff | 30 April 2026 | OpenAI model page |
| Modalities | Text and image in, text out; vision | OpenAI model page |
| Released | 3 Sep 2026 limited, 4 Sep 2026 general | OpenAI announcement |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
model | Selects the model | gpt-6-astra | No alias is published, so pin the id. |
reasoning.effort | How much hidden work before answering | low, medium, high, xhigh, max | Reasoning tokens bill at the output rate of $50 per million, so effort is the main cost dial on this model. Note there is no none here, unlike Sol. |
tools | Function calling | Supported | The agent loop this model is sold for runs through here. |
response_format | Structured outputs | Supported | Schema-constrained output is available. |
Sampling parameters, parallel tool-call defaults and a service-tier flag are not stated on the model page for Astra. They are left out here rather than carried over from Sol.
SamplingShaping the output distribution
The model page for Astra does not document sampling controls. Rather than assume the GPT-5.6 behaviour still applies, treat this as unverified: if you need deterministic decoding, test it against the API before you design around it. Reasoning models in this family have generally moved their quality dial from sampling to reasoning.effort.
ReasoningThinking, effort and budgets
Yes, and it is the model's main axis. reasoning.effort accepts low, medium, high, xhigh and max. Unlike gpt-5.6-sol, there is no none setting — you cannot turn reasoning off, so the floor price of a call is higher than the input rate alone suggests. Hidden reasoning tokens bill at the output rate.
ToolsFunction calling and server tools
Function calling and structured outputs are both supported. Fine-tuning is not. The endpoints are Responses (v1/responses), Chat Completions (v1/chat/completions) and Batch (v1/batch). Realtime, Live sessions, Assistants, Embeddings, image generation, audio and moderation are all listed as unsupported, so Astra is a text-and-vision reasoning endpoint and nothing else.
CostPrice, caching, batching, what drives the bill
List price from OpenAI's pricing page, per million tokens. Unlike Sol, no promotional sentence accompanies these rates, so treat them as standing prices.
| Mode | Input | Cached input | Output |
|---|---|---|---|
| Standard, short context | $10.00 | $1.00 | $50.00 |
| Standard, long context | $20.00 | $2.00 | $75.00 |
| Batch / Flex | $5.00 | $0.50 | $25.00 |
| Fast mode | $20.00 | $2.00 | $100.00 |
The boards on this site quote the Standard short-context band, which is what every other row quotes. Prompt caching is a 90% discount on input. Against gpt-5.6-sol at $4/$20, Astra is 2.5x on input and 2.5x on output — and because reasoning cannot be switched off, the realistic gap on a short task is wider than the headline.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| OpenAI API | Yes | Responses, Chat Completions and Batch. |
| ChatGPT | Yes | Plus, Pro, Business and Enterprise. |
| Microsoft Azure | Yes | Named in the launch coverage. |
| AWS Bedrock | Yes | Named in the launch coverage. |
| Fine-tuning | No | Not supported. |
| Realtime / audio | No | Not supported. |
Rate limits run five tiers, from 500 RPM and 500K TPM at Tier 1 to 15,000 RPM and 40M TPM at Tier 5.
StrengthsWhat it is good at
- State of the art on computer use, browsing and software engineering, per OpenAI.
- Built for long multi-step runs — staying on task is the headline claim, not answer quality.
- Full 1M-token context with 128K output, matching the GPT-5.6 tiers on capacity.
- 90% prompt-cache discount, which matters on agent loops that resend a large stable prefix.
LimitsWhere it falls down
- 2.5x Sol on both input and output, and reasoning cannot be turned off, so there is no cheap mode.
- No fine-tuning, no realtime, no audio — a narrower surface than the GPT-5.6 family.
- Sampling behaviour is undocumented on the model page; verify before depending on it.
- First OpenAI model rated Critical for cybersecurity capability under their Preparedness Framework. That is a vendor capability statement, and it is the reason the release was delayed for extra safeguards. Read their safety overview before granting it tool access in a sensitive environment.
Against its neighboursHow it compares
Against gpt-5.6-sol: same context and max output, same 90% cache discount, 2.5x the price, and a different job. Sol remains the sensible default for hard single answers and tool-moderate work. Astra earns its premium when the model is driving something over many turns.
Against the reasoning flagships from other labs, the distinguishing claim here is computer use and browsing rather than benchmark reasoning. If your workload never touches a browser or a shell, that premium buys you comparatively little.
Getting startedThe smallest call that works
Pin the id, start at reasoning.effort: low, and only raise it when you can show a failure the extra reasoning fixes — effort is billed at $50 per million. Put the stable part of your prompt first so the 90% cache discount applies to it. If the task is a single question with a single answer, price gpt-5.6-sol alongside it before committing.
SourcesWhere every claim above came from
- OpenAI API pricing —
developers.openai.com/api/docs/pricing(read 17 Sep 2026) - OpenAI model reference —
developers.openai.com/api/docs/modelsand thegpt-6-astramodel page (read 17 Sep 2026) - OpenAI launch and safety overview for GPT-6 Astra (Sep 2026)
Every figure on this page was read from a primary source on 17 September 2026. Anything the sources do not state is marked as not stated rather than inferred.
Price and capacity verified 2026-09-17 against https://developers.openai.com/api/docs/pricing. first entry 2026-09-17 (primary sources read today: https://developers.openai.com/api/docs/pricing — Standard, short-context band: input $10.00, cached input $1.00, output $50.00 per million; long-context band $20.00/$2.00/$75.00, Batch and Flex at half of Standard, Fast mode at double. The short band is quoted here because every other row on this board quotes Standard short. https://developers.openai.com/api/docs/models — context 1.05M (~1M), max output 128K tokens, knowledge cutoff Apr 30 2026, text and image input with text output. Released 3 Sep 2026 (limited preview) and 4 Sep 2026 (general availability). RAISED BY THE CURRENCY WATCH ON 9 SEP 2026 and unactioned for 8 days: the proposal read ‘OpenAI has shipped GPT-6 Astra … which the model boards do not yet list.’ No promotional-pricing sentence appears on the pricing page for this model, so promo is null.)
What changedWhat changed here
Updated this page OpenAI launched GPT-6 Sol and Luna, two tiers in the GPT-6 family with different capability and cost balances.
Add GPT-6 Sol and Luna as two tiers in the GPT-6 family alongside Astra, with their capability and cost positioning.
- Introducing GPT-6 Sol and Luna
OpenAI launched GPT-6 Sol and Luna, two models cut from the same cloth as Astra but with different capability/cost balances. Two tiers means you now have a real choice between a cheap workhorse and a stronger model inside one family — worth testing which of your tasks actually needs Sol.
Three kinds of claim, strongest first. Signal runs every morning.