These are published vendor list prices, each confirmed on its own date, 2026-09-05 to 2026-09-18, and every row carries the date it was last checked and the page it was checked against. Vendors change prices, and this table is a record of what they said on those dates — not a quote and not an offer. Confirm the current price with the vendor before you commit a budget to it. Everything here is provided as-is; see the Terms.
What changedWhat changed here
Updated this page A Claude Opus 5.5 model is now callable, which supersedes the Opus 5 and Opus 4.8 entries as the current Opus tier.
The Claude model pages and the costing board must be updated: a Claude Opus 5.5 now exists, so the Opus 5 and Opus 4.8 entries are no longer the current Opus tier and their pricing and status claims need revising.
Updated this page Anthropic released Claude Opus 5.5 at lower prices with Fable-level performance, superseding Opus 5 as the newest Opus model.
Update the Opus page to record Opus 5.5 as the newest Opus release at lower prices, superseding Opus 5 as the recommended default.
Updated this page GPT-6 Sol and Luna are reported at half the price of their GPT-5.6 equivalents, with Luna halved again.
Record that GPT-6 Sol and Luna are priced at half their GPT-5.6 equivalents, with Luna halved again, and update the price figures on both pages.
Updated this page OpenAI has shipped GPT-6 Astra, a new flagship business model with computer use and stronger reasoning, which the model boards do not yet list.
Add GPT-6 Astra to the model fact file and boards as a new frontier model with computer use, and update the newest-entry as-of date.
Updated this page DeepSeek V4.1 Flash is priced at $0.22 per million tokens, below the V4 Flash price the board currently lists.
Replace the DeepSeek V4 Flash entry with V4.1 Flash at $0.22 per million tokens and note that V4 Pro is retired.
Updated this page DeepSeek has shipped a V4.1-Flash model with a 1M-token context window and multimodal input, superseding the V4 Flash entry on the board.
Update the DeepSeek V4 Flash entry to reflect the V4.1-Flash release with its 1M-token context and multimodal input, or add it as a new model row on the board.
Updated this page DeepSeek V4.1 Flash is reported as a 552B-parameter model with 8B active and a 60% cached-input reduction, tying Opus 5, which changes the cost and capability picture for the DeepSeek Flash entry.
Revise the DeepSeek V4 Flash page to reflect the V4.1 Flash specs — 552B parameters with 8B active and a 60% cached-input cut — and update the cost comparison against Opus 5.
Updated this page GPT-5.6 Luna's price is reported to drop 80% to $0.45 per million tokens, which supersedes the page's current per-million figures.
Change the GPT-5.6 Luna page's $0.20/$1.20 per-million pricing to the reported $0.45 per million, and flag the figure as reported rather than confirmed.
Updated this page DeepSeek now serves all V4-Pro API traffic with V4.1-Flash and bills at Flash rates, so the V4-Pro page's pricing no longer describes what a caller pays.
Update the DeepSeek V4 Pro page to state that its API traffic is now routed to V4.1-Flash at Flash rates, so the listed $1.32/$3.96 peak pricing no longer applies to live calls.
Updated this page Google introduced a new Flash model with an introductory price cut, aimed at coding and agentic workloads.
Add Gemini 3.7 Flash to the frontier and costing model boards with its introductory pricing and coding/agentic positioning.
Showing the 10 most recent references. 3 older were dropped — a reference ages, so this list does not grow forever.
Three kinds of claim, strongest first. Signal runs every morning.
Why it mattersCost is a product constraint, not an afterthought
In AI, cost, latency, and quality form an "iron triangle" — push one and the others move. Treat cost the way you treat a pricing decision: modelled up front, owned, and revisited. The good news is AI cost is calculable — it comes down to tokens and volume.
At a glanceHow a bill is built
Every request: input tokens × input price + output tokens × output price = cost per call; multiply by volume; then cut with levers.
The mathHow AI cost is calculated
Providers bill per million tokens, split into input (your prompt + context) and output (the response). Output is typically 3–6× the input rate. So:
The unit formula
- Cost per call= (input tokens ÷ 1M × input price) + (output tokens ÷ 1M × output price).
- Worked examplea ~40-page doc plus a 600-word answer (~30K input + ~800 output) on Claude Sonnet 5 at $2 / $10 ≈ $0.060 + $0.008 = ~$0.07 per query.
- Scale it× calls per user × users × (1 + retry/agent-loop factor) = monthly spend.
The decisionChoosing a model — the cost matrix
Model choice is the biggest cost lever. Score options against these criteria (a defensible matrix you can show leadership) rather than picking what is trending:
| Criterion | What you ask | Cheap model wins if… | Premium model wins if… |
|---|---|---|---|
| Reasoning | How hard is the task? | Simple extract / classify / format | Multi-step logic, legal/medical, code |
| Context size | How much must it read? | Short prompts, small docs | Whole codebases, long documents |
| Latency | How fast must it answer? | Real-time chat, autocomplete | Batch / async is acceptable |
| Cost / 1K tok | What is the unit price? | High volume, thin margins | Low volume, high value per call |
| Privacy | Where can data go? | Public / non-sensitive data | Regulated data → open-weight / VPC |
The field · July 2026Frontier pricing at a glance
What the models actually cost, per million tokens. (Capabilities and use cases are on the Frontier Models page — this is the money view.)
| Provider | Model | Context | Input /1M | Output /1M | Access |
|---|---|---|---|---|---|
| Anthropic | Claude Fable 5 | 1M | $10 | $50 | Closed API |
| Anthropic | Claude Opus 5 | 1M | $5 | $25 | Closed API |
| Anthropic | Claude Opus 4.8 | 1M | $5 | $25 | Closed API |
| Anthropic | Claude Sonnet 5 | 1M | $2 | $10 | Closed API |
| Anthropic | Claude Haiku 4.5 | 200K | $1 | $5 | Closed API |
| OpenAI | GPT-6 Astra (flagship) | ~1M | $10 | $50 | Closed API |
| OpenAI | GPT-5.6 Sol (flagship) | ~1M | $4 | $20 | Closed API |
| OpenAI | GPT-5.6 Terra | ~1M | $2 | $12 | Closed API |
| OpenAI | GPT-5.6 Luna | ~1M | $0.20 | $1.20 | Closed API |
| Gemini 3.1 Pro | 1M | $2 | $12 | Closed API | |
| Gemini 3 Flash | 1M | $0.50 | $3 | Closed API | |
| DeepSeek | DeepSeek V4 Pro | 1M | $1.32 | $3.96 | API + MIT weights |
| DeepSeek | DeepSeek V4 Flash | 1M | $0.44 | $1.32 | API + MIT weights |
| xAI | Grok 4.5 | 500K | $2 | $6 | Closed API |
| Meta | Muse Spark 1.1 | 1M | $1.25 | $4.25 | Closed API |
| Moonshot | Kimi K3 | 1M | $3 | $15 | API only — weights announced, not published |
| Gemini 3.8 Flash | 1M | $0.75 | $3.75 | Closed API | |
| xAI | Grok 4.6 | 500K | $2 | $6 | Closed API |
| Anthropic | Claude Fable 5.1 | 1M | $10 | $50 | Closed API |
One source: every price here, the worked example above and the model matrix all render from build/facts/models.json, so they cannot disagree. Last confirmed 5 Sep 2026 (22 days ago). The sources are watched every morning (64 cited pricing pages, last read 25 Sep 2026). That check reports when a page changes or dies — it does not re-read the numbers, so a price here is still confirmed by a person, and the date above is when one last did. * GPT-5.6 Sol (flagship)'s rate above ($4 / $20) is promotional — the vendor publishes it only through 21 Nov 2026, 55 days away, and does not say what it becomes afterwards. Treat that date as a reminder to re-read the source, not as a forecast.
Unit economicsFrom cents to ROI
Frame cost against the value it replaces
- Roll cost up: per call → per user/session → per month. That is your unit economic.
- Compare to the human/manual cost it replaces. (Worked example: e-commerce content ran $30–80K per season of photographers, copywriters and editors — AI does the first draft for cents, changing the whole business case.)
- Watch the gross margin: if a feature costs $0.30/interaction and you charge $0.10, volume makes it worse, not better.
The leversCutting cost without gutting quality
Cache repeated context (system prompts, docs) — 75–90% off input on most providers, the single biggest lever. Route easy queries to a cheap tier (Haiku/Flash/mini), hard ones to the frontier. Trim output and cap max tokens — output dominates the bill. Shorten context — retrieve fewer, better chunks instead of stuffing everything.
Right-size the model — a small fine-tuned model can beat a giant on a narrow task at a fraction of the cost. Batch non-urgent work for batch-API discounts. Contain agents — every loop is another paid call, so cap steps.
Watch-outsWhere cost quietly hides
The usual surprises
- Agent loopsone user request can become 10+ model calls.
- Long contextstuffing full documents into every prompt multiplies input cost.
- Retries & guardrailsvalidation failures that re-run the model.
- Re-embeddingchanging your embedding model means re-indexing the whole corpus.
- Reasoning modelshidden "thinking" tokens you still pay for.
The playbookHow to run a cost analysis
1) Estimate tokens per interaction (input + output). 2) Pick a candidate model and its input/output price. 3) Multiply to a cost-per-call. 4) Add retrieval, retries and agent-loop overhead. 5) Project to monthly at expected volume. 6) Compare to the value/ROI and target margin. 7) Apply levers and re-model. 8) Instrument cost per feature and per user in production and review it monthly.