Every line here is billed on every request and read on every turn — that combined cost is the problem skills were invented to solve.
Why you'd careThe problem it solves
Someone reports that the assistant used the wrong date format, so a line goes into the system prompt. Someone else notices it skips the disclaimer on financial questions, so another line goes in. A year later the prompt is nine hundred lines, nobody will delete anything because each line was a fix for a real complaint, and new instructions have started contradicting old ones in ways that only show up in production. Two things are true at once and both get missed. You are paying for all of it on every call, including the calls where none of it is relevant. And the model is reading all of it on every turn, so the thirty lines that matter for this question are competing with eight hundred and seventy that do not. Length is not just a cost problem; it is an attention problem.
ConceptWhat it is
The system prompt is the always-on instruction layer. On Anthropic's Messages API it is a top-level system parameter rather than an entry in the message list, and it accepts either a plain string or a list of text blocks. The list form is what lets you place a cache breakpoint on the last stable block. Render order is fixed: tools first, then system, then messages — which is why a breakpoint at the end of system caches the tool definitions with it.
Its defining property is unconditionality. There is no trigger, no matching step, no way for it to be absent on a turn where it does not apply. Everything else in this family exists because that property is sometimes wrong: a skill is the same kind of text with a condition attached, a tool description is the same kind of text scoped to one capability.
One thing complicates the tidy “it is a parameter, not a message” framing, and it is worth knowing rather than glossing. Recent Anthropic models also accept a system role message inside the messages array — a separate channel for operator instructions that arrive mid-conversation. It carries the same authority as the top-level parameter, but it sits after the conversation history rather than in front of it. Two channels, one privilege level, very different caching behaviour. Other vendors differ again: several treat the system prompt as a message role from the start, which is why request bodies do not port cleanly.
How it worksThe mechanics
The list-of-blocks form with a breakpoint on the last stable block:
"system": [
{"type": "text",
"text": "<stable core instructions>",
"cache_control": {"type": "ephemeral"}}
]Cache reads cost roughly a tenth of base input, so a large frozen system prompt is cheap to keep — provided it is genuinely frozen. Any byte that changes invalidates everything rendered after it. Interpolating today's date, a session ID, a feature flag or a user's name into the system prompt destroys caching for the entire conversation, and the symptom is not an error: it is cache_read_input_tokens quietly sitting at zero.
For values that genuinely arrive mid-conversation, the cache-preserving move is a system-role message appended to messages instead of an edit to the top-level parameter:
"messages": [
...history...,
{"role": "user", "content": "..."},
{"role": "system",
"content": "Terse mode enabled. Keep responses under 40 words."}
]The rules are specific. It must follow a user message, or an assistant message that ended in server-tool use. It must be either the last entry or be followed by an assistant turn. It cannot be messages[0] — use the top-level parameter for the opening instruction. As of 2026-09-12 it is available on Claude Fable 5.1, Mythos 5.1, Fable 5, Mythos 5, Opus 4.8 and Opus 5, and notably not on Sonnet 5, where it returns a 400 saying the system role is not supported on this model. No beta header is required. It is also the injection-safe operator channel: text you embed inside a user turn to simulate an operator instruction can be forged by anything that writes user-visible input; a system-role message cannot.
The second failure mode has nothing to do with cost. Recent models follow the system prompt considerably more literally than older ones. Instructions written to overcome an earlier model's reluctance — “CRITICAL: you MUST use this tool”, “if in doubt, search” — now overtrigger. On Claude Opus 5 specifically, scaffolding that tells the model to verify its own work causes over-verification, and the documented fix is to delete those lines rather than soften them. A prompt carried forward across a model upgrade is not neutral; it encodes assumptions about a model you are no longer running.
At a glanceSee it
Where a value goes decides whether the whole conversation reprocesses at full price.
Where it runsSurfaces and availability
| Surface | Status | Notes |
|---|---|---|
| Claude Code | Yes | The harness owns the system prompt; your project instructions arrive through CLAUDE.md and similar files rather than as a parameter you set. |
| Claude API / Messages API | Yes | Top-level system, string or block list. GA, no beta header. |
| Managed Agents | Yes | A system field on the Agent object, fixed for a session’s lifetime. Session-scoped replacement is possible at creation via the agent-with-overrides form; mid-session you append rather than edit. |
| Mid-conversation system messages | Yes | Broader than previously recorded. Quoting “Mid-conversation system messages and tool changes”: “Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud.” GA, no beta header. Model-gated to Claude Fable 5.1, Mythos 5.1, Fable 5, Mythos 5, Opus 4.8 and Opus 5 — explicitly not Sonnet 5, where you use the top-level system field. Placement is constrained: the message must follow a user turn (including one carrying tool_result blocks) or an assistant turn ending in a server tool result, and must either end the array or be followed by an assistant turn; it can never be messages[0]. |
| Claude Desktop / claude.ai | Unverified | Custom instructions and project instructions occupy this role, but the mapping to the API parameter is not documented; treat behaviour as observed, not specified. |
| Claude Agent SDK | Yes | Exposed through the harness’s options rather than as a raw parameter. |
| Amazon Bedrock | Yes | The parameter works, and mid-conversation system messages work here too — Bedrock is named in the availability sentence on the mid-conversation system messages page. Automatic prompt caching is now listed for Bedrock too in that row of the “Features overview”, alongside explicit 5-minute and 1-hour caching. |
| Google Vertex AI | Yes | The parameter works, and mid-conversation system messages work here too — Google Cloud is named in the same availability sentence. Automatic prompt caching is available on Google Cloud as well. |
| Microsoft Foundry | Yes | The parameter is available. Prompt caching is GA rather than beta, including automatic prompt caching. Mid-conversation system messages are the gap: Foundry is not among the three platforms the mid-conversation system messages page lists. |
| OpenAI, Google Gemini, others | Yes | Universal in concept. Whether it is a parameter or a message role, and how it interacts with caching, differs per vendor. |
The base capability is everywhere, so it never constrains a platform choice. What varies is the escape hatch — and it is wider than the reseller-versus-first-party framing suggests. If your application learns things mid-conversation (a mode toggled, context fetched, a permission granted), you can deliver that as a system-role message and keep the cached prefix intact on the Claude API, on Amazon Bedrock, and on Google Cloud alike. The real gate is the model, not the platform: Fable 5.1, Mythos 5.1, Fable 5, Mythos 5, Opus 4.8 and Opus 5 only. On Sonnet 5, and on Microsoft Foundry whatever the model, the fallback is a marked block inside the user turn, which caches the same but is spoofable by anything that writes user input. That is a security difference, not just an ergonomic one, and it is worth weighing when you pick a model as much as when you pick a deployment target.
ExampleIn the real world
A wealth-management assistant has an 1,100-line system prompt. Roughly 200 lines are identity, tone, refusal policy and output format. About 700 are procedures: the KYC escalation ladder, the quarterly-review script, the complaint-handling sequence, the rules for discussing tax. The remaining 200 are a header interpolating today's date, the adviser's name, the client's risk tier and three feature flags.
The rewrite is mechanical. The 200 lines of identity become the frozen system prompt, in a block list with a cache breakpoint on the last block. The 700 lines of procedure become four skills, each with a description naming its trigger — the complaint skill's description mentions complaints, escalation and dissatisfaction, so it loads on those turns and stays out of every other one. The interpolated header moves out of system entirely: the date and adviser name go into the opening user turn, and the risk tier, which is fetched after the conversation starts, arrives as a system-role message appended to messages.
Measured over a week: cache reads go from zero to covering the whole 200-line core plus the tool definitions, and the complaint-handling instructions stop appearing in conversations about tax.
Not thisWhat it is often confused with
- Not a skillthe system prompt is unconditional and the skill body is conditional. That single difference decides where any given paragraph belongs: if you can name the turns where it should apply, it is a skill.
- Not a place for retrieved textpassages change per request, and the system prompt renders ahead of the entire message history. Putting retrieval output there invalidates the cache for the whole conversation and buys nothing in return.
- Not simply a messageon Anthropic's API the classic form is a top-level parameter. The newer system-role message inside
messagesis a distinct, model-gated channel with the same authority and a completely different position in the prefix. Conflating them is how caching regressions get shipped. - Not a secret storeit is not encrypted, it is not hidden from a determined user, and on Managed Agents it persists in session history that the API will hand back. Never put credentials there as a shortcut.
- Not the right home for tool guidanceinstructions about when to call a specific tool belong in that tool's own
description, where prescriptive “call this when…” wording measurably improves whether the model reaches for it.
LimitsWhen not to reach for it
- The instruction applies to a minority of turns.Move it to a skill and write a description that names the trigger. This is the single highest-value edit available to most long prompts.
- The value changes per request.Dates, session IDs, flags, user names. Put them after the last cache breakpoint, or deliver them as a system-role message on a model that supports it.
- You are compensating for an older model.Emphatic “CRITICAL: you MUST” language and self-verification scaffolding overtrigger on recent models. Re-baseline against the model you actually run before adding another guardrail.
- It is reference material rather than instruction.A schema dump, an API catalogue, four hundred product pages. That is a skill's bundled file, or retrieval.
- You want it to be tamper-proof against user input.Text placed inside a user turn to look like an operator instruction can be forged. Use the top-level parameter, or the system-role message where it is available.
Verified 2026-09-12. Stable — the shape of this is unlikely to move. Provider: Cross-vendor.