Home › What Happens After You Hit Enter › System-prompt assembly
Pipeline stage · Operate

System-prompt assembly

Before the user's words, the system stacks its own: identity, rules, tool policy, current date, locale, safety instructions — assembled fresh per request.

In one line

The system prompt is rebuilt from a template on every request, and where you put its one volatile field decides whether thousands of tokens are cached or repriced.

Why you'd careThe thing you have already noticed

Ask a hosted assistant what today's date is and it knows, with no tool call. Ask it something benign and it occasionally declines in language you never wrote. Ship a one-line change and watch the bill and the time-to-first-token both jump while the answers stay the same. All three are this stage. Before your user's first word there is a preamble — identity, house style, refusal policy, tool policy, the current date, locale, and the full JSON schema of every enabled tool — assembled fresh per request and often several thousand tokens long. You never see it in your request payload because you did not write most of it, and it is doing more to shape the answer than your prompt is.

In and outWhat goes in, what comes out

InA versioned base template, plus per-request variables: ISO timestamp, timezone, locale, account tier, feature flags, the enabled tool list with full JSON schemas, the safety policy revision, and any persistent memory blocks the platform maintains for this user.
ProcessRender the template, order layers from most stable to most volatile, append tool schemas, place a cache breakpoint immediately after the last byte that is identical across requests, count the result against the model's context budget, and stamp a prompt-version id for logging.
OutA single system or developer message, commonly one to ten thousand tokens once tool schemas are included, plus cache-control metadata marking the reusable prefix and a version identifier that should appear in every request log and every eval record.

What is preserved is the text. What is lost is provenance: once flattened, platform policy, application instructions and per-user memory are indistinguishable byte ranges at different offsets. The model cannot tell which layer a sentence came from and therefore cannot apply different trust to it, which is why two contradictory rules resolve unpredictably rather than by precedence.

ConceptThe idea underneath

A decoder-only language model computes p(x_t | x_1 ... x_(t-1)): every token's distribution is conditioned on everything before it. A system prompt is not a mode switch or a flag, it is just earlier tokens. Its influence is exactly the influence any prefix has, which is why it can be argued with, diluted by a long context, and overridden by a sufficiently insistent later instruction.

Because attention is causal, the keys and values computed for prefix tokens do not depend on anything that comes after them. That is what makes a prefix cacheable: the server stores the per-layer K and V tensors for the first N tokens and reuses them verbatim next time. It is also why reuse is prefix-exact. Change one token at position 12 and every K and V from position 12 onward is wrong, so a timestamp near the top of the preamble invalidates the entire thing on every single request.

Early positions are not neutral. Work on attention sinks showed that the first few tokens of a sequence absorb a large share of attention mass across heads and layers, acting as somewhere for heads to park probability when they have nothing specific to attend to. Combine that with the normalisation in softmax(QK^T / sqrt(d_k)) V, where every extra context token takes weight away from every other, and the very beginning of the prompt is structurally privileged while the middle is structurally weak.

None of this gives the system role runtime protection. There is no memory boundary between instruction and data; the priority a model gives to system text is a behaviour trained in during post-training, which is what the instruction-hierarchy line of work is trying to make reliable.

At a glanceSee it

System-prompt assembly diagram

Stable template, tool schemas and volatile fields merge into one preamble split by a cache breakpoint.

The knobsHyperparameters and nuance

  • system(Anthropic Messages API), instructions (OpenAI Responses API), or a developer-role message (OpenAI Chat Completions) — the dedicated slot. Putting the same text in a leading user message instead measurably weakens adherence and, on some providers, changes safety behaviour.
  • cache_controlwith type ephemeral (Anthropic) — marks a prefix breakpoint. A small number of breakpoints per request, and a documented minimum cacheable prefix length on the order of one to two thousand tokens, model-dependent. Placed after a volatile field it never hits; placed too early it caches far less than it could.
  • toolsthe tool schema array is usually the largest and most volatile part of the preamble. Adding one tool, or merely reordering the array, changes the prefix and drops the cache for every request that follows.
  • --enable-prefix-caching(vLLM; on by default in recent versions) — the self-hosted equivalent, matching on hashed block boundaries rather than explicit breakpoints. Off, and a shared system prompt is re-prefilled for every request in the batch.
  • timestamp granularity in the templatenot an API parameter, and usually the single most expensive line in the preamble. A date-only variable changes once a day; a second-resolution timestamp changes every request and guarantees a cold prefix forever.

EffectHow this stage moves the answer

Length buys control and spends it. A longer preamble makes formatting and refusals more consistent while making the assistant more cautious in ways nobody asked for: a rule written for one edge case fires on adjacent benign requests, and the observable symptom is a refusal whose wording appears nowhere in your code. Contradictory layers are worse than either layer alone — when the platform prompt says be concise and the application prompt asks for detailed explanations, the resolution is not a precedence rule but whichever phrasing the model finds more salient that moment, so the same request gets different lengths on different days. Tool schemas compete for the same budget: past a few dozen tools the model starts choosing plausible-sounding wrong tools, because tool selection is a reading-comprehension problem over a long list. Style bleeds too — examples in the preamble reappear as templates in unrelated answers.

EvalsWhat it does to your measurements

The system prompt is the highest-leverage variable in most stacks and the least often recorded. Refusal rate, format compliance and instruction-following scores all move with it, and none of them move in a way that looks like a prompt change; they look like a model regression. Record the prompt version id on every eval row, because two runs a week apart are otherwise not comparable. Two specific traps. First, harnesses commonly send no system prompt at all while production sends four thousand tokens, so your offline numbers describe a system you do not run. Second, a second-resolution timestamp makes every request's prefix unique, driving measured cache hit rate to zero and inflating both latency and cost benchmarks — and if you then freeze the timestamp for reproducibility, your latency numbers become cache-warmed and stop describing production.

Failure modesWhen it goes wrong

  • Cost and time-to-first-token double overnight with no change in output qualitya volatile field moved above the cache breakpoint, so the whole prefix is re-prefilled on every request.
  • The assistant refuses a request that worked last weeka safety line added for one case generalised to its neighbours; nothing in your application code changed.
  • A tool is never selected despite an accurate descriptiona sentence in the tool policy contradicts the schema, or the tool list is long enough that selection itself degrades.
  • Two deployments of the same model behave differentlydifferent preamble versions, or different per-request variables such as locale or account tier.
  • Long conversations hit the context limit sooner than arithmetic suggeststhe preamble plus tool schemas are a fixed cost paid before any history exists.

PapersWhere this comes from

  • The Instruction Hierarchy: Training LLMs to Prioritize Privileged InstructionsWallace et al., 2024 (arXiv:2404.13208). Established that role priority is a training objective rather than a property of the message format, and measured how much of it survives adversarial later instructions.
  • Lost in the Middle: How Language Models Use Long ContextsLiu et al., 2023 (arXiv:2307.03172). Measured accuracy as a function of where the relevant material sits and found it highest at the very start or the very end of the input, degrading in between — the empirical basis for treating the top of a preamble as privileged real estate. Note that this is a measured behavioural regularity, not a mechanistic explanation; no published work establishes why primacy holds.
  • Instruction-Following Evaluation for Large Language ModelsZhou et al., 2023 (arXiv:2311.07911). Introduced IFEval, a set of automatically verifiable constraint prompts that make preamble-induced drift measurable rather than anecdotal.
A living map of modern AI — kept current every morning