Your messages array is not what the model sees: a Jinja template shipped with the weights flattens it, and the wrong template costs more quality than any sampling knob.
Why you'd careThe thing you have already noticed
You ran the same open-weight model through two servers with the same messages and the same temperature. One gave clean answers; the other ignored the system prompt, ran past the end of its reply and started writing your next question, or printed <|eot_id|> into the output. Nothing about the weights differed. What differed was the template — the small Jinja program that turns your role-tagged message list into the single flat string the model was post-trained on. Templates are per-model, they ship inside tokenizer_config.json, and they are the thing wrapper libraries most often get subtly wrong. Hosted APIs hide this stage completely, which is why the problem tends to appear the first time you self-host.
In and outWhat goes in, what comes out
| In | An ordered message list with roles — system or developer, user, assistant, tool — plus tool definitions as JSON Schema, any prior tool-call ids and their results, and template flags such as add_generation_prompt and, on reasoning models, whether past thinking blocks are retained. |
|---|---|
| Process | The tokenizer executes the model's Jinja2 template: emit BOS, wrap each turn in that family's control tokens, serialize tool schemas into the syntax the model was trained to read, drop or merge roles the template has no slot for, and append the assistant header. |
| Out | One flat string, or token ids directly, in which role markers are reserved single-id control tokens that cannot be produced by encoding their literal text. The final tokens are the assistant header, so the next prediction is the first token of the reply. |
Role structure survives only as surface markers. After rendering there is no privileged channel: a system instruction and a pasted email are the same kind of bytes at different offsets, which is the structural reason prompt injection works at all. Anything the template has no slot for — a system message on a family whose template lacks one, tool-call ids, prior reasoning blocks — is silently gone, and your application is not told.
ConceptThe idea underneath
This stage has no ML content of its own; it is string formatting. Its power comes from what happened during post-training. Supervised fine-tuning and preference optimisation are run on sequences already laid out in exactly this format, so the model's helpful, instruction-following behaviour is conditioned on the surface markers that precede it. Seeing <|start_header_id|>assistant<|end_header_id|> is what makes the next tokens an answer rather than a continuation.
Feed the same content in a format the model never saw and you have not broken anything, you have moved the input off the post-training distribution. The model falls back toward base language-model behaviour: completing text, mirroring the prompt's style, writing both sides of the conversation. The degradation is graceful and therefore easy to miss, because the output is still fluent English.
Control tokens are reserved vocabulary ids, not strings. <|im_end|> as an id is a single integer the model was trained to emit; the same characters typed by a user encode to several ordinary tokens. That asymmetry is deliberate — it is the only thing stopping a user from forging a turn boundary — and it is why templates must be applied through the tokenizer rather than by string concatenation.
The formats themselves are conventions, not standards. ChatML wraps turns in <|im_start|>role and <|im_end|>; Llama 3 uses header-id markers; Mistral's instruct format uses [INST] and [/INST]; Gemma uses start-of-turn and end-of-turn markers. Several of these historically had no system-role slot at all, so a system message was merged into the first user turn or dropped outright. Read the template that ships with the specific checkpoint rather than assuming a family behaves like its cousins.
At a glanceSee it
A role-tagged message list is flattened by the checkpoint's Jinja template into control tokens the model recognises.
The knobsHyperparameters and nuance
- chat_template(field in
tokenizer_config.json, applied bytokenizer.apply_chat_template) — the Jinja source itself. Pin it to the exact checkpoint revision. A template copied from an earlier revision of the same model is the most common silent regression in self-hosted stacks. - add_generation_prompt(bool, defaults to false in Hugging Face transformers) — appends the assistant header. False at inference time and the model continues the user's turn instead of replying; true when scoring an already-complete conversation and you have appended a header that was never in the training data.
- add_special_tokens(tokenizer argument) — whether to prepend BOS on top of what the template already emitted. Left true after
apply_chat_templateit produces a double BOS, which degrades several Llama-family checkpoints in a way that looks like a bad seed. - continue_final_message(transformers, also exposed by vLLM's OpenAI-compatible server) — resumes an open assistant turn instead of starting a new one, by stripping the final message's end-of-turn tokens. This is the knob for prefilling an answer. It is mutually exclusive with
add_generation_prompt, which adds the tokens that begin a new message: transformers rejects the combination with an error rather than silently emitting two assistant headers, and vLLM's server rejects it the same way. A caller that sets both fails loudly at request time. - --chat-template(vLLM OpenAI-compatible server) — overrides the checkpoint's template from a file path or an inline string. Necessary for base models that ship none, where chat requests otherwise error outright; dangerous for instruct models that already have a correct one.
- enable_thinking(template keyword argument on some reasoning models, Qwen3 among them, where it defaults to true) — whether the template opens a thinking block, and how prior thinking is handled in history. Retaining thinking across turns inflates every later turn's prompt cost.
EffectHow this stage moves the answer
A wrong template changes the class of output, not its polish. The most visible failure is stopping: the model was trained to end its turn with a specific control token, and if the turn was never opened correctly it does not emit that token, so you get an answer followed by an invented user question followed by another answer. The second is instruction adherence — on families whose template has no system slot, your carefully written system prompt is concatenated into the user turn or discarded, and the model reads your rules as something the user said, which it will happily override later. The third is tool use: templates serialize tool schemas in a family-specific syntax, and a mismatch produces JSON printed as prose in the content field rather than a parsed tool call, so your dispatcher sees no tool call at all and the agent loop stalls on step one.
EvalsWhat it does to your measurements
Template mismatch is the leading explanation for an open-weight model missing its published numbers on your own harness. It moves instruction-following scores hardest — IFEval-style verifiable-constraint checks collapse when the model does not know its turn has started — and it moves anything with an exact-match answer format, because the extra invented turn defeats the answer parser. The silent invalidation is comparative: a harness that applies its own generic template to every model is measuring each model's tolerance for the wrong format, not its capability, and will rank a family whose native format resembles the harness default above one that does not. Log the rendered prefix, not the messages array, for at least one sample per run; a template regression is invisible in the request payload and obvious in the rendered string. Cached-prefix metrics break here too, since any template change moves every token id.
Failure modesWhen it goes wrong
- The model answers, then writes your next question and answers that toothe turn was never opened correctly, so the end-of-turn token is off-distribution and the server's stop-token list does not match either.
- Control tokens such as
<|eot_id|>appear in the visible outputthey were rendered as literal text rather than encoded as reserved ids, or the detokenizer was not told to skip special tokens. - The system prompt has no observable effectthe family's template has no system slot and merged it into the first user turn or discarded it.
- Quality is slightly but consistently below the model carddouble BOS, or a template revision that does not match the weights revision.
- Tool calls arrive as prose JSON in the content fieldthe template serialized tools in a syntax the model was not trained on, or the server's tool-call parser expects a different one.
PapersWhere this comes from
- No research literature defines chat templatesthe format is a per-model convention, and the authoritative artefact is the
chat_templateJinja source inside each checkpoint'stokenizer_config.json, with Hugging Face transformers'apply_chat_templateas the reference implementation. Read the template, not a blog post about the family. - Training language models to follow instructions with human feedbackOuyang et al., 2022 (arXiv:2203.02155). Established the supervised-fine-tuning-plus-preference-optimisation recipe that makes instruction following a learned property of formatted training sequences, which is why format mismatch costs capability rather than just tidiness.
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged InstructionsWallace et al., 2024 (arXiv:2404.13208). Showed that role priority has to be trained in rather than enforced by the format, the precise reason a rendered template gives the system role no runtime protection.