The guard is always reading a prefix, so the only real setting is how much text you are willing to show before you can still take it back.
Why you'd careThe thing you have already noticed
You watch an answer type itself out, get a paragraph and a half in, and then the whole thing disappears and is replaced by a refusal. Or the stream simply stops mid-word and the response carries a filter-shaped terminal reason instead of a normal one. Either way, a second model has been reading the output alongside you. It scores the partial completion every few dozen tokens, and when a category crosses its threshold it pre-empts the decoder mid-generation. The reason this can happen after text is already on screen is that the guard is structurally behind the generator: it is judging a prefix, so the verdict on the paragraph you just read arrives only after you have read it.
In and outWhat goes in, what comes out
| In | The completion as it is produced, in chunks — SSE delta frames or raw token IDs, commonly 16 to 256 tokens per scoring window — plus the request context the guard needs to judge them: system prompt, user turn, and sometimes the retrieved spans. |
|---|---|
| Process | A second, much smaller model scores the prefix so far against a category taxonomy. A hold-back buffer withholds the newest chunks until they have been scored. If any category crosses its threshold, the decoder is pre-empted, its KV blocks freed, and the stream terminated. |
| Out | Either the frames pass through unchanged, or the stream ends early with a typed terminal reason — content_filter on OpenAI, a refusal-typed stop reason elsewhere — optionally with the already-sent text replaced by a canned message in clients that render a turn atomically. |
Anything already flushed to the socket is unrecoverable in a token-by-token UI; you can send a retraction event, but the user has read it. The suffix that was never generated is unrecoverable in a different sense — you cannot know whether the completion would have turned benign again. And the abort tells you a category, not which span of output triggered it, so debugging a false positive means re-running with the guard in annotate-only mode.
ConceptThe idea underneath
Ordinary text classification sees a finished document. A streaming guard never does. At token t it is being asked whether text that has not been written yet will be harmful, given the prefix that has. The decision rule is release chunk t if max over categories of P(harmful | tokens 1..t) is below tau, and the quantity being estimated is a prediction about a distribution over futures, not a property of the text in hand. That is strictly harder than classifying the same completion after the fact, and it explains both characteristic error modes: over-firing on benign preambles that pattern-match to a bad continuation, and under-firing on payloads that only become harmful in the final paragraph.
The useful frame is optimal stopping, not classification. Every chunk you release lowers latency and raises regret if the verdict later flips; every chunk you hold raises latency and lowers regret. The hold-back buffer is that tradeoff made concrete: release chunk t only once chunk t+k has been scored. There is no free setting, and the choice is really a product decision about whether truncation or retraction is the less bad experience.
The systems half matters as much as the modelling half. A naive implementation re-scores the whole prefix at every interval, which is quadratic in output length. Production guards keep their own KV cache and score incrementally, so the marginal cost per chunk is one small forward pass over the new tokens. Guards are typically one to two orders of magnitude smaller than the generator, which makes this affordable but never free — and because the guard competes with the generator for the same GPU, enabling it moves inter-token latency more than the raw FLOP count suggests.
At a glanceSee it
The decoder runs ahead while the guard scores the prefix behind it and can pre-empt mid-stream.
The knobsHyperparameters and nuance
- streamProcessingModeAWS Bedrock Guardrails exposes synchronous or asynchronous. Synchronous buffers each chunk until the guard has judged it and adds latency; asynchronous lets tokens reach the user immediately and flags afterwards, which is exactly what produces disappearing text. The naming is provider-specific; the tradeoff is universal.
- scoring interval / chunk sizetokens accumulated between guard calls, typically 16 to 256. Small intervals multiply guard cost by output length and add per-chunk latency; large intervals mean much more unscored text has already shipped by the time a verdict lands.
- hold-back buffer lengthhow many scored-but-unreleased chunks you keep. Zero gives the best time-to-first-token and no ability to retract; two or three chunks makes a clean abort possible at the cost of visible stutter in the stream.
- per-category abort thresholdguards emit a score per category, not one number. Tightening a single category is the usual cause of a sudden over-refusal spike, and because it is one config value it almost never appears in a changelog.
- abort behaviourtruncate in place, replace the entire turn with a canned refusal, or roll back to the last safe chunk. Replacement requires a client that renders turns atomically; in a token-streaming UI, truncation is the only honest option.
EffectHow this stage moves the answer
Long answers are affected out of all proportion, because every additional chunk is another chance for a threshold to trip. A 200-token reply is almost never aborted; a 3000-token security write-up, a red-team summary, or a code file containing a plausible-looking credential is a much fatter target. What the user sees is a coherent answer that ends abruptly, sometimes mid-word, with no explanation beyond a terse notice. If the turn was structured output, the abort lands inside a JSON object and the parser receives a fragment. The second-order effect is larger and almost never measured: once a team learns which topics trip the guard, prompts get rewritten to steer around them, and the model's usable output distribution narrows in ways nobody logged. In annotate-after modes the answer is not truncated at all — it is fully rendered and then withdrawn, which users experience as the product retracting something it already said.
EvalsWhat it does to your measurements
Two metrics move in opposite directions and you need both: harmful-completion rate and over-refusal rate. XSTest is the standard probe for the second — a suite of benign prompts that look superficially unsafe — and it is the benchmark most sensitive to an output guard whose threshold has been tightened. On the latency side, the hold-back buffer shows up in inter-token latency and time-to-last-token, not in time-to-first-token, so a benchmark reporting only TTFT will score a guard change as free. The silent invalidation here is mode mismatch: many guards run only on streaming responses, so an eval harness that requests non-streaming completions never exercises them and reports the ungated model's quality. Cost accounting drifts too, since providers differ on whether tokens generated before an abort are billed, which makes cost-per-eval move without a single prompt changing.
Failure modesWhen it goes wrong
- Text appears, then vanishes and is replaced by a refusalthe guard is running in annotate-after mode, so chunks ship before any verdict exists.
- Answers stop mid-word at roughly the same length every timea scoring-window boundary combined with a threshold tripped by a recurring section of the output, routinely misdiagnosed as a max_tokens cap.
- Streaming JSON arrives truncated and unparseablethe abort landed inside an object and the client had no branch for the terminal reason.
- p99 inter-token latency spikes while p50 stays flatthe hold-back buffer plus guard forward passes contending with the generator for the same GPU.
- The guard misses harm that only exists in the last paragrapha prefix classifier is predicting the future from a benign preamble, and a short hold-back leaves nothing still in hand to re-score.
PapersWhere this comes from
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI ConversationsInan et al., 2023 (arXiv:2312.06674). Introduced a small instruction-tuned safeguard model that applies a single safety risk taxonomy to two distinct tasks — prompt classification and response classification — with different guidelines supplied for each. It is the open reference implementation of the second-model-reads-the-answer pattern this stage runs.
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingSharma et al. (Anthropic), 2025 (arXiv:2501.18837). Trained input and output classifiers on synthetic data generated from an explicit constitution, and describes an output classifier that supports streaming prediction — judging the harmfulness of the eventual full output at each token, without waiting for generation to finish. That is the specific capability that makes mid-generation abort possible rather than post-hoc filtering.
- A Holistic Approach to Undesired Content Detection in the Real WorldMarkov et al., AAAI 2023 (arXiv:2208.03274). Documents the taxonomy, data pipeline and active-learning loop behind a production moderation classifier; the part that matters here is its treatment of threshold selection and per-category precision, which is what you are actually tuning.
- No literature on the streaming tradeoff itselfthe hold-back-versus-latency decision has no published treatment to cite. The closest reference material is provider documentation on streaming guardrail modes and the model cards for open guard models. Treat buffer length and thresholds as engineering choices to be measured locally, not as settled results.