Home › What Happens After You Hit Enter › Repetition, frequency and presence penalties
Pipeline stage · Operate

Repetition, frequency and presence penalties

Docking points from tokens the model has already used, so it stops saying the same thing.

In one line

Penalties never look at meaning, only at token ids and counts, so anything tuned to stop a repeated phrase also punishes the closing brace your JSON needs.

Why you'd careThe thing you have already noticed

You had a summariser that looped, repeating the same sentence with small variations until it hit max_tokens. Someone set frequency_penalty to 1.0 and the loop stopped. Two weeks later the same endpoint is writing “the aforementioned methodology hitherto utilised” and occasionally dropping the word “the” altogether. Both are this stage. The penalty operates on token ids and counts, not on meaning. It does not know that the appears four hundred times because English requires it, or that } appears often because JSON has nesting. It docks whatever it has already seen, and the tokens it has already seen most are exactly the ones the language needs most.

In and outWhat goes in, what comes out

InThe current logits vector, plus per-sequence occurrence data: a set of seen token ids for the presence term, integer counts per id for the frequency term, and the union of prompt and output ids for repetition_penalty. Which of prompt or output is counted is engine-specific.
ProcessApplied in place before temperature scaling, as one fused kernel over the batch. Repetition penalty divides positive logits by θ and multiplies negative ones by θ. Frequency and presence penalties subtract alpha × count and beta × 1[count > 0] respectively across the whole vocabulary.
OutA same-length logits vector with previously emitted ids depressed, handed to the truncation and sampling stage. Nothing else in request state changes except the running count table, updated after the token is actually chosen.

What survives is the relative ordering among tokens the model has not yet emitted. What is destroyed is the model's calibrated opinion about repeating. If the correct next token is one already used — a variable name, a closing bracket, a person's surname in a biography — the penalty erases the evidence that it was correct, and no stage downstream can distinguish a legitimate repeat from a degenerate one.

ConceptThe idea underneath

Two different formulations are in circulation and they behave differently. The CTRL formulation is multiplicative: z_i' = z_i / theta if z_i > 0 else z_i * theta, with θ around 1.0 to 1.2. Note the asymmetry — dividing a positive logit and multiplying a negative one both push toward zero, so this is a shrinkage toward the middle of the logit range rather than a fixed subtraction. That makes it sensitive to the model's logit scale, which is why θ=1.15 can be gentle on one model and destructive on another.

The OpenAI-style formulation is additive: z_i' = z_i - alpha_frequency * count_i - alpha_presence * 1[count_i > 0]. Additive in logit space is multiplicative in probability space, and it is independent of logit scale, so the same numbers transfer between models far better. presence fires once per token regardless of count, so it pushes topic diversity; frequency grows linearly with count, so it pushes word diversity and gets harsher the longer you generate.

Why the problem exists at all is a training question, not a decoding one. Welleck et al. showed that the maximum-likelihood objective assigns probability to repetition because human text does repeat, and that low-temperature decoding turns a small self-reinforcing bias into a loop: a phrase appears, attention to it makes it more likely, it appears again. Unlikelihood training fixes this at the source. Penalties are a hand-tuned, non-learned patch applied at inference to a defect created during training.

That mismatch is the key insight. The model's distribution is conditioned on the context; the penalty adds a second, hidden conditioning on a count table the model has never been trained with. That is why the failure mode reads as unnatural rather than merely wrong.

At a glanceSee it

Repetition, frequency and presence penalties diagram

Counts of already-emitted tokens become a penalty vector subtracted from the logits before sampling.

The knobsHyperparameters and nuance

  • repetition_penalty(HuggingFace generate, vLLM) — 1.0 is off, 1.05–1.15 is useful, above 1.2 vocabulary visibly degrades. Multiplicative and logit-scale-sensitive. Most implementations count prompt tokens too, which is why it can stop the model quoting its own input. llama.cpp implements the same idea under a different name: repeat_penalty in the server JSON API (default 1.1), --repeat-penalty on the command line. Sending repetition_penalty to llama.cpp is ignored, not applied.
  • frequency_penalty(OpenAI-compatible) — −2.0 to 2.0, default 0. Above roughly 0.5 you get rare-word drift; above 1.0 you get dropped function words. Negative values force repetition, which is occasionally the right tool for a rigid template.
  • presence_penalty(OpenAI-compatible) — −2.0 to 2.0, default 0. Fires once per distinct token, so 0.3–0.6 broadens a brainstorm. Above 1.0 the model drifts away from the question because staying on topic requires reusing its nouns.
  • no_repeat_ngram_size(HuggingFace generate) — a hard ban on repeating any n-gram, typically 3 or 4. Catastrophic for code, tables, and any text containing a repeated proper noun; it makes correct output literally unreachable rather than merely unlikely.
  • repeat_last_n(llama.cpp server API; --repeat-last-n on the CLI) — how many recent tokens feed the count table. Default 64, 64–256 typical, 0 disables and -1 means the entire context. Applying it over the whole context is almost always wrong on long generations, because early legitimate repeats accumulate into a permanent handicap. penalty_last_n is the internal C++ field name for the same setting, not a parameter you can send.

EffectHow this stage moves the answer

Structured output degrades first, and fastest. A JSON array whose objects all share the keys name and value is, to a frequency penalty, a runaway repetition; by the fourth element the penalty is actively fighting your schema and the model starts renaming keys or omitting them. Code degrades second, because identifiers, keywords, brackets and indentation tokens are all repeats by definition — output loses closing braces and consistent indentation. Prose degrades last but most visibly: the model reaches for a synonym it barely knows, then for one that does not exist, then drops articles and prepositions because those have the highest counts of all. From the outside this looks like the model got worse, or got “more creative”. It is neither. It is being scored on an objective it was never trained against. The practical rule is simple: never apply penalties to a JSON, tool-calling or code endpoint.

EvalsWhat it does to your measurements

Penalties move exactly the metrics they were designed to move — rep-n, distinct-n, self-BLEU — and quietly cost you exact match everywhere else. On GSM8K the correct answer usually restates numbers and entities from the question; a frequency penalty pushes the model to paraphrase them and the string match fails even when the reasoning was right. HumanEval drops because function signatures echo the prompt. The dangerous pattern is organisational: a team tunes penalties against a chat-quality eval, ships one sampling config across every endpoint, and the structured-output suite loses two points with no code change and no obvious cause. Penalties must be evaluated per endpoint, not per model, and their values must be logged alongside every score. If you inherit a benchmark number without knowing the penalty settings, you do not know what was measured.

Failure modesWhen it goes wrong

  • Bizarre synonyms and missing articlesrepetition_penalty above about 1.2, or frequency_penalty above about 1.0; the highest-count tokens are the function words.
  • JSON goes malformed a few array elements inthe presence or frequency penalty is punishing the repeated key set that the schema requires.
  • The model refuses to quote text from its own promptthe repetition penalty is counting prompt tokens, so verbatim extraction is penalised by construction.
  • Loops reappear after switching providersrepetition_penalty was sent to an OpenAI-compatible endpoint that only reads frequency_penalty and presence_penalty, and the unknown field was silently dropped.
  • Generated code loses indentation or closing bracketswhitespace and bracket tokens carry the highest counts in any code output, so they are penalised hardest.

PapersWhere this comes from

  • CTRL: A Conditional Transformer Language Model for Controllable GenerationKeskar et al., 2019 (arXiv:1909.05858). Introduced the multiplicative repetition penalty as “penalized sampling” in Section 4.1 of the main text, discounting the scores of already-generated tokens with a suggested factor of about 1.2. That formulation is now the default implementation in nearly every open inference engine, which is why the parameter is scale-sensitive rather than additive.
  • Neural Text Generation with Unlikelihood TrainingWelleck et al., 2019 (arXiv:1908.04319). Showed that repetition is a property of the maximum-likelihood objective itself, not merely of the decoding strategy, and offered a training-time fix. It explains why an inference-time penalty is a patch rather than a solution.
  • The Curious Case of Neural Text DegenerationHoltzman et al., 2019 (arXiv:1904.09751). Quantified how badly maximisation-based decoding degenerates and provided the measurement vocabulary (repetition rate, distinct-n) that penalties are still tuned against.
A living map of modern AI — kept current every morning