Citation quality is mostly a formatting property: a model can only cite an identifier you actually placed in front of it, in a form short enough to copy exactly.
Why you'd careThe thing you have already noticed
Your RAG system returns the right fact attributed to the wrong document. Or it cites a document id that does not exist in your corpus. Or it answers from the first chunk and behaves as if the other seven were not there. You go and inspect the retriever, and the retriever is fine — the answer-bearing chunk was in the prompt, ranked second. The problem lives downstream of retrieval and upstream of the model: how the chunks were deduplicated, labelled, delimited and ordered before they were handed over. A meaningful share of what teams gain from months of retrieval tuning is available immediately from formatting the evidence block properly.
In and outWhat goes in, what comes out
| In | Top-k chunks after retrieval and reranking, each carrying roughly two hundred to fifteen hundred tokens of text plus metadata: document id, title, url, page or section, timestamp, relevance score, and ideally a trust tier separating first-party content from scraped or user-supplied text. |
|---|---|
| Process | Drop near-duplicates above a similarity threshold, merge overlapping windows from the same document, truncate to a token budget, wrap each surviving span in a delimiter carrying a short stable identifier, and order the spans deliberately rather than by whatever the reranker returned. |
| Out | One evidence block of labelled spans, the identifier vocabulary the model is instructed to cite from, and — if you use one — an explicit trust marker on every span whose content originated outside your control. |
Relevance scores are almost always dropped at this boundary. Unless you print them, the model perceives no ranking at all; order is the only rank signal it receives, which is why reordering the same chunks changes the answer. Document structure usually flattens too: tables become pipe-delimited text, headings lose their level, and the cue that two paragraphs belonged to different sections disappears.
ConceptThe idea underneath
Grounded answering is not lookup. The model is still predicting the next token; a citation appears when the tokens spelling an identifier are the highest-probability continuation, which happens when that identifier is present in context, unique, and short. The relevant mechanism is the copying behaviour attributed to induction heads: attention heads that locate an earlier occurrence of the current token and promote whatever followed it. That machinery copies short exact strings reliably and long ones with errors, which is why [D3] comes back correct and a forty-character URL comes back with a mangled path.
Delimiters do a different job. They give the model a consistent segmentation cue, so the boundary between two documents is a token pattern rather than an inference from topic change. Which delimiter helps is provider- and model-specific — Anthropic documents XML-style tags for Claude, other families were post-trained on markdown-shaped documents — and the honest general rule is that a distinctive, consistent delimiter that never appears inside content beats a clever one.
Then the arithmetic of attention applies. Softmax weights sum to one, so each additional chunk reduces the weight available to the answer-bearing one. The cost is not uniform: passages semantically close to the query that do not contain the answer are far more damaging than passages that are plainly unrelated, because they compete for the same attention while supplying plausible wrong content. That is why raising top-k past its useful point degrades answers rather than leaving them flat.
Finally, this block is a trust boundary, and the format has to carry it because nothing else will. After tokenization there is no difference between an instruction you wrote and a sentence inside a retrieved web page.
At a glanceSee it
Ranked chunks are deduplicated, ordered, labelled and trust-tagged into an evidence block the model can cite.
The knobsHyperparameters and nuance
- top_k / top_n(retriever and reranker) — 3 to 10 chunks is the usual productive band. Too few and the answer is simply absent; too many and near-miss passages crowd it out while your token bill grows linearly.
- chunk_size / chunk_overlap256 to 1024 tokens with 10 to 20 percent overlap is typical. Small chunks fragment an answer across retrieval boundaries; large ones drag paragraphs of irrelevant neighbouring text along with the hit.
- deduplication thresholdcosine similarity around 0.95, or MinHash for near-exact duplicates. Too loose and overlapping windows from one document eat half the budget; too tight and you discard genuinely distinct passages that merely share boilerplate.
- span orderingscore-descending, score-ascending so the best chunk sits nearest the question, or original document order. LangChain ships
LongContextReorder, which puts the strongest chunks at both ends specifically to exploit the U-shaped position curve. - identifier schemeshort unique copyable tokens such as
[D1]beat urls and uuids for citation accuracy. Keep the mapping from identifier back to the real source in your application, not in the model's output.
EffectHow this stage moves the answer
Formatting decides which of three answer shapes you get. With per-span identifiers present and an instruction to use them, the model cites, and the citations are checkable. With sources named only in prose — according to the documentation — it produces confident attributions that are frequently to the wrong neighbouring document, because it inferred the source from proximity. With no source markers at all, it invents ones that look right, since a plausible filename is a perfectly good next-token prediction. Order changes the content itself: in a long strongest-first block, later evidence is effectively absent, so contradicting or updating information sitting at rank six never reaches the answer. Unlabelled duplicates cause their own artefact — the model reads repetition as corroboration and states one source's claim with the confidence of several.
EvalsWhat it does to your measurements
Faithfulness and citation precision move here more than almost anywhere else, and they move independently of retrieval recall, which is what makes this stage easy to misattribute. If you measure retrieval with recall@k and answers with accuracy, a formatting regression looks exactly like a model regression. The trap specific to this stage is oracle context: a harness that hands the model the gold chunk alone, correctly labelled, is measuring a pipeline you do not run and cannot expose any dedup, ordering or trust-marking defect. The second trap is order sensitivity — shuffle the same retrieved set between runs and answer accuracy typically moves by several points, so an unpinned reranker makes your eval noise floor larger than most of the improvements you are trying to detect. Report citation precision separately from answer accuracy; the two dissociate cleanly and the gap between them is diagnostic.
Failure modesWhen it goes wrong
- Right fact, wrong source citedidentifiers were attached to the block as a whole rather than to each span, so the model guessed attribution from adjacency.
- Citations to document ids that do not existno copyable identifier was present, and a plausible-looking source name is a valid next-token prediction.
- Only the first two chunks influence the answera long strongest-first block puts everything else in the weak middle of the context.
- Half the evidence budget spent on the same paragraphoverlapping retrieval windows were never deduplicated or merged.
- The model follows instructions found inside a retrieved pagethe block carries no trust marker, and after tokenization retrieved text is indistinguishable from your own.
PapersWhere this comes from
- Lost in the Middle: How Language Models Use Long ContextsLiu et al., 2023 (arXiv:2307.03172). Showed that answer accuracy depends strongly on where the answer-bearing document sits among the retrieved set, which makes span ordering a quality parameter rather than a cosmetic one.
- Large Language Models Can Be Easily Distracted by Irrelevant ContextShi et al., 2023 (arXiv:2302.00093). Introduced GSM-IC and demonstrated that adding irrelevant but topically related material reduces accuracy on problems the model otherwise solves, quantifying the cost of a loose top-k.
- The Power of Noise: Redefining Retrieval for RAG SystemsCuconasu et al., 2024 (arXiv:2401.14887, SIGIR 2024). Found that the retriever's high-scoring but non-answer-bearing passages hurt RAG accuracy more than plainly unrelated ones — unrelated documents even raised accuracy in their setup — and that placement relative to the query matters. Direct evidence that what you include and where you put it rivals retrieval quality itself.
- In-context Learning and Induction HeadsOlsson et al., 2022 (arXiv:2209.11895). Identified attention circuits that perform prefix matching and copying — completing
[A][B] ... [A]with[B]— with causal evidence in small attention-only models and correlational evidence in larger ones. Copying a short identifier out of context back into an answer is the same shape of operation, which makes this a plausible account of why terse copyable tags outperform long urls; note that the paper studies in-context learning generally and does not measure citation behaviour, so treat the link as motivation rather than a measured result.