Home › Security › Audit logs
🔒 · Operate

Audit logs

Recording enough of each answer that you can reconstruct why the system said it, months later.

In one line

An audit log for an AI system has to capture what was retrieved, not just what was asked and answered, or the decision cannot be reconstructed.

ConceptWhat it is

An audit log records what the system did in a form someone can inspect later without your help. For a conventional application that is a request, an actor and a result. For a model, the request and the result together are not enough to explain the behaviour, because the same question asked twice can be answered from different context.

The reconstructable unit is the whole turn: who asked, what was retrieved and from where, what prompt was assembled, which model and version answered, what it said, and what any guardrail did about it. Anything less and an investigation ends in a shrug.

How it worksThe mechanics

Each turn writes one structured record keyed by a trace id that follows the request across services. The retrieved chunk ids and their source documents go in, along with scores; the prompt is recorded by template id and variable values rather than as one opaque string, so it stays queryable and so secrets are not copied into the log.

Retention and access are part of the design, not an afterthought. The log now contains the sensitive material the rest of the tile is protecting — questions people asked, and passages from documents they were entitled to see. It inherits the same access control, the same encryption and the same deletion obligations as the corpus.

At a glanceSee it

Audit logs diagram

One record per turn, joined by a trace id. The retrieved chunk ids are the field that makes an answer explainable after the fact.

When to use itWhere it fits

  • Any system whose answers inform a decision someone may later contest.
  • Regulated domains, where the ability to reconstruct a decision is the requirement rather than a nicety.
  • During error analysis — the same records are the raw material the improvement loop runs on.
  • Whenever you will need to answer whether a specific document ever reached a specific person.

When NOT to use itLimits & anti-patterns

  • Logging full prompt text indiscriminately, which copies secrets and personal data into a second, usually less protected, store.
  • As a performance-monitoring substitute; a per-turn audit record is not a metrics pipeline and is expensive to aggregate.
  • In throwaway prototypes over public data, where the retention liability outweighs the value.
  • As a compliance answer on its own — recording that you did the wrong thing is not a control against doing it.

Trade-offsAdvantages & costs

Advantages
  • Turns an unexplainable answer into a reconstructable one, which is the difference between an incident and a mystery.
  • The same records feed error analysis, so the control pays for itself in product improvement rather than only in assurance.
  • Makes it possible to answer exposure questions precisely after a permissions bug, instead of assuming the worst.
  • Chunk ids plus scores expose retrieval regressions that output-only logging cannot see.
Trade-offs & costs
  • Volume is significant and grows with traffic, so retention policy is a cost decision as much as a legal one.
  • The log becomes a sensitive asset in its own right and needs every control the corpus has.
  • Deletion obligations now apply in two places, and the log is the one people forget.
  • Recording prompts carelessly is a common way to leak secrets into a system with broader read access.

ExampleIn the real world

Three months after a policy assistant gave a manager the wrong parental-leave figure, the question is whether the model was wrong or the corpus was. The record shows the retrieved chunk was from a superseded handbook still present in the index. The model reported what it was given, and the fix is an ingestion problem, which is only knowable because the chunk ids were kept.

ToolsHow to implement it

  • OpenTelemetrytrace ids that already flow through your services, so the record joins rather than sits alone.
  • Langfuse or Phoenixpurpose-built trace stores that understand retrieval and prompt structure rather than generic spans.
  • A columnar store such as ClickHouse or BigQuerywhen volume makes per-turn records something you query in aggregate.
  • Your existing SIEMfor the access-relevant subset, so security reviews one stream and not a new one.

Cost & effortWhat it takes

Storage is the visible cost and is usually modest next to inference — a rich per-turn record is kilobytes, not megabytes. The real expense is retention at scale and the engineering to keep prompts queryable without copying secrets. Effort is low if tracing already exists and moderate if it does not, since the value depends on the trace id being present everywhere.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page Presents a protocol for recording each runtime governance decision as a signed object that any party can verify offline.

    arXiv cs.AI · 24 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning