Home › Guardrails & Responsible AI › Content moderation
🛡️ · Operate

Content moderation

Filtering harmful, unsafe, or policy-violating content before and after generation.

In one line

Content moderation stops a model from generating or surfacing what it shouldn't, before it reaches a user.

ConceptWhat it is

Content moderation is the layer of classifiers and rules that screen both user input and model output for categories like violence, hate speech, sexual content, self-harm, and off-brand material. It exists because a generative model produces whatever its training and prompt make statistically likely, with no innate sense of what a company is willing to publish under its name.

Modern systems run moderation on two passes: pre-generation on the prompt, and post-generation on the completion, since a clean prompt can still yield an unsafe answer.

How it worksThe mechanics

Text passes through a purpose-built classifier such as OpenAI's Moderation endpoint, Azure AI Content Safety, or a fine-tuned model like Llama Guard, which returns per-category scores; a policy engine then blocks, rewrites, or flags for review based on thresholds tuned per category and per surface.

At a glanceSee it

Content moderation diagram
Content moderation diagram 1

Instead of a single block-or-pass gate, mature systems route each item by its highest-severity category into graduated actions — from hard block to untouched pass-through.

Content moderation diagram 2

Flagged and borderline items feed a human-review loop whose verdicts become labeled training data, continuously retuning the classifiers that produced them.

When to use itWhere it fits

  • Any consumer-facing chat or content-generation product open to the public.
  • Platforms handling user-generated prompts that could probe for unsafe completions.
  • Regulated industries where offensive or unsafe output creates legal exposure.
  • Multi-tenant apps where one tenant's content must not leak into another's session.

When NOT to use itLimits & anti-patterns

  • Purely internal, single-user developer tools where the only user is trusted staff and moderation adds latency without risk reduction.
  • Highly technical domains, like security research assistants, where over-aggressive filters block legitimate content and cause false-positive frustration.

Trade-offsAdvantages & costs

Advantages
  • Blocks the worst-case outputs before they reach an end user or get logged.
  • Off-the-shelf APIs mean no need to train a classifier from scratch.
  • Configurable thresholds let teams tune strictness per category and per surface.
Trade-offs & costs
  • Adds latency, typically 50 to 200ms per pass.
  • False positives frustrate legitimate users and false negatives still slip through.
  • Category taxonomies rarely match a company's exact policy, requiring custom rules on top.

ExampleIn the real world

Character.AI runs input and output moderation on every chat turn to catch self-harm and sexual content involving minors, escalating flagged conversations to human review and crisis resources.

ToolsHow to implement it

  • OpenAI Moderation APIfree, low-latency baseline classifier for common harm categories.
  • Azure AI Content Safetyenterprise-grade with severity scores and custom blocklists.
  • Llama Guardopen-weight model for self-hosted, customizable moderation.
  • Perspective APItoxicity scoring tuned for comment and community platforms.

Cost & effortWhat it takes

Cheap per call, often free or fractions of a cent; the real cost is engineering time tuning thresholds and building the escalation workflow for flagged content.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page Microsoft published a humanist AI code of conduct and opened a six-week public consultation on the draft.

    The Verge AI · 13 Sep 2026 · source

RecentAuto-linked from the brief, not a rewrite of this page

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning