Home › Evals & Testing › Key term › Red-teaming
Key term · Operate

Red-teaming

Deliberately attacking a model to surface unsafe or broken behavior before users do.

In one line

Red-teaming is paying people or models to try to break your model on purpose.

DefinitionWhat it means

Red-teaming is the structured practice of deliberately probing a model or product with adversarial, edge-case, or malicious inputs, run by internal specialists, external experts, or automated attacker models, to surface unsafe, biased, or failing behavior before it reaches real users.

Why it mattersWhy you should care

Regulators, enterprise buyers, and safety frameworks increasingly expect documented red-team results before an AI system launches or scales, so red-teaming has moved from a nice-to-have research exercise to a standard release gate alongside functional testing for any consumer-facing or high-stakes AI product.

At a glanceSee it

Red-teaming diagram
Red-teaming diagram 1

Opening up the adversarial-prompt black box — the main families of attack technique, split between human-crafted tricks and machine-generated exploits.

Red-teaming diagram 2

What fix-and-retest really involves — triage each finding by whether it is exploitable now, mitigate the blockers, add them to a regression suite, then loop into the next red-team round.

Where you see itIn the wild

  • Pre-launch safety reviews at frontier labs before a model release.
  • Bug-bounty style programs paying researchers to find jailbreaks.
  • Planning discussions on designing an internal red-team process.
A living map of modern AI — kept current every morning