Home › Agents & Tool Use › Reflection / critique
🕹️ · Build

Reflection / critique

A critic reviews the agent's draft and the agent revises before returning a final answer.

In one line

Generate a draft, critique it against explicit criteria, then revise, trading an extra pass for higher-quality output.

ConceptWhat it is

The reflection pattern (also called self-critique or critique-and-revise) inserts a review step between an agent's first attempt and its final answer. A generator produces a draft, a critic — often the same model prompted differently, sometimes a separate model or an objective check like a test suite — inspects it for errors, gaps, or policy violations, and the agent then revises before finishing.

It exists because a model's first pass is frequently good but rarely its best. Forcing an explicit evaluation stage surfaces mistakes the generator glossed over and gives the system a structured chance to fix them, rather than shipping whatever fell out of a single forward pass.

How it worksThe mechanics

The generator produces an initial output; the critic then evaluates it against explicit criteria — a checklist, a rubric, unit tests, or a plain instruction to find flaws — and emits concrete feedback rather than a bare pass or fail. If the feedback flags problems, the agent revises the draft using that feedback and the critic reviews again, looping until the output passes or a maximum number of rounds is reached, at which point the best revision is returned.

At a glanceSee it

Reflection / critique diagram
Reflection / critique diagram 1

Reflection is only as reliable as its critic — an objective check beats a model grading its own work, which tends to rubber-stamp its own mistakes.

Reflection / critique diagram 2

The loop needs a governor — an iteration budget and a saved best draft stop revision from looping forever or quietly degrading the answer.

When to use itWhere it fits

  • Quality-sensitive outputs where an obvious error is costly, such as code, legal or financial text, or customer-facing copy.
  • Tasks with checkable criteria, like passing tests, matching a schema, or satisfying a rubric.
  • Long or complex generations where a single pass reliably misses edge cases.
  • Situations where one or two extra passes are affordable in exchange for fewer downstream fixes.

When NOT to use itLimits & anti-patterns

  • Latency- or cost-sensitive paths where an extra round-trip is unacceptable.
  • Simple, low-stakes outputs the model already gets right on the first try.
  • Tasks with no checkable signal, where an unreliable critic may reject good work or approve bad work.
  • Cases where a cheaper fix — a better prompt, a schema validator, or a tool — removes the error more directly.

Trade-offsAdvantages & costs

Advantages
  • Catches obvious mistakes and gaps before they reach the user.
  • Raises average output quality without retraining the model.
  • Makes quality criteria explicit and inspectable in the critic prompt.
  • Composes with tools such as tests, linters, and validators as objective critics.
Trade-offs & costs
  • Every review-and-revise round adds latency and token cost.
  • The critic can be wrong, over-flagging good work or waving through real errors.
  • Self-critique with the same model can share the generator's blind spots.
  • Needs a stopping rule, since loops can churn without ever converging.

ExampleIn the real world

A code-generation assistant drafts a function, then runs a critic pass that reads the draft against the task description and the project's test suite. The critic reports that an edge case for empty input is unhandled; the assistant revises to add the guard, the tests pass on the second round, and the reviewed version is returned instead of the flawed first draft.

ToolsHow to implement it

  • LangGraphbuild an explicit generate-then-critique graph with a reviser node and a loop-back edge.
  • AutoGenmulti-agent setups where a reviewer agent critiques a writer agent's output.
  • Reflexiona research technique where an agent writes verbal self-reflection on failures and retries.
  • DSPy assertionsdeclare constraints that trigger self-correction and retries when an output violates them.

Cost & effortWhat it takes

Reflection roughly multiplies cost by the number of passes: a single critique-and-revise cycle typically adds two to three extra model calls beyond the base generation, and more if the loop runs several rounds. Engineering effort is modest — a critic prompt or rubric, a revise step, and a stopping rule — but most of the work goes into tuning the critic so it is neither too harsh nor too lenient, and capping the loop so it terminates.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page A study finds that agentic scaffolding such as reflection amplifies sycophancy and lowers accuracy, so reflection is not a guaranteed quality improvement.

    Revise sub/agt-reflection to note that iterative critique can amplify sycophancy and reduce accuracy (a study reports a 6.3-point drop), so an extra pass does not guarantee higher-quality output.

    arXiv cs.CL · 24 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning