Home › Agents & Tool Use › Reflexion
🕹️ · Build

Reflexion

Reflexion: a single agent that acts, self-checks for failure, and retries from an episodic memory.

In one line

An agent tries a task, and when a success check fails it writes a verbal lesson to memory and tries again.

ConceptWhat it is

Reflexion is a single-agent reasoning loop that turns failure into signal. Instead of accepting a first attempt or blindly resampling, the agent runs its output against a checkable success signal — unit tests, a grader, a tool result, an environment reward — and, when that signal says it failed, generates a written self-reflection explaining what went wrong. That reflection is stored in an episodic memory buffer and fed back into the next attempt, so the model improves within a single task without any weight updates.

It exists because plain retries waste the most useful thing a failure produces: a diagnosis. Introduced by Shinn et al. (2023) as verbal reinforcement learning, Reflexion treats natural-language critique as the gradient — the agent learns from its own mistakes in-context. The pattern separates three roles: an Actor that produces attempts, an Evaluator that scores them, and a self-reflection step that converts a bad score into an actionable lesson for the next round.

How it worksThe mechanics

The Actor takes the task plus any accumulated reflections and produces an attempt — an answer, a code patch, or a sequence of tool calls. The Evaluator then applies a concrete pass/fail or scalar signal: running the tests, inspecting the final environment state, comparing against a rubric. On success the loop ends and the answer is returned. On failure, a self-reflection prompt asks the model to explain why the attempt failed and what to do differently, and that short verbal lesson is appended to episodic memory. The Actor retries with memory in context, and the cycle repeats until the check passes or a retry cap is reached.

At a glanceSee it

Reflexion diagram
Reflexion diagram 1

Reflexion is really three cooperating LLM roles — actor, evaluator, self-reflector — wired to two memory tiers: a short-term trajectory that resets each try and a long-term lesson store that persists and re-primes the actor.

Reflexion diagram 2

The quality of a reflection depends on how checkable the task is — crisp pass/fail signals teach the most, while tasks with no reliable grader leave the loop with nothing solid to learn from.

When to use itWhere it fits

  • A reliable, automatable success signal exists — tests, a compiler, a grader, or an environment reward the agent can query.
  • First attempts are often close but wrong, as in code generation, math, or multi-step tool use.
  • You can trade extra latency and cost for a materially higher solve rate on hard tasks.
  • Decision-making agents that interact with an environment across repeated trials and can learn from each one.

When NOT to use itLimits & anti-patterns

  • There is no dependable way to tell success from failure — reflection has nothing solid to anchor to.
  • Latency- or cost-sensitive paths where extra retry rounds are unacceptable.
  • Simple tasks a single pass already solves — the loop just burns tokens and invites over-thinking.
  • Open-ended or subjective outputs where failure is a matter of taste rather than a check.

Trade-offsAdvantages & costs

Advantages
  • Learns within a single task from its own mistakes, with no fine-tuning or weight updates.
  • Often lifts solve rates on hard reasoning and coding benchmarks over single-shot prompting.
  • Transparent — the verbal reflections are human-readable and auditable.
  • Model-agnostic and drop-in: it wraps an existing Actor prompt rather than replacing it.
Trade-offs & costs
  • Multiplies calls and tokens — every retry is another full generation plus a reflection call.
  • Only as good as the success signal; a weak or noisy check sends reflection in the wrong direction.
  • Can over-think — spiral, second-guess a correct answer, or thrash without a firm retry cap.
  • Adds orchestration and memory state that must be built, tuned, and maintained.

ExampleIn the real world

A coding agent is asked to implement a function against a hidden unit-test suite. Its first attempt compiles but fails 3 of 12 tests. Instead of resampling blindly, the Evaluator feeds the failing test names and assertion errors into a reflection step, which writes: "I assumed the input list was already sorted, but the failing cases pass unsorted input, so I must sort first and also handle the empty-list edge case." That lesson is appended to episodic memory, the Actor rewrites the function with it in context, and the second attempt passes all 12 tests. A retry cap of three rounds keeps an unfixable task from looping forever.

ToolsHow to implement it

  • The original Reflexion reference implementation (Shinn et al., the noahshinn/reflexion repo).
  • LangGraph — its reflection and Reflexion agent tutorials build the loop as an explicit graph.
  • LlamaIndex — introspective / self-reflective agents that critique and revise their own output.
  • DSPy — Refine and assertion-driven retry-with-feedback loops implement the same idea (closely related).

Cost & effortWhat it takes

Cost scales roughly linearly with retry rounds: each failed attempt spends a full Actor generation plus a reflection call, and episodic memory grows the context on every subsequent try, so the later rounds are the priciest. Budget for something like 2 to 4 times the tokens of a single-shot baseline in the common case, and more if the cap is high. Engineering effort is moderate — the loop itself is simple, but the real work is building a trustworthy Evaluator and picking a sensible retry cap. Without a solid success check the extra spend buys little, so reserve the pattern for high-value tasks where a higher solve rate clearly justifies the added calls.

A living map of modern AI — kept current every morning