Home › Agents & Tool Use › Code-as-action (CodeAct)
🕹️ · Build

Code-as-action (CodeAct)

An agent pattern where the model's action is code it writes and runs in a sandbox.

In one line

Instead of picking from a fixed menu of tools, the agent writes and executes code to act on the world.

ConceptWhat it is

Code-as-action, usually shortened to CodeAct, is an agent pattern where each action the model takes is a snippet of executable code rather than a call to one predefined tool. Instead of choosing a single function from a fixed menu, the model writes a short program, runs it in a sandbox, and reads the result back as its observation.

It exists because real tasks rarely reduce to one atomic call: filtering a list, looping over pages, or chaining an API result into a calculation all take several steps. Expressing an action as code lets the model compose logic, control flow, and existing libraries in a single move, which is far more expressive than emitting one structured tool call at a time. The named CodeAct framework showed this can also cut the number of turns a task needs. The trade-off is unavoidable: running model-generated code demands a secure, isolated execution environment.

How it worksThe mechanics

The model receives a task plus the available libraries or API definitions, then generates a block of code as its action; a runtime executes that code inside a sandboxed interpreter, capturing stdout, return values, and any errors; the captured output is fed back into the context as an observation, and the model either writes more code to continue or refine the work or emits a final answer. Persisting variables across turns lets later code build on earlier results, and a step or time limit caps how long the write-run-observe loop can run.

At a glanceSee it

Code-as-action (CodeAct) diagram
Code-as-action (CodeAct) diagram 1

Why CodeAct pays off — one code action chains many steps in a single turn, where tool-calling spends a round trip per step.

Code-as-action (CodeAct) diagram 2

The defensive stack — blocking network, jailing the filesystem, capping resources and using throwaway containers is what turns a weak sandbox into a safe one.

When to use itWhere it fits

  • Data, math, and analysis tasks where the model must filter, aggregate, or transform structured results.
  • Workflows that chain several API calls or loop over many items, where one tool call per step would be slow and verbose.
  • Work that benefits from reusing existing libraries, such as a dataframe or plotting library, instead of hand-rolling each operation.
  • Tasks where intermediate values must persist and compose across steps.

When NOT to use itLimits & anti-patterns

  • Environments where you cannot provide a hardened sandbox, since running generated code is a genuine security risk.
  • Simple, single-shot actions that a plain tool call handles just as well with less exposure.
  • Highly regulated or deterministic flows that require every action to be a pre-approved, auditable operation.
  • Low-latency paths where spinning up an interpreter each turn adds unacceptable overhead.

Trade-offsAdvantages & costs

Advantages
  • Highly expressive: a single action can carry loops, conditionals, and multi-step logic.
  • Composes naturally with existing code libraries and APIs.
  • Often needs fewer turns than one-tool-per-step calling, because one code block does more work.
  • Errors return as real stack traces, giving the model concrete signal to self-correct.
Trade-offs & costs
  • Requires a secure, isolated sandbox; without one, generated code is an attack surface.
  • Debugging and reproducibility are harder when behavior lives in free-form generated code.
  • Wider blast radius: a bad snippet can exhaust resources, reach the network, or corrupt state.
  • Sandbox infrastructure and dependency management add real engineering and operational cost.

ExampleIn the real world

A data-analysis agent is asked which of a company's regions grew fastest last quarter. Rather than calling a fixed lookup tool repeatedly, it writes a few lines of Python that load the sales table into a dataframe, group by region, compute quarter-over-quarter growth, and sort the result. The sandbox runs the code and returns the ranked table; the agent reads it, notices two regions are nearly tied, writes a follow-up snippet to break the tie by absolute revenue, and then reports the answer alongside a small chart it generated in the same environment.

ToolsHow to implement it

  • Hugging Face smolagentsits default CodeAgent writes and executes Python as its action.
  • E2Ban open-source secure cloud sandbox purpose-built for running AI-generated code.
  • OpenAI Code Interpreterthe Assistants code-interpreter tool runs model-written Python in a sandbox.
  • Anthropic code execution toollets Claude run Python in an isolated environment as part of a turn.

Cost & effortWhat it takes

Token cost sits in the medium range: code actions are compact, so a single block often replaces several round-trips of one-tool-per-step calling, but each execution still adds a turn and the model must read outputs back, which can be large. The dominant cost is engineering and operations rather than tokens, standing up a hardened sandbox, managing dependencies and time or memory limits, and monitoring for abuse. Model-call volume is moderate and scales with how many refine-and-rerun loops a task needs, so a step cap is essential to bound both spend and blast radius.

A living map of modern AI — kept current every morning