Home › Prompt Engineering › ART (auto reasoning + tools)
✍️ · Ground

ART (auto reasoning + tools)

Automatically assembling reasoning demonstrations and tool calls from a library for each new task.

In one line

ART picks matching worked examples and tools from a library, so each new task reuses proven tool-augmented reasoning instead of a hand-written prompt.

ConceptWhat it is

ART (Automatic Reasoning and Tool-use) is a prompting technique that, given a new task, automatically retrieves worked reasoning demonstrations from a task library and pulls the relevant tools from a tool library, assembling them into a few-shot prompt for a frozen model. It exists to remove the per-task hand-crafting that plain chain-of-thought or ReAct prompts demand: instead of writing a bespoke exemplar every time, ART reuses a curated library so that tool-augmented reasoning generalizes across related tasks.

It builds on the same think-act-observe idea as ReAct, but shifts the effort from writing prompts to curating libraries. Because the model stays frozen with no fine-tuning, the technique is cheap to extend: adding a tool or a better demonstration changes behavior without retraining, and a human can edit a flawed reasoning trace to steer future runs.

How it worksThe mechanics

At inference, ART matches the incoming task to similar tasks in the demonstration library and pulls their multi-step reasoning traces, which already show where tool calls belong; it appends the relevant tool definitions and runs the frozen model, and whenever the generated trace reaches a tool-call marker it pauses generation, executes the external tool, splices the returned observation back into the trace, and resumes, repeating until the model produces a final answer.

At a glanceSee it

ART (auto reasoning + tools) diagram
ART (auto reasoning + tools) diagram 1

The reliability view — ART can break at two junctions, a missing task cluster or an unknown tool, and each is repaired library-side rather than by rewording the prompt.

ART (auto reasoning + tools) diagram 2

The improvement flywheel — a human-corrected trace becomes a reusable demo and a new tool, so an unrelated later task improves with no weight update to the frozen model.

When to use itWhere it fits

  • Repeated, tool-heavy tasks of the same shape, like calculation, search, or code execution, where a reusable library pays for itself.
  • Teams standardizing how reasoning and tool calls are done across many related tasks rather than reinventing a prompt each time.
  • Settings where the model must stay frozen and fine-tuning is off the table, but you still need reliable multi-step tool use.
  • Cases where non-engineers should fix behavior by editing an example trace instead of touching code.

When NOT to use itLimits & anti-patterns

  • One-off or highly novel tasks, where building and curating a library costs more than a single hand-written prompt.
  • Simple single-shot questions that need no tools and no multi-step reasoning at all.
  • When a strong native tool-use or reasoning model already handles the task with a plain ReAct loop or function calling.
  • Domains where you cannot assemble representative demonstrations, so retrieval returns poor matches.

Trade-offsAdvantages & costs

Advantages
  • Reuses proven reasoning-plus-tool patterns across tasks instead of one-off prompt engineering.
  • Works with a frozen model, so no training data or fine-tuning is required.
  • Extensible by design: add tools or demonstrations to the library without retraining.
  • Human-editable traces give a direct, low-code lever to correct mistakes.
Trade-offs & costs
  • Requires building and maintaining two libraries, a real upfront and ongoing cost.
  • Quality hinges on demonstration retrieval; a bad match produces a bad prompt.
  • Retrieved demos plus tool definitions inflate the prompt, raising token cost and latency.
  • It still runs a multi-step tool loop, inheriting ReAct's runaway-loop and error-propagation risks.

ExampleIn the real world

An internal team building a quantitative-reasoning assistant curates a small library of demonstrations for tasks like unit conversion, table lookup, and multi-step arithmetic, each trace showing exactly where a calculator or search tool is invoked. When an analyst asks a new but structurally similar question, ART retrieves the closest demonstrations, wires in the calculator and search tools, and runs the frozen model through the interleaved reasoning-and-tool steps, so the team gets consistent tool-grounded answers without authoring a fresh prompt per question, much as the original ART research was evaluated across the BigBench and MMLU task suites.

ToolsHow to implement it

  • DSPyprogrammatically bootstraps and selects few-shot demonstrations for a pipeline, the closest modern embodiment of ART's automatic exemplar selection.
  • LangChain and LangChain Huba shared library of prompts plus a tool registry the agent draws tools from.
  • Microsoft Semantic Kernela plugin library that a planner selects tools from to compose multi-step solutions.
  • Guidancestructured generation that can pause a reasoning trace to run a tool and resume, close to ART's execution model.

Cost & effortWhat it takes

The main cost is upfront and recurring library curation: authoring demonstration traces and wiring up tools, then maintaining them as tasks drift. Runtime cost tracks a ReAct-style loop, since retrieved demos and tool definitions enlarge the prompt and each tool call adds a round trip, so expect higher token and latency cost than a single call, offset by not re-engineering a prompt for every new task.

A living map of modern AI — kept current every morning