ART picks matching worked examples and tools from a library, so each new task reuses proven tool-augmented reasoning instead of a hand-written prompt.
ConceptWhat it is
ART (Automatic Reasoning and Tool-use) is a prompting technique that, given a new task, automatically retrieves worked reasoning demonstrations from a task library and pulls the relevant tools from a tool library, assembling them into a few-shot prompt for a frozen model. It exists to remove the per-task hand-crafting that plain chain-of-thought or ReAct prompts demand: instead of writing a bespoke exemplar every time, ART reuses a curated library so that tool-augmented reasoning generalizes across related tasks.
It builds on the same think-act-observe idea as ReAct, but shifts the effort from writing prompts to curating libraries. Because the model stays frozen with no fine-tuning, the technique is cheap to extend: adding a tool or a better demonstration changes behavior without retraining, and a human can edit a flawed reasoning trace to steer future runs.
How it worksThe mechanics
At inference, ART matches the incoming task to similar tasks in the demonstration library and pulls their multi-step reasoning traces, which already show where tool calls belong; it appends the relevant tool definitions and runs the frozen model, and whenever the generated trace reaches a tool-call marker it pauses generation, executes the external tool, splices the returned observation back into the trace, and resumes, repeating until the model produces a final answer.
At a glanceSee it
The reliability view — ART can break at two junctions, a missing task cluster or an unknown tool, and each is repaired library-side rather than by rewording the prompt.
The improvement flywheel — a human-corrected trace becomes a reusable demo and a new tool, so an unrelated later task improves with no weight update to the frozen model.
When to use itWhere it fits
- Repeated, tool-heavy tasks of the same shape, like calculation, search, or code execution, where a reusable library pays for itself.
- Teams standardizing how reasoning and tool calls are done across many related tasks rather than reinventing a prompt each time.
- Settings where the model must stay frozen and fine-tuning is off the table, but you still need reliable multi-step tool use.
- Cases where non-engineers should fix behavior by editing an example trace instead of touching code.
When NOT to use itLimits & anti-patterns
- One-off or highly novel tasks, where building and curating a library costs more than a single hand-written prompt.
- Simple single-shot questions that need no tools and no multi-step reasoning at all.
- When a strong native tool-use or reasoning model already handles the task with a plain ReAct loop or function calling.
- Domains where you cannot assemble representative demonstrations, so retrieval returns poor matches.
Trade-offsAdvantages & costs
Advantages
- Reuses proven reasoning-plus-tool patterns across tasks instead of one-off prompt engineering.
- Works with a frozen model, so no training data or fine-tuning is required.
- Extensible by design: add tools or demonstrations to the library without retraining.
- Human-editable traces give a direct, low-code lever to correct mistakes.
Trade-offs & costs
- Requires building and maintaining two libraries, a real upfront and ongoing cost.
- Quality hinges on demonstration retrieval; a bad match produces a bad prompt.
- Retrieved demos plus tool definitions inflate the prompt, raising token cost and latency.
- It still runs a multi-step tool loop, inheriting ReAct's runaway-loop and error-propagation risks.
ExampleIn the real world
An internal team building a quantitative-reasoning assistant curates a small library of demonstrations for tasks like unit conversion, table lookup, and multi-step arithmetic, each trace showing exactly where a calculator or search tool is invoked. When an analyst asks a new but structurally similar question, ART retrieves the closest demonstrations, wires in the calculator and search tools, and runs the frozen model through the interleaved reasoning-and-tool steps, so the team gets consistent tool-grounded answers without authoring a fresh prompt per question, much as the original ART research was evaluated across the BigBench and MMLU task suites.
ToolsHow to implement it
- DSPyprogrammatically bootstraps and selects few-shot demonstrations for a pipeline, the closest modern embodiment of ART's automatic exemplar selection.
- LangChain and LangChain Huba shared library of prompts plus a tool registry the agent draws tools from.
- Microsoft Semantic Kernela plugin library that a planner selects tools from to compose multi-step solutions.
- Guidancestructured generation that can pause a reasoning trace to run a tool and resume, close to ART's execution model.
Cost & effortWhat it takes
The main cost is upfront and recurring library curation: authoring demonstration traces and wiring up tools, then maintaining them as tasks drift. Runtime cost tracks a ReAct-style loop, since retrieved demos and tool definitions enlarge the prompt and each tool call adds a round trip, so expect higher token and latency cost than a single call, offset by not re-engineering a prompt for every new task.