Few-shot prompting teaches by example, letting the model pattern-match the format you want.
ConceptWhat it is
Few-shot prompting includes a small number of example input-output pairs directly in the prompt before the actual task, letting the model infer the desired pattern, format, and style. It exists because zero-shot instructions alone often fail to convey subtle formatting or stylistic requirements that are easier to show than describe.
The examples act as in-context training data, steering the model without any actual fine-tuning or weight updates.
How it worksThe mechanics
Two to five example pairs are written into the prompt, each showing an input and the ideal output, followed by the real input; the model attends to the examples as a pattern template and generates output consistent with their style, structure, and reasoning depth.
At a glanceSee it
The escalation ladder — add examples only when zero-shot misses, and reach for fine-tuning or per-query retrieval only when a fixed handful still fails.
Dynamic few-shot swaps the fixed example set for the k nearest cases retrieved per query, so each prompt carries its most relevant demonstrations.
When to use itWhere it fits
- Tasks requiring a specific, consistent output format.
- Domain-specific style or tone that is hard to describe in words.
- Classification tasks with subtle category boundaries.
- Situations where fine-tuning is overkill but zero-shot underperforms.
When NOT to use itLimits & anti-patterns
- Very long tasks where examples consume too much context budget.
- Highly variable tasks where no small set of examples generalizes well.
- When examples might bias the model toward surface-level pattern copying instead of real reasoning.
Trade-offsAdvantages & costs
Advantages
- Reliably improves output format consistency.
- No training or fine-tuning infrastructure needed.
- Easy to iterate by swapping example sets.
- Effective for niche or company-specific conventions.
Trade-offs & costs
- Uses more tokens, raising cost and latency.
- Quality depends heavily on example selection.
- Can overfit to example patterns rather than generalizing.
- Requires curating and maintaining good examples.
ExampleIn the real world
A legal-tech startup extracting clause types from contracts includes three labeled clause examples in every prompt so GPT-4 consistently outputs the same JSON schema.
ToolsHow to implement it
- LangChain FewShotPromptTemplatestructured way to manage and inject example sets.
- PromptLayerversion-controls few-shot prompt variants over time.
- DSPyautomatically selects and optimizes few-shot examples.
- Weights and Biases Promptstracks few-shot prompt performance experiments.
Cost & effortWhat it takes
Higher token cost than zero-shot due to embedded examples; latency grows modestly, effort is in curating a small, representative example set.