Home › Prompt Engineering › Skeleton-of-Thought
✍️ · Ground

Skeleton-of-Thought

Skeleton-of-Thought: outline the answer first, then expand every bullet in parallel for faster structured output.

In one line

Have the model sketch a point-by-point skeleton first, then fill each point independently so long answers assemble faster and stay organized.

ConceptWhat it is

Skeleton-of-Thought (SoT) is a prompting pattern that splits answer generation into two stages: first the model produces a lightweight skeleton — a numbered list of concise point headers that map out the answer — and then it expands each header into full prose. Introduced in research on parallel decoding for LLMs, its motivating insight is that much of the wall-clock cost of a long answer comes from generating tokens strictly left-to-right, even when the sections are logically independent.

Because the expansions do not depend on one another, they can run concurrently — as separate API calls or batched together — collapsing what would be one long sequential generation into a short outline pass plus a fan-out. The payoff is twofold: lower end-to-end latency on long outputs, and a cleaner, more consistent structure because the outline is committed before any section is written.

How it worksThe mechanics

A first prompt instructs the model to answer only with a short skeleton — typically a handful of terse point headers, one per line, no elaboration. The application parses that list, then issues one expansion request per point, each carrying the original question, the full skeleton for context, and the single point to flesh out. These requests are dispatched in parallel and their results stitched back together in skeleton order, often with light post-processing to normalize headings and remove duplicated preambles before the assembled answer is returned.

At a glanceSee it

Skeleton-of-Thought diagram
Skeleton-of-Thought diagram 1

Why SoT saves wall-clock time — naive decoding cost scales with total token count, while batching the independent point-expansions makes latency track only the longest single point.

Skeleton-of-Thought diagram 2

Keeping a skeleton answer coherent — scope the headers so they do not overlap, feed every expansion the full skeleton, and add a light stitch pass so transitions do not fray.

When to use itWhere it fits

  • The expected output is long and naturally decomposes into independent sections — listicles, comparisons, checklists, multi-part explanations.
  • Latency on long answers matters and you can afford several concurrent calls to shrink wall-clock time.
  • You want the model to commit to an outline before writing, improving coverage and reducing rambling.
  • Sections are loosely coupled, so each can be written with only the skeleton as shared context.

When NOT to use itLimits & anti-patterns

  • The reasoning is tightly sequential — later steps depend on the results of earlier ones, as in math derivations or multi-hop logic.
  • The answer is short, where the outline pass adds overhead without saving time.
  • Sections overlap heavily, so independent expansion produces redundancy or contradictions.
  • Strict narrative flow or a single continuous argument is required, which parallel fragments tend to fracture.

Trade-offsAdvantages & costs

Advantages
  • Cuts end-to-end latency on long outputs by overlapping section generation instead of decoding one long stream.
  • Produces well-organized, well-covered answers because structure is decided up front.
  • Modular — an individual weak section can be regenerated without redoing the whole answer.
  • Works on top of any chat model through ordinary prompting; no fine-tuning needed.
Trade-offs & costs
  • Coordination overhead — you must parse the skeleton, fan out calls, and reassemble reliably.
  • Parallel expansions share no context beyond the outline, so transitions and cross-references can read disjointed.
  • Total token spend rises because each expansion re-sends the question and skeleton.
  • Actively harmful for problems that require step-by-step dependent reasoning.

ExampleIn the real world

A support team wants an LLM to generate a "Top 10 troubleshooting steps" article from a product manual. A skeleton pass returns ten one-line step titles in roughly the time of a single short generation. The application then fires ten concurrent expansion calls, each given the question, all ten titles, and one title to write up in two or three sentences. The full article assembles in roughly the time of the single slowest expansion rather than the sum of all ten, and every step is present and consistently formatted because the outline fixed scope before any prose was written. A follow-up pass trims a repeated "To do this" opener that several sections independently generated.

ToolsHow to implement it

  • LangChainRunnableParallel / LCEL and map-reduce chains to fan out and gather section expansions.
  • LlamaIndexquery pipelines and async execution for splitting then recombining sub-responses.
  • vLLMhigh-throughput batched decoding so parallel expansions share GPU efficiently when self-hosting.
  • Provider async SDKs with asyncio.gather — the minimal way to dispatch concurrent expansion calls.

Cost & effortWhat it takes

SoT trades money and orchestration complexity for speed. Latency on long answers drops toward the cost of a single section plus the outline, but total token usage climbs — every expansion re-sends the question and skeleton, so you pay for repeated context across N calls. Concurrent requests also press against rate limits and concurrency caps, and the glue code for parsing, dispatch, and reassembly is real engineering. It pays off when responses are long, decomposable, and latency-sensitive; for short or tightly sequential answers the overhead is not worth it.

A living map of modern AI — kept current every morning