Home › Agent Skills › Progressive disclosure
Agent Skills · Build

Progressive disclosure

Three loading tiers: metadata always resident, body on trigger, bundled files only when reached for.

In one line

Three loading tiers, and the published ceiling for the middle one contradicts itself across Anthropic's two authoring guides — 5,000 words, 3,000 words, or 500 lines.

Why you'd careThe problem it solves

You install a plugin, then another, then a marketplace bundle, and now there are thirty-odd skills on your machine. The obvious worry is that you are paying context rent on all of them before you type a word. You are not, and the reason is worth understanding precisely, because it also tells you where the real budget goes. What is always resident is roughly a hundred words per skill — a name and a description. Thirty skills is a few thousand tokens of menu, not thirty bodies. The cost you actually control is what happens after one fires: the entire body lands in context, in full, every time. That is the number worth managing, and it is the number Anthropic's own guidance gives three different answers for.

ConceptWhat it is

Progressive disclosure is the loading policy that makes a large skill library affordable. It has three tiers.

Level 1 is the frontmatter name and description. These are resident for every installed skill from the moment the session starts, whether or not the skill is ever used. Roughly a hundred words each. This is the menu the model chooses from.

Level 2 is the SKILL.md body. It loads in full when the skill triggers — not partially, not summarized, not the relevant section. Every word you put in the body is a word you spend on every invocation.

Level 3 is bundled files: references/, scripts/, assets/. These load only when the model reaches for one by name, and the guidance calls this tier “unlimited” with an asterisk. The asterisk is that scripts can be executed without being read, so their line count never touches the context window — only their output does. A reference file, by contrast, is unlimited in the sense that you are not forced to open it, not in the sense that opening it is free.

The boundary that carries the design is between Level 2 and Level 3. Level 2 is what the model always pays for; Level 3 is what it pays for only on demand. Deciding what sits on each side is the actual craft of skill authoring, and it is a routing decision, not a tidiness one.

How it worksThe mechanics

The ladder, with the numbers as published:

code
Level 1  name + description     always resident     ~100 words per skill
Level 2  SKILL.md body          loaded on trigger   ceiling disputed
Level 3  scripts / references   on demand           "unlimited"

The Level 2 ceiling is where the guidance stops agreeing with itself. Checked on 2026-07-25 against the first-party plugin marketplace, one file — plugin-dev/skills/skill-development/SKILL.md — gives three different figures in three places: “<5k words” in its anatomy diagram, “Keep under 3,000 words, ideally 1,500-2,000” in its progressive-disclosure section, and “1,500-2,000 words ideal, <5k max” in its validation checklist. Meanwhile skill-creator, the other first-party authoring guide, does not count words at all. It says keep the body under 500 lines, then adds that the counts are approximate and you should feel free to go longer.

The practice matches the ambiguity. Measuring the bodies of all 25 first-party skills in that marketplace: only five sit inside the 1,500-2,000 band. Four exceed 3,000 words. Five exceed 500 lines. The guide that tells you to stay under 3,000 words is itself 3,143 words and 632 lines, breaching both of its own thresholds. The skill-creator body is 5,151 words.

Read that as signal rather than as sloppiness. There is no enforced limit anywhere in the toolchain — no validator checks body length, nothing truncates, nothing warns. The advice is advice. The useful reformulation is not a word count at all: the body should contain what is needed on every invocation of this skill, and nothing that is needed on only some of them. If a section applies to one branch in four, it belongs in references/ with a pointer.

At a glanceSee it

Progressive disclosure diagram

The three loading tiers and the two places a skill can cost you nothing at all.

Where it runsSurfaces and availability

SurfaceStatusNotes
Claude CodeYesAll three tiers observable, with two documented exceptions to "every skill appears". A skill with disable-model-invocation: true has its "Description not in context" at all, and paths globs restrict automatic loading to matching files. The listing itself is capped: combined description and when_to_use is truncated at 1,536 characters. Once loaded, SKILL.md stays in the conversation across later turns and is not re-read; after compaction it is re-attached only within a token budget. Source: code.claude.com/docs/en/skills, "Frontmatter reference" and the invocation-control table.
Agent SDKYesSame harness and loader: the SDK "discovers skill metadata at startup from user and project directories, and loads the full content when Claude invokes the skill". The skills option filters what reaches Level 1, but it is not a sandbox: unlisted skills' files "remain on disk and stay reachable through Read and Bash". Source: code.claude.com/docs/en/agent-sdk/skills.
Claude API / Messages APIYesCorrected: Level 1 does apply. The overview's tier table gives "Level 1: Metadata — Always (at startup) — ~100 tokens per Skill", and states the description "is what Claude matches your request against when determining whether to trigger the Skill". You can attach up to 20 Skills per request, so descriptions compete here too — the earlier claim that a weak description costs nothing on the API is wrong. Sources: agents-and-tools/agent-skills/overview, "How Skills work"; build-with-claude/skills-guide.
Managed AgentsYesCorrected for the same reason. "Each skill you add incurs a modest cost on the session's context window, adding instructions and metadata that help the model use the skill," and "your agent invokes them automatically when they are relevant to the task." A session supports up to 500 skills, and mounting more slows sandbox start — so the resident menu is real and can be large. Source: managed-agents/skills.
Claude Desktop / claude.aiYesWas Unverified. claude.ai is documented as a full Skills surface, and the three-level loading model is stated for Skills generally rather than per-surface, so the tiers apply. The published figures (~100 tokens of metadata per Skill, under 5k for a loaded SKILL.md) are generic and not broken out for claude.ai; the surface-specific variable is network access, which "may be full, partial, or no network access" depending on settings. Source: agents-and-tools/agent-skills/overview, "How Skills work" and "Runtime environment constraints".
Amazon BedrockNoConfirmed. Agent Skills are listed under "Features not supported", so there is no loading policy to speak of. Source: build-with-claude/claude-in-amazon-bedrock.
Google Vertex AINoConfirmed, identical listing under "Features not supported". Source: build-with-claude/claude-on-vertex-ai.
Microsoft FoundryYesWas Unverified; now confirmed by construction. The overview states that Microsoft Foundry "inherit[s] the same Skills behavior as the Claude API in all following sections", which includes the progressive-disclosure model. Applies to Hosted on Anthropic deployments only. Sources: agents-and-tools/agent-skills/overview; build-with-claude/claude-in-microsoft-foundry.
OpenAI Codex CLI / ChatGPTYesWas Unverified; a three-tier equivalent is documented. "ChatGPT and Codex start with each skill's name and description, then load the full SKILL.md instructions when they decide to use that skill," with bundled scripts/, references/ and assets/ read further down. The resident listing is explicitly budgeted: at most 2% of the model's context window, or 8,000 characters when that window is unknown. Source: developers.openai.com/codex/skills.
Google Antigravity CLIYesWas Unverified; the same pattern is documented. "When a conversation starts, the agent sees a list of available skills with their names and descriptions. If a skill looks relevant to your task, the agent reads the full SKILL.md content," then follows it. Source: antigravity.google/docs/skills.

The split this page originally drew — Level 1 as a harness-only behaviour — does not survive checking. Metadata for every attached skill is resident on the hosted surfaces too: the API preloads roughly 100 tokens per Skill and lets you attach twenty, and a Managed Agents session can carry up to 500, each one paying context rent and each one competing to be selected. What actually differs is who assembles the menu, not whether a menu exists. Locally the harness enumerates whatever is on disk; on the API and on an Agent you choose the shortlist. That makes descriptions cheaper to get wrong on the API only in the sense that the shortlist is shorter — with twenty skills attached, or five hundred on an Agent, a vague description still loses. The genuine harness-only levers are the Claude Code ones: disable-model-invocation, which removes a skill from Level 1 entirely, and paths, which makes residency conditional on what you are editing.

ExampleIn the real world

You are writing a bigquery skill for an analytics team. The first draft is 4,200 words: a workflow section, then the schemas for eleven tables, then a troubleshooting appendix about quota errors. It works, and every single invocation — including “how many signups yesterday”, which touches one table — pays for all eleven schemas and the quota appendix.

You re-cut it along the every-time boundary. The body keeps the workflow, the join conventions the team always gets wrong, and two paragraphs of pointers: schemas live in references/schema.md, quota handling in references/quotas.md. That body is about 900 words.

Now “how many signups yesterday” loads 900 words, the model reads that it should open references/schema.md for table shapes, greps it for the signups table, and writes the query. The quota appendix is never touched. A harder question that hits a quota error loads the same 900 words and then opens quotas.md because the body told it when to.

Nothing got deleted. The same total content is installed. What changed is that the rare material stopped being charged to the common case.

Not thisWhat it is often confused with

  • Not RAGthere is no index, no embedding, no similarity search. Level 3 files are opened because the SKILL.md body names a path and says when to read it. It is a hyperlink the model follows, not a retrieval system.
  • Not context compactioncompaction summarizes and discards material already in the window under pressure. Progressive disclosure decides what never enters it. One is lossy cleanup after the fact; the other is admission control.
  • Not lazy loading in the software sensenothing is memoized and nothing is resolved by a dependency graph. The model reads a file because your prose told it to, which means a reference file you forget to mention is a file that will never be opened.
  • Not an enforced budgetno validator checks body length and nothing truncates an oversized SKILL.md. Every ceiling you have read is guidance, and the first-party skills routinely exceed it.

LimitsWhen not to reach for it

  • The skill has one short body and no bundled files.Splitting a 400-word skill across three files buys nothing and costs the model an extra read. Leave it as one file.
  • Every branch needs the same material.If the reference file gets opened on all paths, moving it out of the body added a round trip and saved zero tokens. Inline it.
  • You are using tiering to avoid scoping.A skill so broad it needs eight reference files is usually two or three skills. Splitting by concern beats splitting by file.
  • The content is genuinely dynamic.Deferring a load does not make a stale file fresh. If the material changes daily, the fix is a tool or a query, not a lower tier.
  • You only ever invoke the skill explicitly.If a human always names it, the Level 1 economics do not apply to you and the description length debate is moot.
Checked

Verified 2026-09-12. Moves on a scale of months. Re-check before you depend on it. Provider: Anthropic.

A living map of modern AI — kept current every morning