Home › Agent Skills › scripts / references / assets
Agent Skills · Build

scripts / references / assets

Three conventional subdirectories, distinguished by whether their contents get executed, read into context, or copied into your output.

In one line

The three directories differ only by destination — scripts get executed, references get read into context, assets get copied into your output — and nothing enforces the names.

Why you'd careThe problem it solves

You watch a session and notice the same waste twice in a week. The model rewrites the same twenty-line PDF rotation snippet it wrote on Tuesday, gets an argument order slightly wrong, and fixes it after one failed run. Later it re-derives your warehouse table relationships from three sample queries, correctly, spending two thousand tokens on something that is written down in a document nobody gave it.

Both are the same failure: knowledge that exists in stable form is being regenerated at inference time. The three conventional subdirectories are the fix, and choosing between them is not a filing decision. It determines whether a thing costs you an execution, a chunk of context window, or nothing at all. Put the schema in scripts/ and it never gets read; put the rotation code in references/ and you pay for it in tokens every time.

ConceptWhat it is

All three are optional subdirectories of a skill folder, and the distinction between them is what happens to their contents at runtime.

scripts/ holds executable code — Python, Bash, whatever runs. Its purpose is deterministic reliability and avoiding repeated regeneration. The token argument is that a script can be executed without being loaded into context: only its output enters the window. The official guidance attaches an important qualification, that scripts may still need to be read for patching or environment-specific adjustment, so the saving is real but conditional rather than guaranteed.

references/ holds documentation intended to be loaded into context when needed — schemas, API docs, policies, detailed procedures. This is material the model must actually read to reason with. It costs tokens on every open. The point is not that it is free but that it is conditional: it is out of the body, so you pay only on the invocations that need it.

assets/ holds files that are never read at all, only used in what the model produces — a .pptx template, a logo, a font, HTML boilerplate. The model copies or fills them.

The boundary worth holding on to: references/ is the only one of the three whose contents enter the context window as text. That single fact should drive every placement decision you make.

How it worksThe mechanics

The layout, with the fourth and fifth directories you will meet in real repositories:

code
skills/my-skill/
├── SKILL.md
├── scripts/      executed    output enters context, the file need not
├── references/   read        spends context tokens each time opened
├── assets/       copied      never read, lands in your output
├── examples/     read        created by plugin-dev's own scaffold
└── evals/        neither     stripped out at packaging time

Two things about that list. First, examples/ is not one of the canonical three, yet plugin-dev's own setup command is mkdir -p plugin-name/skills/skill-name/{references,examples,scripts} — creating examples/ and never creating assets/. Second, evals/ is a skill-creator convention, and package_skill.py excludes it specifically at the skill root when building a .skill archive, alongside __pycache__, node_modules, *.pyc and .DS_Store. Your test harness stays in the repo and does not ship.

Now the part that surprises people: none of these directory names are enforced anywhere. The loader requires SKILL.md and nothing else. Bundled files are discovered because your SKILL.md body names their paths in prose. Anthropic's guidance lists an unreferenced resource as an explicit mistake, with the reason stated plainly — Claude does not know the file exists. Rename references/ to docs/ and everything still works, provided the body points at docs/. You lose the shared vocabulary, not the mechanism.

On sizing, the two first-party guides again disagree in units. plugin-dev says reference files can be 2,000–5,000+ words and that anything over 10k words should ship grep patterns in SKILL.md so the model searches rather than reads. skill-creator says anything over 300 lines should carry a table of contents. Both are aiming at the same thing: give the model a way to open part of a large file instead of all of it.

At a glanceSee it

scripts / references / assets diagram

The same skill folder, three destinations, and only one of them spends context tokens.

Where it runsSurfaces and availability

SurfaceStatusNotes
Claude CodeYesAll three work: scripts execute through Bash, reference files are read with the file tools, assets are copied. ${CLAUDE_SKILL_DIR} resolves to the skill's own directory in both the body and allowed-tools, so a bundled script can run without a permission prompt. Note the directory names are a convention, not a contract — Anthropic's own examples use scripts/ plus reference/ (singular) and top-level markdown. Sources: code.claude.com/docs/en/skills; agent-skills/best-practices.
Agent SDKYesSame harness, same filesystem and execution access. One difference that bites bundled scripts: the skill's allowed-tools frontmatter is ignored in the SDK, so pre-approval must come from the query's allowedTools instead. Source: code.claude.com/docs/en/agent-sdk/skills.
Claude API / Messages APIYesConfirmed — this is what the code-execution container exists for, and scripts run through bash with only their output entering context. Two constraints to design around: the container has no network access and no runtime package installation, so a bundled script may only use pre-installed packages. Sources: agents-and-tools/agent-skills/overview, "Runtime environment constraints"; agent-skills/best-practices, "Package dependencies".
Managed AgentsYesWas Unverified. Skills mount into the session's sandbox — "Mounting more skills increases the time it takes for the session's sandbox to start" — and the sandbox is provisioned when the session first needs it. Skills require the read tool, so clearing an agent's tools while skills are attached returns a 400. Execution of a bundled script still depends on the agent's tool set including a code or bash tool, so verify that before depending on one. Sources: managed-agents/skills; managed-agents/sessions.
Claude Desktop / claude.aiYesWas Unverified. Custom skills require code execution to be enabled on the account, and the runtime is more permissive than the API's: on claude.ai skills "can install packages from npm and PyPI and pull from GitHub repositories". Network access varies by user and admin settings, from full to none. Sources: agent-skills/best-practices, "Package dependencies"; agent-skills/overview, "Runtime environment constraints".
Amazon BedrockNoConfirmed twice over: both "Server-side tools (code execution, web search, web fetch, advisor)" and "Agent infrastructure (Agent Skills, ...)" are listed as not supported. Scripts have no runtime and the bundle has no consumer. Source: build-with-claude/claude-in-amazon-bedrock.
Google Vertex AINoConfirmed, same two listings. Source: build-with-claude/claude-on-vertex-ai.
Microsoft FoundryYesWas Unverified. Foundry "inherit[s] the same Skills behavior as the Claude API", including the code-execution container, on Hosted on Anthropic deployments. On Azure-hosted deployments, server-side code execution and Agent Skills both return 400, so bundled scripts have no runtime there. Sources: agents-and-tools/agent-skills/overview; build-with-claude/claude-in-microsoft-foundry.
OpenAI Codex CLI / ChatGPTYesWas Unverified, and the assumption to correct runs the other way: scripts/, references/ and assets/ are exactly the documented Codex layout — "scripts/ (optional: executable code), references/ (optional: documentation), assets/ (optional: templates, resources)" — alongside SKILL.md and an optional agents/openai.yaml. Source: developers.openai.com/codex/skills.
Google Antigravity CLIYesWas Unverified. Bundled resources are supported but the names differ: scripts/ for helper scripts, examples/ for reference implementations, resources/ for templates and assets. Source: antigravity.google/docs/skills.

The division that matters when choosing still holds: references/ and assets/ need only a filesystem, while scripts/ needs an execution environment, so a skill built entirely from prose and reference files is the most portable thing you can write. Two corrections to how you act on that. First, the three directory names are OpenAI's documented convention, not Anthropic's — Anthropic's examples use scripts/ and reference/, Antigravity uses scripts/, examples/ and resources/, and on Claude surfaces the names carry no special meaning at all because the model simply reads paths you point it at. Second, "can run a script" is not one capability but several: the Claude API container has no network and cannot install packages, claude.ai can install from npm and PyPI, and Claude Code has whatever the user's machine has. A script that shells out to pip install is portable in the same nominal sense that a bridge is portable. Bedrock and Vertex remain the only surfaces where the question does not arise.

ExampleIn the real world

You maintain a quarterly-deck skill. Three things kept going wrong: the model regenerated a fragile python-pptx snippet for the chart slide each quarter, it guessed at revenue table column names, and the output never matched brand typography.

You split them by destination. scripts/build_chart_slide.py takes a CSV path and a title and emits a slide — the model runs it and sees only a success line, never the 180 lines of pptx wrangling. references/metrics.md holds the warehouse column names and the definition of net revenue the finance team insists on; roughly 1,200 words, opened only when the deck needs numbers. assets/brand-template.pptx carries the master slides and fonts; the model opens it as a starting file and never reads a byte of it as text.

SKILL.md is under 500 words and names all three explicitly: run the script for chart slides, read references/metrics.md before writing any revenue figure, start from assets/brand-template.pptx rather than a blank deck.

A run that produces a five-slide deck now spends context on the body plus 1,200 words of metrics. The chart code and the template — by far the larger files — cost nothing at all.

Not thisWhat it is often confused with

  • Not toolsa script in scripts/ has no schema and the model cannot call it as a function. It runs it through Bash because your prose said to. A tool is registered with the API and invoked structurally; a script is just a file with a path.
  • Not a RAG corpusreferences/ is not indexed, chunked or embedded. Files are opened because SKILL.md names them. Ten thousand reference documents would not work; ten would.
  • Not an MCP serverbundled files ship inside your skill and are static. An MCP server is a separate process exposing live capability, and it stays current without you editing anything.
  • Not enforced directory namesnothing in the loader knows what references/ means. The convention buys you readability among humans and consistency with first-party skills, not behaviour.
  • Not automatically discovereda bundled file that SKILL.md never mentions will sit there unread forever. The body is the index.

LimitsWhen not to reach for it

  • The content is short.A 200-word reference file costs an extra tool call to open and saves almost nothing. Put it in the body.
  • The script would need reading anyway.If the model must open a script to adapt its arguments every time, you have paid the token cost and added an execution step. Inline the logic as instructions instead.
  • The data changes between runs.A reference file is a snapshot. If the schema shifts weekly, you have built a document that will be confidently wrong; query the live source through a tool.
  • You are targeting a surface with no execution.Bundled scripts on Bedrock or Vertex are dead weight. Write the procedure out and let the model do it.
  • You are splitting for neatness.Three near-empty directories signal organisation without providing any. Add a directory when a specific file has a specific destination.
Checked

Verified 2026-09-12. Moves on a scale of months. Re-check before you depend on it. Provider: Anthropic.

A living map of modern AI — kept current every morning