Home › Build
Build

Agent Skills

What a skill actually is, which surface it runs on, and what it is constantly confused with — every claim dated, because provider surfaces move monthly.

At a glanceAgent Skills

Agent Skills diagram

Only a skill's description stays resident; a description match is what loads the body, then any bundled files and scripts.

CompareEvery entry, side by side

A skill is packaged instruction the model loads when the task matches — not a tool, not an MCP server, not RAG, and the difference decides which one you reach for. Filter by kind, or by PROVIDER, or search across surface, vendor and gotcha. Every row carries the date it was checked and how fast it goes stale, because an undated answer is the one you cannot act on.

LandscapeTypes & approaches

Click a highlighted type to open its own page — concept, use case, and diagram.

Why it mattersThe 40-line preamble you paste into every session

You have a block of text you keep re-pasting: how migrations work in this repo, which of the four PDF paths does not mangle form fields, the house rules for a redline. Forty lines, correct, pasted again because the model has no way to know it — and on the sessions you forget, you pay for it in corrections.

Packaging that block as a skill changes two mechanical things. The cost moves from every request to the ones where it is relevant: only name and description stay resident, and the body loads when the task matches. And the procedure becomes reachable by machinery you are not sitting in front of — a teammate's checkout, a scheduled run, a CI job — because it is a file at a known path rather than a paste.

Then you go to write one and find the word has come loose. “Skill” now means at least four different things, and which one someone means changes what you have to build.

  • A folder on a filesystem.A directory containing SKILL.md that a coding harness discovers at session start. No upload, no ID, no version bump — edit the file and the next session picks it up.
  • A registry object.An org-scoped server-side record created by POST /v1/skills, with an ID and versions. A folder on your laptop is invisible to it; converting one into the other is a deliberate step, not a sync.
  • A hosted capability the vendor wrote.The pre-built document skills — pptx, xlsx, docx, pdf — switched on by naming a skill_id. You author nothing; you wire something.
  • A loose synonym for what the model can do.The sense used in launch posts and sales calls. No file exists anywhere.

The differences show up as errors, not as nuance. Skills do not sync across surfaces: one uploaded to claude.ai is not available via the API, and filesystem skills are separate from both. So the first question about any claim here is not whether it is true. It is where, and when.

AnatomyA folder, one required file, and three loading tiers

Take the narrowest sense first, because the other three are built on it. A skill is a directory whose only required member is a file named SKILL.md. The file name is the fixed part, capitalised exactly like that; the directory name is the part you choose, and it is what people type. The file opens with a --- fenced YAML block; everything after it is Markdown read as procedural guidance.

  • SKILL.mdrequired. Frontmatter carries name and description, required everywhere. name must match ^[a-z0-9-]+$, no leading, trailing or consecutive hyphens, 64 characters max.
  • scripts/executable code for anything needing deterministic reliability or that the model keeps rewriting. Token-efficient because it can run without being read, though not perfectly opaque: the model may still open a script to patch it for your environment.
  • references/material pulled in only when the task reaches for it. A schema, a spec, a long checklist.
  • assets/files copied into your output rather than read as instruction: templates, boilerplate, a logo.

Progressive disclosure loads those in three stages: name plus description always resident, the body when the skill triggers, bundled files only when something reaches for them. A description is capped at 1024 characters, so the resident cost of an installed-but-unused skill has a hard ceiling of about a paragraph. That is the entire reason fifty installed skills do not cost fifty skills' worth of context.

Two rough edges will otherwise cost you an afternoon. How long may the body be? Anthropic's guidance answers twice in one file, and the two answers are not even in the same unit — a checklist item capping it at 5k, a section further down saying keep it under 3,000 words. Nothing on the page reconciles them. Take the smaller and move on. Second, there are two frontmatter schemas: the standalone validator enforces a closed allowlist of name, description, license, allowed-tools, metadata and compatibility, while the Claude Code runtime's is not closed. The check takes a minute — run that validator over an official plugin skill that loads fine in a session and watch it fail on keys it has never heard of. When the two disagree the runtime wins, because the runtime is the thing that loads the file. Even the template shows the seam: its example frontmatter reads name: Skill Name, title case with a space, failing the regex that same toolchain enforces. Two smaller things: version is decoration, and a .skill file is a zip.

Not thisSix ways to put something in front of a model, and what separates them

Most arguments about skills are arguments about which of six mechanisms someone meant. All six end up as tokens in a context window, which is why they blur; they differ on who loads them, when they arrive, and whether they can make anything happen.

MechanismWhat it isWho loads itWhen it enters contextWhat it can do
Skill A directory containing SKILL.md, plus optional scripts, references and assets. The harness or runtime, matching the task against the description. Description always; body on trigger; bundled files only when reached for. Instructs. Executes nothing itself — its scripts run through a tool the surface already has.
Tool A named function with a JSON Schema: name, description, input_schema. You, declared on the request. The full schema on every request, used or not. Model returns tool_use; your code runs it and replies with tool_result. The only one of the six that changes state outside the model.
MCP server A separate process advertising tools, prompts and resources over a defined protocol. The client, at connect time. As ordinary tool definitions, once discovered. Whatever its tools do. It supplies transport and discovery, not a new capability class.
RAG Passages you searched out of a corpus and injected. Your code, per query. At request time, as message or document content. Supplies facts. The model cannot distinguish retrieved text from any other text.
System prompt Always-on instruction at the front of every request. You, on every call. Every turn, relevant or not, after tools and before messages. Shapes behaviour, and competes for attention with everything else you put there.
Subagent A fresh model instance with its own context window, given a scoped task. The harness — not an API primitive. Not in your context at all. Only its summary returns. Runs the task and returns a summary. Buys separate context, not new capability.

Three confusions cause most of the wasted work. Skill versus tool: a skill can tell the model how to file an expense report; only a tool can file it. Knowledge and procedure → skill; a connection to a system → tool. Skill versus RAG: skills load whole files on a trigger, while RAG selects a few chunks out of something far too large to load at all — which is why putting a knowledge base in a skill is the common wrong instinct. Skill versus subagent: builders reach for “a reviewer skill and a researcher skill” when what they want is a clean context window for each; a skill adds instructions to this conversation and cannot give you a second one.

Where they runAvailability is a matrix, not a yes or no

“Does Claude support skills?” has four answers on Anthropic's own surfaces before another vendor is involved. The useful question has three parts — which surface, for which skill, on what date — and it resolves through the same three checks every time.

How does this surface find a skill? Two loaders, and they do not talk to each other. A filesystem loader discovers a directory on disk: no upload, no skill_id, no version bump, edit and re-run. A registry loader needs the object to exist server-side first — create it, cut a version, reference it by ID. A local folder is invisible to the API, and that is not a gap waiting to be closed; converting between them is a step you take on purpose.

Does this surface have a code-execution container? The pre-built document skills run inside one, which is why enabling them means sending the skill reference and the code-execution tool and the beta headers together — omit one and you get a plain text answer with no file. It also explains the hard negatives: where there is no code execution there are no skills, so the path is missing rather than disabled. Amazon Bedrock and Google Vertex AI were in that category as of the check date carried on their rows in the grid below.

Who wrote the skill? Availability is per skill per surface, not per surface: the same harness can ship one first-party skill and explicitly not ship the document ones. So “skills work here” is never a complete sentence.

Underneath all three sits a scope question that catches people once: a skill living only in your personal skills directory does not exist to a cloud session or a scheduled routine — each starts fresh and reports it as not found. And one naming trap turns an availability question into an argument: Amazon Bedrock is not the same product as Claude Platform on AWS. Same cloud, two products; Bedrock is partner-operated, uses anthropic.-prefixed model IDs, and carries a reduced feature set. The model ID is the cheapest tell — read the prefix in whatever code is under discussion before anyone argues about features.

VendorsThe format travelled. The contents did not.

Anthropic defined the shape and other vendors adopted it. OpenAI ships the same SKILL.md folder in Codex CLI and inside ChatGPT's Code Interpreter sandbox. Google's terminal agent reads the same folders and pushed the vendor-neutral .agents/skills/ path, with four discovery tiers from built-in up through extension, user and workspace skills.

The convergence is at the file level and stops there. A SKILL.md written for one harness names that harness's tools and assumes its permission prompts; another vendor's agent will read it happily and follow instructions about tools it does not have. The folder is portable; the body is a document about one runtime. The hosted document skills are not portable at all — a vendor capability with an ID is not a format.

Google's row carries the sharpest warning here, and it is not about skills. The CLI that shipped Google's skills support is no longer available to individuals on free or consumer-paid accounts, and the agent has since shipped under a different name — the grid row carries that name and the date it was checked — while the skills documentation is still live and carries no deprecation banner. Follow that tutorial today and you will carefully configure a tool that does not answer on your account. It is the best argument on this page for dating claims: the doc was not wrong when written, and nothing about it looks stale.

Two conventions sit beside skills rather than competing with them. AGENTS.md is a schema-less Markdown file at the repository root — no frontmatter, read as project context at session start, nearest file winning where they nest. It is stewarded by the Agentic AI Foundation under the Linux Foundation, contributed by OpenAI, and it is the one file most agents will read. MCP is wired through OpenAI's Responses API, Agents SDK and ChatGPT connectors, but assume nothing about coverage: a client may implement tools and none of resources, prompts, sampling or elicitation, and the only reliable check is to connect one and read what it advertises.

Then the mechanism most teams already have checked in and forgot: IDE rules files — Cursor's .mdc rules, whose frontmatter produces four activation modes, and Copilot's instruction files under .github. Grep both before writing a skill; the procedure is often already encoded there and will keep firing alongside whatever you add. One caveat on this family: the exact skills paths and enablement flag for Codex are December 2025 launch-state details that the sweep behind this page did not re-verify. Treat them as unverified.

AuthoringThe description is the only part that is always read

The classic failure is not a bad skill. It is an excellent 2,000-word skill that never fires, followed by the conclusion that skills do not work. The body was never the problem — the model never saw it. Only name and description stay in context, and the model chooses from that listing without opening anything.

Concretely: a description reading Helps with PDFs. names no trigger, so it competes for every PDF-adjacent request and reliably wins none of them. The repair adds the missing half — Fills and flattens AcroForm PDFs without losing field values. Use when the user asks to fill, merge or flatten a PDF form, or reports blank fields after a merge. The second sentence is the one doing the selection work.

  • Write both halves.What the skill does and when to reach for it. The trigger conditions are the half people leave out, and the half that decides whether it fires.
  • The cap is not a target.The limit is 1024 characters. Across the 25 official first-party skills the median description is 403 characters, the longest 906, the shortest 105 — counted from the published skill folders, so you can re-run it.
  • Name it as a phrase someone would say.Kebab-case, 64 characters or fewer. It is what gets invoked and what shows up in a listing.
  • Pick the scope deliberately.Four locations, and the default people reach for is usually wrong: personal (every project on your machine), project (committed to the repo, so the team and CI get it), plugin (wherever it is enabled), organization-wide via managed settings. Anything a teammate or a scheduled run needs cannot live in personal.
  • Push determinism into scripts/.If the model is re-deriving the same snippet every session, that is a script, not a paragraph.
  • Distribute with a pluginonce more than a couple of teams want it; a marketplace is a git repo with a catalog JSON at a fixed path in the root, named on its grid row. Two traps: the marketplace name is a global-per-user key, so a second one with the same name silently replaces the first, and you must reload plugins after installing or the skills are missing from the session you installed from — which reads exactly like a broken install.
  • Test it like a prompt, because it rots like one.Run a handful of realistic prompts in a fresh session with the skill available and again with it disabled, then compare. Fresh matters: leftover context from authoring masks whether the skill did any work. And watching a skill trigger tells you the model found it, not that it helped.

The gridThirty-five entries, and a date on every row

The grid holds 35 entries in six families: Anatomy (what a skill is made of), Where they run (the surfaces), Pre-built (the hosted first-party skills), Authoring & distribution, Neighbours (the six mechanisms above, one row each) and Other vendors. A provider filter ANDs with the family picker, so you can ask for Google rows about where things run and nothing else. Some entries carry more than one provider, because some rows describe a convention nobody owns; the values come from each row's sourced attribution rather than from picking a winner.

The pair of columns worth reading first is Freshness. Checked is the date someone verified the row. Goes stale in is a volatility class — a claim about half-life, not about confidence:

  • stableunlikely to move. Four rows, the conceptual ones.
  • monthsre-check before you depend on it. Twenty-one rows, the bulk of the set.
  • weekstreat anything specific as a starting point, not a fact. Ten rows, clustered where you would expect: API surfaces, beta headers, distribution mechanics.

The set carries two checked-dates, not one: 31 rows from the first pass and four added by a second, and the older 31 were deliberately not re-stamped — moving a date without re-reading the source is the one thing a dated catalogue exists to prevent. So the grid shows its own seams, which is the intended reading. When a date gets old the rows do not become wrong; they become unverified, which is a different and more useful state. The Watch out for cell carries the gotcha or the correction, and several rows record a repair: an earlier draft claimed something, it was checked, it was wrong. Those stay in deliberately.

DisagreementsWhat to do when the grid and the docs say different things

It will happen. These surfaces move on a scale of weeks and this page is a snapshot with a timestamp on it. When a row and a vendor's current documentation conflict, work in this order.

  1. Read the date before the claim.If the row was checked before the vendor shipped the change, the row is old and the vendor wins. Nothing was dishonest; it expired. That is the case most of the time, and it takes ten seconds to establish.
  2. Suspect the documentation too.Vendor docs disagree with themselves more often than people assume, and three examples are already named above: two body-length ceilings in one guidance file, in two different units (Anatomy); a validator that rejects official skills the runtime loads without complaint, and an official template whose example name fails that same validator (Anatomy); and a live tutorial with no deprecation banner for a CLI individuals can no longer obtain (Vendors). “The docs say” is not a single oracle.
  3. Prefer what you can run.Most of these disputes resolve in one request that either returns a 400 or does not. A validation error is a primary source, dated by definition.
  4. Treat anything marked unverified as a question, not an answer.Where the sweep behind this page could not confirm something it says so in place rather than guessing. An unverified line is a lead: it names the page to open and what to look for. It is not permission to assume in either direction.
  5. Then produce a date, not an opinion.If you check something and it has moved, the useful output is the new fact and the day you saw it.

The point of dating a claim is not that a dated claim is permanently right. It is that its wrongness is discoverable. You can ask any model what a skill is and get a fluent, confident, well-organised answer with no date attached — and you will not be able to tell whether it is describing this month's API surface or one from a year ago. Neither will it.

GlossaryKey terms

SKILL.mdThe one required file in a skill directory. That exact name, capitalized — a case-sensitive volume will walk straight past skill.md and the skill simply never appears.
Progressive disclosureOnly the skill's name and description sit in context until the task matches; the body loads on demand. It is why a hundred skills do not cost a hundred preambles.
The description fieldThe one thing that decides whether a skill fires. Written for the model deciding whether to load it, not for a human browsing a list.
SurfaceThe specific place a skill runs — Claude Code, the Messages API, Managed Agents, a reseller cloud. 'Does it support skills' has no answer; 'which surface, on what date' does.
VolatilityHow fast a row goes stale: stable, months, or weeks. Printed beside the date it was checked, because the two together are the claim.

CheckCheck your understanding

When is a skill the wrong answer, and a tool the right one?

A skill is instruction — it changes what the model knows to do. A tool is capability — it lets the model cause an effect outside itself. If your problem is that the model does the job wrong, a skill can fix it. If your problem is that the model cannot reach the system at all, no amount of instruction helps.

Why does this catalogue date every row?

Because a reader can ask any model what a skill is and get a fluent answer. What they cannot get is a claim someone checked, on a date, and said so. Provider surfaces move monthly; an undated answer is the one you cannot act on.

A row says Unverified. Is that a gap?

No — it is the honest state. It means a primary source could not confirm it either way at the date shown. Treat it as a question to settle before you depend on it, not as a soft yes.

What is the difference between a skill and an MCP server?

A skill is content the model reads; an MCP server is a process the harness connects to that exposes tools and resources. One ships as a folder of Markdown, the other as a running endpoint. They compose — a skill can tell the model how to use the tools an MCP server provides.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page Claude Cowork now runs in Anthropic's Chrome extension, extending agentic workflows to in-browser tasks with skills and plugins.

    Anthropic · 12 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

With thanksCaptures learning from Elite AI Assisted Coding on Maven, taught by Eleanor Berger. Thanks & what I learned →
A living map of modern AI — kept current every morning