Home › Agent Skills › The description field
Agent Skills · Build

The description field

The one line that is always in context, and therefore the only thing deciding whether your skill ever fires.

In one line

Across Anthropic's own 25 first-party skills the median description is 357 characters, not the 1024 the limit allows — the cap is a guardrail, not a target.

Why you'd careThe problem it solves

You wrote a genuinely good skill. Two thousand words, a worked example, a reference file with the real schemas. You ask for exactly the thing it covers and the model does it from scratch, badly, the way it always did. Nothing errored. Nothing warned you. The skill is installed and it did not fire.

The instinct is to go back and improve the body, which is the one thing that cannot possibly help, because the body was never read. The model made a yes-or-no decision from one line of text and answered no. That line is the description. It is the only part of your skill that is in context at the moment the decision gets made, which means it is not a summary of your skill — it is your skill's entire interface to the model that has to choose it.

ConceptWhat it is

The description is the second of the two required frontmatter keys, and functionally it is the only one that does work at runtime on a local harness. It sits in Level 1 of progressive disclosure: resident from session start, for every installed skill, whether or not that skill is ever used.

That residency creates the constraint. At selection time the model is looking at a list of one-liners and nothing else. It cannot peek at your body to check whether the skill is relevant, because the body is exactly what loading costs. So the description has to answer two questions that most people only answer one of: what does this do, and when should I reach for it. Descriptions that only answer the first are the standard failure mode. “Provides guidance for working with hooks” is accurate, well-formed, and will lose every close call.

Anthropic's house style pushes hard in one direction here. The convention is third person — “This skill should be used when the user asks to…” — followed by literal phrases a user would type. And skill-creator goes further, stating outright that Claude currently tends to undertrigger skills, and instructing authors to make descriptions “a little bit pushy”: name the adjacent situations too, including ones where the user does not use your vocabulary.

The boundary: the description decides whether. It does not decide how. Procedure belongs in the body. Trigger conditions belong here and nowhere else.

How it worksThe mechanics

Two hard limits are enforced by quick_validate.py, and the second one surprises people. The description must be at most 1024 characters. And it may not contain < or > anywhere — the check is a flat substring test, so “handles files <5MB” or any XML-looking fragment is a hard failure.

That second rule catches Anthropic itself. The example-command skill, whose stated purpose is to demonstrate frontmatter options, has a description reading “…demonstrates frontmatter options and the skills/<name>/SKILL.md layout”. Remove its other validation problem and it still fails on the angle brackets.

code
Rejected &mdash; angle brackets in the description:
  description: An example user-invoked skill that demonstrates
    frontmatter options and the skills/<name>/SKILL.md layout

Accepted &mdash; 204 chars, says what and when:
  description: Guidance for distinctive, intentional visual design
    when building new UI or reshaping an existing one. Helps with
    aesthetic direction, typography, and making choices that don't
    read as templated defaults.

On length, here is the measured distribution across all 25 first-party skills as of 2026-07-25: minimum 105 characters (example-command), maximum 906 (project-artifact), mean 388, median 357. This corrects a figure of 403 that appeared in an earlier draft of this entry — the median is 357. Either way the shape of the finding holds and is the useful part: real skills use about a third of the available budget. Nobody is writing to the cap.

If you want this settled empirically rather than by taste, skill-creator ships the machinery. Its improve_description.py scores a candidate description against a labelled set of prompts split into should-trigger and should-not-trigger, then rewrites and rescores. The guidance is specific that the valuable negatives are near-misses: queries sharing vocabulary with your skill that genuinely need something else.

At a glanceSee it

The description field diagram

The description is read before the body exists in context, so it alone decides whether the body ever loads.

Where it runsSurfaces and availability

SurfaceStatusNotes
Claude CodeYesThe description is the trigger surface, but the absolutes need correcting. description is documented as "Recommended", not required, and "If omitted, uses the first paragraph of markdown content" — so the body can feed the decision. The listing truncates combined description and when_to_use at 1,536 characters, and when_to_use exists specifically to carry trigger phrases. Source: code.claude.com/docs/en/skills, "Frontmatter reference".
Agent SDKYesSame selection mechanism: "The description field determines when Claude invokes your Skill", and the troubleshooting section's first remedy for a skill that never fires is "Check the description". Source: code.claude.com/docs/en/agent-sdk/skills.
Claude Desktop / claude.aiYesWas Unverified. The docs describe selection generically for all Skills surfaces — metadata preloaded into the system prompt, "the description is what Claude matches your request against" — and claude.ai is documented as a full Skills surface where Claude "uses it automatically when relevant to your request". Source: agents-and-tools/agent-skills/overview.
Claude API / Messages APIYesCorrected from No. It is a trigger surface: "The description is critical for skill selection: Claude uses it to choose the right Skill from potentially 100+ available Skills." You may attach up to 20 Skills per request and Claude picks among them; naming a skill_id makes a skill available, it does not force its use. Sources: agents-and-tools/agent-skills/best-practices, "Writing effective descriptions"; build-with-claude/skills-guide.
Managed AgentsYesCorrected from No. "Both work the same way: your agent invokes them automatically when they are relevant to the task," with up to 500 skills per session — the description is the gate at that scale. The session-scoped override is confirmed: agent_with_overrides accepts a skills field that replaces the agent's set for one session. Sources: managed-agents/skills; managed-agents/sessions, "Override agent configuration for a session".
Amazon BedrockNoConfirmed. Agent Skills appear under "Features not supported". Source: build-with-claude/claude-in-amazon-bedrock.
Google Vertex AINoConfirmed, same listing. Source: build-with-claude/claude-on-vertex-ai.
Microsoft FoundryYesWas Unverified. Foundry "inherit[s] the same Skills behavior as the Claude API", so selection semantics follow the API: description-driven, on Hosted on Anthropic deployments only. Sources: agents-and-tools/agent-skills/overview; build-with-claude/claude-in-microsoft-foundry.
OpenAI Codex CLI / ChatGPTYesWas Unverified. The analogous concept is confirmed: description "explains when the skill should and should not trigger", and Codex "can choose a skill implicitly when your task matches the skill description". The caution about limits holds and is now specific — OpenAI documents no per-field character cap and no XML/angle-bracket rule, only a whole-list budget of 2% of context or 8,000 characters, where Anthropic caps description at 1,024 characters and forbids XML tags. Source: developers.openai.com/codex/skills.
Google Antigravity CLIYesWas Unverified. The agent "decides based on context" from the name-and-description listing, and users can mention a skill by name to force it. Note the inversion: on Antigravity description is the only required frontmatter field and name is optional. Source: antigravity.google/docs/skills.

The surface split this page was built on is much narrower than claimed. The description is load-bearing on every surface that actually runs Skills, hosted ones included: the API picks among up to eight attached skills by description, and a Managed Agent picks among up to five hundred. What you control on the hosted surfaces is the candidate set, not the selection. So the failure mode is not "author on the API, ship to Claude Code, discover the description matters" — it is subtler: with one skill attached to a request, a bad description rarely costs you anything visible, and the problem only appears when the set grows or the same skill lands in a harness that enumerates everything on disk. Write the description as the trigger it is everywhere, and use the Claude Code-only levers (when_to_use, disable-model-invocation, paths) as refinements rather than substitutes.

ExampleIn the real world

You ship a postmortem skill. Description: “Provides guidance for incident documentation.” For three weeks it fires zero times, while people ask things like “write up what happened with the checkout outage” and “we need the incident review for Tuesday”.

Nothing in that description contains the words those people used. “Incident documentation” is a category label; “write up what happened” is a request. The model was matching against a taxonomy entry, not a situation.

You rewrite it to about 380 characters, in the house third-person form: this skill should be used when the user asks to “write a postmortem”, “do the incident review”, “write up what happened”, or refers to an outage, a sev, a regression that reached customers, or a retro on a production failure — and it names the deliverable, a timeline plus contributing factors plus action items with owners.

Same body. Same reference files. Next Monday someone types “can you write up the checkout thing” and the skill fires. Nothing about the skill's capability changed; what changed is that its advertisement now uses the vocabulary of the person asking, rather than the vocabulary of the person who wrote it.

Not thisWhat it is often confused with

  • Not a human-facing summarynobody browses these. It is a routing input consumed by a model at selection time, and it should be optimised for that reader, not for a docs page.
  • Not the skill's titlethat is name. Restating the name in prose wastes the only line you get; the description exists to add the trigger conditions the name cannot carry.
  • Not a system prompta system prompt shapes behaviour on every turn unconditionally. The description shapes exactly one decision, once, and then stops mattering.
  • Not keyword matchingthe model reads it semantically, so you are not stuffing keywords. But it also cannot infer coverage you never mentioned, which is why naming concrete phrases beats naming an abstract category.
  • Not documentationhow the skill works goes in the body. Anything explanatory in the description is paying always-resident rent for information nobody needs at selection time.

LimitsWhen not to reach for it

  • The skill should only run when explicitly invoked.Some first-party skills set disable-model-invocation: true in frontmatter precisely so the model never auto-selects them. If that is your intent, tuning the description is effort spent on a decision that no longer happens.
  • Your problem is overtriggering.A pushy description that fires on adjacent work is worse than one that never fires, because it costs tokens and produces confidently wrong shapes. Narrow it and name what the skill is not for.
  • You are padding toward 1024 characters.The measured median is 357. Length past the point of specificity adds resident cost with no selection benefit.
  • The real fault is scope.If you cannot describe when to use the skill in one sentence, the skill probably does two things. Split it; two crisp descriptions beat one hedged one.
Checked

Verified 2026-09-12. Moves on a scale of months. Re-check before you depend on it. Provider: Anthropic.

A living map of modern AI — kept current every morning