Watching a skill trigger tells you the description matched, not that the skill worked; the only signal worth anything is the same prompts run with it and without it.
Why you'd careThe problem it solves
You rewrote the skill because the first draft rambled. The new one is tighter, it fires on every prompt you try, and you ship it. A week later the output is quietly worse: the model has started copying one of your examples too literally, or it skips a step it used to do because you cut the sentence that mentioned it. Nothing in the workflow would have caught that. Watching the skill get picked up feels like a test and is not one — it confirms the description matched, which is the easiest part to get right and the part you were not worried about. Skills rot the way prompts rot: quietly, and in the direction of whatever you edited last.
ConceptWhat it is
Two things travel under this name and they are worth keeping apart. One is software: Anthropic publishes a skill-creator, distributed through its plugin marketplace and also readable as a plain directory in Anthropic's public skills repository. The other is the method it encodes — a comparison loop that treats a skill edit as a change to be measured rather than a document to be proofread.
The method has three parts. A fixed prompt set: a handful of realistic requests, written once and then frozen, including at least one request the skill must not fire on. A control: the same prompts run where the skill is unavailable, or where the previous version is installed instead. And a bar: for each prompt, a sentence written in advance saying what a pass looks like — the file exists, the timeline table is present, the deprecated flag was not used. Written in advance matters, because a bar written after reading the output is a description of the output.
What the loop produces is not a score. It is an attribution. Every failing prompt lands in exactly one of three buckets: the skill never loaded, it loaded and the model ignored it, or it loaded and was followed and the result is still wrong. Each bucket points at a different part of the file — the description, the instructions, the content. Without a control run you cannot tell them apart, and you will spend an afternoon rewriting the description when the problem was in the body.
How it worksThe mechanics
Install first; the old surprise is mostly gone. If the install summary says Run /reload-plugins to activate., Claude Code now runs that reload for you (v2.1.268+). You act only if the reload warns that your next message would re-read the conversation — then /reload-plugins --force makes the plugin's skills available in the current session. Starting a fresh session has the same effect.
Then the loop, which needs no tooling at all:
- Write the prompt set into a folder next to the skill. Nothing reads these Markdown files — they are a convention for you, not a product feature (skill-creator keeps its own test cases in
evals/evals.jsoninside the skill directory). - Run each prompt in a fresh session. This is not fussiness. The session you authored the skill in already contains the skill's text, your explanations of it, and the file you were editing, all of which mask a description that would never have matched on its own.
- Run the control. Same prompts, skill absent or previous version installed.
- Diff, then attribute.
skills/incident-review/
├── SKILL.md
└── evals/
├── 01-normal-case.md prompt plus what a pass looks like
├── 02-no-timeline-yet.md
└── 03-must-not-fire.md a request this skill should stay out ofFor the control run you need the skill genuinely unavailable. Claude Code's settings expose a per-skill override for exactly this — skillOverrides — though it does not reach plugin skills, which you disable through /plugin instead. If that key is not in your version's settings reference, rename the skill directory for the duration of the control run: crude, slower, and impossible to get wrong.
Attribution rules, in order. Skill did not load → the description is the problem; rewrite it to name the triggering situation in the user's words, not yours. Loaded but ignored → the body is the problem; it is probably too long, too hedged, or buries the instruction under context. Loaded, followed, still wrong → the content is the problem, and no amount of prompt-craft fixes a rule that is wrong.
One note about the tooling: what skill-creator ships has changed over its short life, and today's build runs the comparison for you. It spawns a subagent per test case so each run starts clean, grades each assertion pass or fail into grading.json, aggregates with-skill against without-skill pass rate, time and tokens into benchmark.json, and can run a blind A/B between two versions of the skill. For a skill that ships in a plugin, claude plugin eval runs a similar with-and-without loop from the command line. Read its own SKILL.md after installing to see what your build does.
At a glanceSee it
Same prompts run with and without the skill, then each failure attributed to one part of the file.
Where it runsSurfaces and availability
| Surface | Status | Notes |
|---|---|---|
| Claude Code | Yes | Where the plugin installs (/plugin install skill-creator@claude-plugins-official; Claude Code runs the reload itself) and where the loop is practical: fresh sessions are cheap, and skillOverrides plus the /skills menu disable a non-plugin skill without deleting it (plugin skills go through /plugin), which is what a control run needs. Both confirmed at code.claude.com/docs/en/skills. |
| Claude API / Messages API | No | This is an authoring-time activity. The skill it produces can be uploaded to /v1/skills and used there; the refinement loop itself does not run there. |
| Managed Agents | No | A consumption surface, not an authoring one. The session-scoped override is real and useful for A/B once a skill is deployed: agent_with_overrides takes a skills field that replaces the agent's skills for one session, or clears them with null/[] for the control arm (platform.claude.com/docs/en/managed-agents/sessions). |
| Claude Desktop / claude.ai | Unverified | No confirmed install path for this plugin. The control run itself is not the blocker, though: claude.ai skills are individually enabled and disabled from the skills settings on claude.ai or Customize in the Desktop sidebar, and Cowork sessions load whichever are enabled at session start (code.claude.com/docs/en/skills). |
| Agent SDK | Yes | The method transfers cleanly and is easy to script, and the plugin loads: pass its directory to the plugins option (CLI-installed copies live under ~/.claude/plugins/). This is independent of settings sources — plugins have their own option, and with default options the SDK loads the user and project sources in any case. Sources: code.claude.com/docs/en/agent-sdk/plugins and .../agent-sdk/skills. |
| Amazon Bedrock | Yes | The loop is client-side and works with Claude Code pointed at Bedrock: code.claude.com/docs/en/feature-availability lists skills, plugins and subagents among the features available on every provider. Nothing skill-shaped exists on the Bedrock API itself to evaluate against. |
| Google Vertex AI | Yes | Same confirmation and same gap as Bedrock, from the same feature-availability page, which lists it as "Google Cloud's Agent Platform". |
| Microsoft Foundry | Yes | Also covered by the every-provider list on code.claude.com/docs/en/feature-availability. Foundry does have server-side skills via the Skills API, so a deployed skill there can be re-run against the same prompt set. |
| OpenAI Codex CLI / Google Antigravity CLI | No | Anthropic's plugin does not install there. The method transfers, and Codex has gone further: it ships its own bundled skill-creator as a system skill available to every user (developers.openai.com/codex/skills). Both tools use SKILL.md under .agents/skills and carry the same trigger-versus-outcome ambiguity. |
The one-row answer is that evaluation lives where authoring lives, and authoring lives in the harness — the terminal, or the Agent SDK if you want the comparison loop scripted. If your skills are destined for a server-side agent, do the comparison loop locally and then re-run the same frozen prompt set once against the deployed agent, because the harness around the model differs and a skill tuned against a tool-rich local session can behave differently where those tools are absent.
ExampleIn the real world
An SRE writes incident-review: given a set of alerts and a Slack export, produce a postmortem. The first version works well enough. She then adds a line telling Claude to always include a timeline table, because one review came back without one.
The skill now triggers on everything, so it looks healthy. Her prompt set says otherwise. Six frozen prompts, each run fresh with the skill on and again with it off. Prompt 2, a one-line incident that resolved in four minutes, comes back with a three-row timeline table padded out with invented intermediate steps — the control run, with no skill, produced a clean paragraph. Prompt 5 is worse: it is a request to summarise a design doc, which the skill should never have touched, and the new phrasing pulled it in and wrapped a design doc in postmortem structure.
Two edits follow, in different places. The timeline instruction gets a condition attached rather than being deleted. The description gets narrowed so it names incidents and alerts explicitly instead of the vaguer “analysing what happened”. She reruns the same six prompts, prompt 5 stops firing, prompt 2 goes back to a paragraph, and the four that were already good stay good — which is the part that would have been invisible without the control.
Not thisWhat it is often confused with
- Not a statistical evalthe method is a structured way of looking at a handful of transcripts. The plugin does grade assertions pass or fail and report a with-versus-without pass rate, but over the prompts you wrote, not a dataset with confidence behind it. If you need statistical confidence over hundreds of cases, build a real eval; this is the thing you do instead of nothing.
- Not a linterit will not tell you a frontmatter key is misspelled or a referenced file is missing. Those failures show up as the skill silently not loading, which is exactly why the control run matters.
- Not a trigger guaranteepassing the loop today does not make the skill fire tomorrow on a phrasing you never tested. The prompt set is a floor, not a proof.
- Not a test suite for your scriptsif the skill bundles executable code, that code needs ordinary unit tests. Judging a Python helper by reading an agent transcript is the slowest possible debugger.
- Not a substitute for reading the transcriptthe useful information is usually in what the model did between loading the skill and producing output, not in the final answer alone.
LimitsWhen not to reach for it
- The skill is ten lines and has one user.Use it for a week. Personal friction is a faster signal than a formal comparison at that size.
- You need the behaviour every time.No amount of refinement makes model-invoked instructions deterministic. If the step is mandatory, move it into a hook, a script, or a CI check and let the skill handle the judgement-shaped part.
- The failure is inside bundled code.Write a unit test against the script. Rerunning the whole agent loop to discover that a date parser is wrong wastes an hour per iteration.
- You are actually comparing models.If the same skill behaves differently across model versions, that is a model question. Hold the skill fixed and vary one thing, or you will learn nothing from either.
- The prompt set was written after reading the output.Stop and rewrite it from the user's request, not from what you got. A bar reverse-engineered from a transcript will confirm every future change.
Verified 2026-09-12. Moves in weeks. Treat anything specific here as a starting point, not a fact. Provider: Anthropic.