Try β€œClaude Code skills”, β€œMCP servers for Cursor” or β€œCodex” Β· Esc to close

Progressive Disclosure: How Agent Skills Stay Out of Your Way

By Kelvin Β· 29 September 2026 Β· Updated 29 Sep 2026 Β· 8 min read

agent-skills skill-md context-window claude-code mcp

Progressive Disclosure: How Agent Skills Stay Out of Your Way
Photo by Daniil Komov on Pexels

Progressive Disclosure: How Agent Skills Stay Out of Your Way

Install ten Agent Skills into a coding agent and something you'd expect to happen doesn't: the agent's context window doesn't fill up with ten skills' worth of instructions. That's not an accident, and it's not a limit you're about to hit later. It's the core design decision behind the Agent Skills format, and it's called progressive disclosure β€” the idea that an agent should only pay the token cost for a capability at the moment it actually needs it, not the moment it becomes available.

If you've written a skill and wondered why your SKILL.md "isn't triggering," or you've watched a skills-heavy agent stay surprisingly fast despite a folder full of .md files, this is the mechanism underneath both. Understanding it changes how you write skills, how many you can reasonably install, and how you debug the ones that misbehave.

The problem progressive disclosure solves

Every token an agent's system prompt carries is a token it can't spend on your actual conversation, your codebase, or its own reasoning. Coding agents already compete for that budget with the file you're editing, your AGENTS.md or CLAUDE.md instructions, recent shell output, and β€” critically β€” the full tool schemas for every MCP server you've connected. None of that is free.

Agent Skills exist to add capability without adding to that fixed cost. A skill can be as large as it needs to be β€” a full TDD methodology, a 40-page document-formatting reference, a validator script and three example files β€” because almost none of that content sits in the agent's context by default. Only a sliver of it does, all the time. The rest loads conditionally, and most of it may never load in a given session at all.

This is what separates a skills folder from, say, pasting ten reference documents into your system prompt. The documents approach charges you for content whether or not it's relevant to the current task. Progressive disclosure charges you only when the content earns its place.

The three-stage loading model

The Agent Skills specification β€” the open format Anthropic released and that a growing list of clients now implement β€” defines loading in three concrete stages. Each stage moves more content into context, but only after the previous stage decided it was warranted.

StageWhat loadsWhenTypical size
1. DiscoveryThe skill's name and description from its frontmatterAt session start, for every installed skillA line or two per skill
2. ActivationThe full body of SKILL.mdWhen the agent judges the task matches the descriptionUp to a few thousand words
3. ExecutionBundled scripts, reference files, templates, assetsWhen the loaded instructions tell the agent to read or run themUnbounded β€” as large as the task needs

Stage 1 is the only cost every skill imposes unconditionally, and it's deliberately tiny. An agent running twenty skills pays for twenty short descriptions, not twenty methodologies. Stage 2 is where the real content shows up, and only for the skill (or skills) that matched. Stage 3 is where a skill can genuinely be enormous β€” a full eval harness, a set of JSON schemas, a directory of examples β€” without that size ever touching the context window unless the agent specifically reaches for one of those files.

Why the description is the load-bearing part

Because stage 1 is the only thing evaluated for every skill, on every relevant turn, the description is doing more work than most people writing a SKILL.md initially assume. A vague description ("helps with code") either never triggers or triggers constantly for the wrong tasks, and both failure modes are more common than a skill that's actually broken. A precise one β€” naming the concrete situation, the file types, or the phrase a user is likely to type β€” is what makes stages 2 and 3 possible in the first place.

Skill Creator, Anthropic's own skill for building skills, treats this as the central problem rather than a footnote: it ships scripts specifically for improving a skill's description so it triggers at the right time, alongside tooling for grading whether a candidate skill actually behaves as intended once it does load.

Skill Creator 🧩 SkillFree

A skill for writing new skills: structure, descriptions and testing

β˜… 179k Β· +1.2k this week

The same logic explains why well-designed skill collections stay small per-skill rather than building one sprawling do-everything file. Matt Pocock's Skills, for example, splits alignment, planning, building, and review into more than a dozen narrowly-scoped skills β€” tdd, to-spec, code-review, diagnosing-bugs β€” each with a description specific enough that the agent (or the developer typing a slash command) reaches for the right one instead of a single skill trying to cover every case and matching none of them precisely.

Matt Pocock's Skills 🧩 SkillOpen source

Small, composable engineering skills: grilling, TDD, specs, tickets and architecture review

β˜… 274k Β· +4.9k this week

Close-up of someone sketching a plan in a spiral notebook at a desk
Photo by Ivan S on Pexels

Skills, MCP tools, and subagents handle this differently

Progressive disclosure is specific to the Agent Skills format β€” it's not how every extension mechanism in the ecosystem manages context, and mixing them up is a common source of surprise when a session feels heavier than expected.

MechanismWhat's always in contextWhat's conditional
Agent SkillsName + description per skillFull instructions, scripts, references
MCP serversThe full tool schema for every connected serverOnly the tool call's result, not the schema
SubagentsNothing β€” the parent sees only a summaryThe subagent's entire working context is separate and discarded when it finishes

An MCP server's tool definitions β€” names, parameter schemas, descriptions β€” are sent to the model up front for every connected server, every turn, whether or not you use them that session. That's a real, ongoing cost that scales with how many servers you connect, which is one reason a directory like this one separates "MCP servers" from "Skills" rather than treating them as interchangeable ways to extend an agent; see the full MCP and skills ecosystem for how many of each a given app currently supports. Subagents go the other direction entirely: a subagent's context is walled off from the parent's, so nothing about how it did its work β€” not even the equivalent of a "description" β€” persists in the main thread once it hands back a result.

Knowing which bucket a given integration falls into tells you where to look when a session gets sluggish. A slow, verbose response after installing five skills is probably a stage-2 problem β€” something matched and loaded when it shouldn't have. The same slowness after connecting three new MCP servers is a stage-1-equivalent problem for MCP: you're paying their schema cost on every turn regardless of use.

A minimal SKILL.md that respects the budget

The format itself is simple enough that the discipline is almost entirely about restraint, not syntax. A skill folder needs at minimum a SKILL.md with frontmatter and a body:

---
name: changelog-entry
description: Write a changelog entry for a merged PR, following this repo's CHANGELOG.md format and tone. Use when asked to update the changelog or summarize a release.
---

# Changelog Entry

1. Read the most recent entries in CHANGELOG.md to match tone and format.
2. Summarize the PR's user-facing effect in one line, not its implementation.
3. Group under the correct heading (Added / Changed / Fixed) based on CHANGELOG.md's existing convention.
4. If the format is unclear from recent entries, check reference.md for the full style guide.

Notice what's missing from the body: the full style guide. That lives in a separate reference.md file inside the same folder, mentioned only by name. It only gets read β€” and only then does it cost any context β€” if step 4 actually triggers, which for most changelog entries it won't. That's the pattern in miniature: keep the always-loaded description sharp, keep the body short enough to always be worth loading once matched, and push everything else one level further down.

Auditing a skills folder that's grown too large

Skills accumulate the same way browser tabs do, and the failure modes are predictable enough to check for directly:

SymptomLikely causeFix
Skill never triggersDescription too vague or too narrowRewrite the description around the exact phrases a real task would use
Skill triggers on unrelated tasksDescription too broadNarrow it to the specific file types, commands, or situations it actually handles
Session feels slower after loading a skillStage 2 body is doing stage 3's jobMove reference material, examples, and long lookups into separate files the body only points to
Two skills fight over the same taskOverlapping descriptionsMerge them, or make each description name what specifically distinguishes its case

None of this requires special tooling to check by hand, but Skill Creator's scripts exist precisely because "does this description trigger correctly" is hard to judge just by reading it β€” it benefits from running actual test prompts with and without the skill installed and comparing the results.

Someone arranging books on minimalist white shelves
Photo by Thirdman on Pexels

Where progressive disclosure still breaks down

The model isn't foolproof. A description can be precise and still lose to a more confident-sounding neighbor when two skills' triggers genuinely overlap β€” the agent has to pick one interpretation of an ambiguous request, and it won't always pick the one you intended. Bundled reference files are only as useful as the body's pointers to them: a SKILL.md that never tells the agent "check pricing-tiers.json for exact numbers" leaves that file dead weight, technically stage-3-eligible but functionally unreachable. And because discovery happens per-skill but the agent reasons about the whole set at once, a folder with dozens of similarly-scoped skills can slow down decision-making even though each individual skill's stage-1 cost is negligible β€” the aggregate judgment call, not the token count, becomes the bottleneck.

None of that is an argument against installing skills liberally. It's an argument for treating the description as the interface you're actually designing, since it's the one part of the skill that's never optional.

The takeaway

Progressive disclosure is why Agent Skills scale the way MCP tool schemas don't: cost follows relevance instead of following installation. If a skill in your setup feels like it's either never showing up or showing up everywhere, the fix is almost never a bigger SKILL.md β€” it's a sharper description, with everything else pushed one stage further down. Start there before you start trimming your skills folder, and check the Agent Skills specification itself if you want the mechanism in the format's own words rather than secondhand. For a broader sense of how skills, subagents and MCP servers divide up context differently across apps, the ecosystem overview breaks down what each tool in the directory actually supports, and the glossary is worth a look if any of the terms here β€” SKILL.md, MCP, subagent β€” are still new.

Mentioned in this post

Matt Pocock's Skills 🧩 SkillOpen source

Small, composable engineering skills: grilling, TDD, specs, tickets and architecture review

β˜… 274k Β· +4.9k this week

Skill Creator 🧩 SkillFree

A skill for writing new skills: structure, descriptions and testing

β˜… 179k Β· +1.2k this week

More from the blog

02 Oct 2026 Β· 8 min read

A Vibe Coding Stack for a DevOps Engineer

Claude Code, a DevOps skill pack, a scanning layer, an agent dashboard, and three scoped MCP servers for errors, clusters and tickets.

devops claude-code mcp-servers

Building your vibe coding stack?

Browse 296 apps, skills, subagents and MCP servers, mapped to the apps they work with.

See the ecosystem map