02 Oct 2026 Β· 8 min read
A Vibe Coding Stack for a DevOps Engineer
Claude Code, a DevOps skill pack, a scanning layer, an agent dashboard, and three scoped MCP servers for errors, clusters and tickets.
devops claude-code mcp-serversBy Kelvin Β· 29 September 2026 Β· Updated 29 Sep 2026 Β· 8 min read

Install ten Agent Skills into a coding agent and something you'd expect to happen doesn't: the agent's context window doesn't fill up with ten skills' worth of instructions. That's not an accident, and it's not a limit you're about to hit later. It's the core design decision behind the Agent Skills format, and it's called progressive disclosure β the idea that an agent should only pay the token cost for a capability at the moment it actually needs it, not the moment it becomes available.
If you've written a skill and wondered why your SKILL.md "isn't triggering," or you've watched a skills-heavy agent stay surprisingly fast despite a folder full of .md files, this is the mechanism underneath both. Understanding it changes how you write skills, how many you can reasonably install, and how you debug the ones that misbehave.
Every token an agent's system prompt carries is a token it can't spend on your actual conversation, your codebase, or its own reasoning. Coding agents already compete for that budget with the file you're editing, your AGENTS.md or CLAUDE.md instructions, recent shell output, and β critically β the full tool schemas for every MCP server you've connected. None of that is free.
Agent Skills exist to add capability without adding to that fixed cost. A skill can be as large as it needs to be β a full TDD methodology, a 40-page document-formatting reference, a validator script and three example files β because almost none of that content sits in the agent's context by default. Only a sliver of it does, all the time. The rest loads conditionally, and most of it may never load in a given session at all.
This is what separates a skills folder from, say, pasting ten reference documents into your system prompt. The documents approach charges you for content whether or not it's relevant to the current task. Progressive disclosure charges you only when the content earns its place.
The Agent Skills specification β the open format Anthropic released and that a growing list of clients now implement β defines loading in three concrete stages. Each stage moves more content into context, but only after the previous stage decided it was warranted.
| Stage | What loads | When | Typical size |
|---|---|---|---|
| 1. Discovery | The skill's name and description from its frontmatter | At session start, for every installed skill | A line or two per skill |
| 2. Activation | The full body of SKILL.md | When the agent judges the task matches the description | Up to a few thousand words |
| 3. Execution | Bundled scripts, reference files, templates, assets | When the loaded instructions tell the agent to read or run them | Unbounded β as large as the task needs |
Stage 1 is the only cost every skill imposes unconditionally, and it's deliberately tiny. An agent running twenty skills pays for twenty short descriptions, not twenty methodologies. Stage 2 is where the real content shows up, and only for the skill (or skills) that matched. Stage 3 is where a skill can genuinely be enormous β a full eval harness, a set of JSON schemas, a directory of examples β without that size ever touching the context window unless the agent specifically reaches for one of those files.
Because stage 1 is the only thing evaluated for every skill, on every relevant turn, the description is doing more work than most people writing a SKILL.md initially assume. A vague description ("helps with code") either never triggers or triggers constantly for the wrong tasks, and both failure modes are more common than a skill that's actually broken. A precise one β naming the concrete situation, the file types, or the phrase a user is likely to type β is what makes stages 2 and 3 possible in the first place.
Skill Creator, Anthropic's own skill for building skills, treats this as the central problem rather than a footnote: it ships scripts specifically for improving a skill's description so it triggers at the right time, alongside tooling for grading whether a candidate skill actually behaves as intended once it does load.
A skill for writing new skills: structure, descriptions and testing
The same logic explains why well-designed skill collections stay small per-skill rather than building one sprawling do-everything file. Matt Pocock's Skills, for example, splits alignment, planning, building, and review into more than a dozen narrowly-scoped skills β tdd, to-spec, code-review, diagnosing-bugs β each with a description specific enough that the agent (or the developer typing a slash command) reaches for the right one instead of a single skill trying to cover every case and matching none of them precisely.
Small, composable engineering skills: grilling, TDD, specs, tickets and architecture review

Progressive disclosure is specific to the Agent Skills format β it's not how every extension mechanism in the ecosystem manages context, and mixing them up is a common source of surprise when a session feels heavier than expected.
| Mechanism | What's always in context | What's conditional |
|---|---|---|
| Agent Skills | Name + description per skill | Full instructions, scripts, references |
| MCP servers | The full tool schema for every connected server | Only the tool call's result, not the schema |
| Subagents | Nothing β the parent sees only a summary | The subagent's entire working context is separate and discarded when it finishes |
An MCP server's tool definitions β names, parameter schemas, descriptions β are sent to the model up front for every connected server, every turn, whether or not you use them that session. That's a real, ongoing cost that scales with how many servers you connect, which is one reason a directory like this one separates "MCP servers" from "Skills" rather than treating them as interchangeable ways to extend an agent; see the full MCP and skills ecosystem for how many of each a given app currently supports. Subagents go the other direction entirely: a subagent's context is walled off from the parent's, so nothing about how it did its work β not even the equivalent of a "description" β persists in the main thread once it hands back a result.
Knowing which bucket a given integration falls into tells you where to look when a session gets sluggish. A slow, verbose response after installing five skills is probably a stage-2 problem β something matched and loaded when it shouldn't have. The same slowness after connecting three new MCP servers is a stage-1-equivalent problem for MCP: you're paying their schema cost on every turn regardless of use.
The format itself is simple enough that the discipline is almost entirely about restraint, not syntax. A skill folder needs at minimum a SKILL.md with frontmatter and a body:
---
name: changelog-entry
description: Write a changelog entry for a merged PR, following this repo's CHANGELOG.md format and tone. Use when asked to update the changelog or summarize a release.
---
# Changelog Entry
1. Read the most recent entries in CHANGELOG.md to match tone and format.
2. Summarize the PR's user-facing effect in one line, not its implementation.
3. Group under the correct heading (Added / Changed / Fixed) based on CHANGELOG.md's existing convention.
4. If the format is unclear from recent entries, check reference.md for the full style guide.
Notice what's missing from the body: the full style guide. That lives in a separate reference.md file inside the same folder, mentioned only by name. It only gets read β and only then does it cost any context β if step 4 actually triggers, which for most changelog entries it won't. That's the pattern in miniature: keep the always-loaded description sharp, keep the body short enough to always be worth loading once matched, and push everything else one level further down.
Skills accumulate the same way browser tabs do, and the failure modes are predictable enough to check for directly:
| Symptom | Likely cause | Fix |
|---|---|---|
| Skill never triggers | Description too vague or too narrow | Rewrite the description around the exact phrases a real task would use |
| Skill triggers on unrelated tasks | Description too broad | Narrow it to the specific file types, commands, or situations it actually handles |
| Session feels slower after loading a skill | Stage 2 body is doing stage 3's job | Move reference material, examples, and long lookups into separate files the body only points to |
| Two skills fight over the same task | Overlapping descriptions | Merge them, or make each description name what specifically distinguishes its case |
None of this requires special tooling to check by hand, but Skill Creator's scripts exist precisely because "does this description trigger correctly" is hard to judge just by reading it β it benefits from running actual test prompts with and without the skill installed and comparing the results.

The model isn't foolproof. A description can be precise and still lose to a more confident-sounding neighbor when two skills' triggers genuinely overlap β the agent has to pick one interpretation of an ambiguous request, and it won't always pick the one you intended. Bundled reference files are only as useful as the body's pointers to them: a SKILL.md that never tells the agent "check pricing-tiers.json for exact numbers" leaves that file dead weight, technically stage-3-eligible but functionally unreachable. And because discovery happens per-skill but the agent reasons about the whole set at once, a folder with dozens of similarly-scoped skills can slow down decision-making even though each individual skill's stage-1 cost is negligible β the aggregate judgment call, not the token count, becomes the bottleneck.
None of that is an argument against installing skills liberally. It's an argument for treating the description as the interface you're actually designing, since it's the one part of the skill that's never optional.
Progressive disclosure is why Agent Skills scale the way MCP tool schemas don't: cost follows relevance instead of following installation. If a skill in your setup feels like it's either never showing up or showing up everywhere, the fix is almost never a bigger SKILL.md β it's a sharper description, with everything else pushed one stage further down. Start there before you start trimming your skills folder, and check the Agent Skills specification itself if you want the mechanism in the format's own words rather than secondhand. For a broader sense of how skills, subagents and MCP servers divide up context differently across apps, the ecosystem overview breaks down what each tool in the directory actually supports, and the glossary is worth a look if any of the terms here β SKILL.md, MCP, subagent β are still new.
Small, composable engineering skills: grilling, TDD, specs, tickets and architecture review
A skill for writing new skills: structure, descriptions and testing

02 Oct 2026 Β· 8 min read
Claude Code, a DevOps skill pack, a scanning layer, an agent dashboard, and three scoped MCP servers for errors, clusters and tickets.
devops claude-code mcp-servers
01 Oct 2026 Β· 8 min read
A verified walkthrough of AGENTS.md rules, SKILL.md packs, subagents, MCP servers, hooks and plugins in Google Antigravity.
google-antigravity setup-guide mcp
30 Sep 2026 Β· 8 min read
A practical walkthrough of Copilot instructions, agent mode, MCP servers and the CLI, with verified commands from the docs.
github-copilot agent-mode mcpBrowse 296 apps, skills, subagents and MCP servers, mapped to the apps they work with.
See the ecosystem map