SKILL.md in 2026: Authoring, Portability, Security
Agent Skills became one of the most widely adopted ways to package portable agent capabilities in 2026. A production-grade SKILL.md is not a prompt file; it is a permissioned, progressively disclosed instruction surface, and the portability decision you cannot skip is whether you will run third-party skills with your agent’s full ambient authority.
Anthropic introduced Agent Skills on Oct 16, 2025, then published the format as an open standard on Dec 18, 2025 at agentskills.io. That two-step history matters: one vendor proved the pattern, then the specification made it portable across harnesses.
Why skills won 2026 in four numbers
Skills spread fast in 2026 because they standardized portable agent capabilities with tiny startup cost, broad harness support, and massive open-source reuse. Anthropic launched the format in October, opened the standard in December, and registries plus star counts now prove adoption. Four numbers below explain why teams treat SKILL.md as infrastructure.
First, context economy: about 100 tokens per skill at startup. Second, distribution: a client roster that includes Hermes Agent, Claude, Cursor, GitHub Copilot, Goose, OpenHands, Amp, Tabnine, Roo Code, JetBrains Junie, Qodo, Spring AI, Snowflake Cortex Code, Pulumi Neo, OpenClaw, and Command Code, listed on agentskills.io/clients. Third, reuse at scale measured via the GitHub API on 2026-10-05: obra/superpowers 295,590 stars, anthropics/skills 179,764, addyosmani/agent-skills 101,480, wshobson/agents 40,216, agentskills/agentskills 25,924, microsoft/agent-skills 3,081. For contrast, modelcontextprotocol/servers measured 91,017 stars the same day. Fourth, risk surface: of the 3,984 skills Snyk scanned, 36.82% carried at least one security flaw and none run sandboxed by default — the reason this guide pairs authoring with security.
What a skill actually is
A skill is a directory containing at minimum a SKILL.md file, not a single prompt. Optional scripts/, references/, and assets/ subdirectories add executable code, on-demand docs, and templates or data. That folder contract lets any compatible harness discover, advertise, load, and run the same skill portably.
Per the Agent Skills specification, the minimum shape is:
my-skill/
SKILL.md
scripts/ # executable code, run on demand
references/ # docs loaded on demand
assets/ # templates, images, data
Only SKILL.md is required. scripts/ may be executed by the agent, references/ is loaded as needed to save context, and assets/ holds files the skill produces or consumes. Keep paths relative and explicit so Microsoft Agent Framework, Claude, Cursor, and other harnesses resolve them identically. If you are comparing transport costs, see our MCP token-cost audit.
Frontmatter that fails validation
Most validation failures come from name and description frontmatter, not body prose. Name and description are required, license and compatibility and metadata are optional, and allowed-tools is experimental. The skills-ref checker enforces naming and frontmatter validity only, so authors must get these six fields right before publishing.
The six frontmatter fields are:
name— REQUIREDdescription— REQUIREDlicense— optionalcompatibility— optional, max 500 chars for environment requirementsmetadata— optional, string-to-string mapallowed-tools— optional, EXPERIMENTAL, space-separated pre-approved tools
Exact name rules: at most 64 chars; lowercase a-z, 0-9, and hyphen only; must not start or end with a hyphen; no consecutive hyphens; must match the parent directory name. Exact description rule: at most 1,024 chars; must say what the skill does AND when to use it.
Validate naming and frontmatter only with the skills-ref reference library:
- Name the folder to match
name, for examplepdf-fill/. - Keep
nameat 64 characters or fewer with the hyphen rules above. - Write
descriptionwith capability plus trigger intent in 1,024 characters or fewer. - Run
skills-ref validate ./my-skillbefore publishing. - Remember it checks frontmatter validity and naming conventions ONLY — not body quality or security.
Progressive disclosure is the product
Progressive disclosure keeps hundreds of skills affordable by loading metadata first, the full SKILL.md body only on activation, and resources on demand. Level one costs about one hundred tokens per skill at startup, level two recommends under five thousand tokens, and level three stays lazy. The product is context economy.
The spec levels are:
- Level 1 — metadata: about 100 tokens per skill loaded at startup for ALL skills. This is
nameplusdescription. - Level 2 — full SKILL.md body: under 5,000 tokens recommended, loaded only on activation.
- Level 3 — resources:
references/,assets/, andscripts/loaded or run as needed.
Operational rule: keep SKILL.md under 500 lines and move detail to referenced files. A 900-line SKILL.md with embedded API dumps defeats disclosure, inflates every activation, and hides malicious instructions in noise. Lean bodies with linked references are cheaper, more portable, and easier to audit.
Cross-harness portability: the four-stage variant
The Agent Skills specification and Microsoft Agent Framework share the same disclosure idea but differ in tooling. The specification defines three levels, while Microsoft implements a four-stage variant with Advertise, Load, Read resources, and Run scripts. Understanding both lets you author skills that port cleanly across harnesses.
Microsoft Agent Framework implements a four-stage variant, not the official agentskills.io spec verbatim. Microsoft notes load_skill is always advertised; the other two tools appear only when a skill has resources or scripts. Author against the spec, then test the Microsoft tool mapping.
| Stage | Agent Skills spec (agentskills.io) | Microsoft Agent Framework | Load trigger | Budget |
|---|---|---|---|---|
| Startup metadata / Advertise | name + description frontmatter | Advertise | At startup for all installed skills | ~100 tokens per skill |
| Full skill body / Load | SKILL.md body | Load (load_skill tool) | On activation / intent match | under 5,000 tokens recommended |
| Referenced resources / Read | references/ and assets/ loaded as needed | Read resources (read_skill_resource) | When a referenced file is needed | On demand |
| Optional scripts / Run | scripts/ may be executed | Run scripts (run_skill_script) | When a process must run | On demand |
Client adoption listed on agentskills.io/clients includes Hermes Agent, Claude, Cursor, GitHub Copilot, Goose, OpenHands, Amp, Tabnine, Roo Code, JetBrains Junie, Qodo, Spring AI, Snowflake Cortex Code, Pulumi Neo, OpenClaw, and Command Code. For marketplace dynamics across harnesses, see our multi-harness plugin marketplace analysis.
Authoring a portable SKILL.md body
A portable body triggers on user intent, stays lean, and delegates detail to references, assets, and scripts. Write a precise description that states what the skill does and when to use it, keep SKILL.md under five hundred lines, and use explicit file references. Portability comes from structure, not clever prompting.
Follow this body pattern:
- Write the
descriptionfor retrieval: verb plus object plus when-to-use, for example fills PDF forms when a user provides a template plus field data. - Open SKILL.md with purpose, prerequisites, and one minimal happy-path example.
- Push long schemas, edge cases, and transcripts to
references/. - Put templates and sample data in
assets/with relative links. - Put deterministic work in
scripts/with explicit invocation commands and expected inputs. - Declare environment needs in
compatibilityso non-matching harnesses skip cleanly.
Keep instructions imperative and tool-agnostic. Keep harness-specific tool names out of the body and declare environment requirements in compatibility instead, so the same skill activates cleanly on every compatible client. For injection-resistant instruction design, see our prompt-injection defense guide.
The security decision most guides skip
The security decision is whether third-party skills should run with your agent’s full ambient authority. An installed skill inherits shell access, filesystem rights, credentials, messaging, and memory, while public corpora already contain critical flaws and confirmed malicious payloads. You must design trust boundaries before you install anything.
The Snyk ToxicSkills study, published Feb 5, 2026, scanned 3,984 skills from ClawHub and skills.sh. It found 13.4% with at least one CRITICAL issue, 36.82% with at least one security flaw, and 76 confirmed malicious payloads covering credential theft, backdoors, and exfiltration. An installed skill inherits the agent’s full permissions: shell access, filesystem read/write, credentials in env/config, messaging, and persistent memory. The publishing barrier on ClawHub is a SKILL.md plus a one-week-old GitHub account, with no code signing, no security review, and no sandbox by default.
Published security research agrees. arXiv 2510.26328 shows malicious instructions hidden in long skill files and referenced scripts can exfiltrate internal files or passwords, and a benign task-specific approval using the Don’t ask again option can carry over to closely related harmful actions. Long files plus overbroad approvals are the exploit path. For the parallel tool-trust problem, read our MCP tool-poisoning trust-model breakdown.
A production install checklist
Production installation is a governance workflow covering source trust, content review, sandboxing, allowlisting, and version pinning. Public registries lower the publishing barrier dramatically, so every install needs explicit approval criteria and runtime limits. This checklist gives platform teams a repeatable gate for third-party skills without blocking adoption.
- Trust the source, not the registry: prefer first-party or vendored skills; treat public downloads as untrusted input.
- Review triggers: always review
SKILL.mdplusscripts/andreferences/for network calls, credential access, and exfiltration. - Least privilege: run with scoped tools, read-only mounts where possible, and short-lived tokens. See our credential-scoped agents guide.
- Sandbox and allowlist: sandbox third-party code by default; where supported, pre-approve tools via
allowed-tools(advisory and experimental, never a security boundary) and deny ambient shell. - Pin versions: pin commit hashes, re-review on update, and log skill version per run.
- Registry risk: neither skills.sh nor askill.sh should be treated as curated or safe given the Snyk ToxicSkills findings.
Do not disable third-party skills entirely. Frame the call as trust plus sandbox plus allowlist plus pinning, so useful skills ship safely.
Publishing for portability
Publishing for portability means declaring compatibility clearly, adding structured metadata, and treating allowed-tools as experimental. Compatibility documents environment requirements in under five hundred characters, metadata provides string-to-string routing hints, and registries distribute your skill. Precise frontmatter helps every compatible harness activate correctly.
Set compatibility to environment requirements in at most 500 chars, for example requires python 3.10 plus poppler for PDF work. Use metadata only as a string-to-string map for routing hints like team, maturity, or region. Treat allowed-tools as EXPERIMENTAL space-separated pre-approved tools and never rely on it as a security boundary. Distribution options include skills.sh with its add-skills CLI and askill.sh, an SKILL.md registry that states it works with 40+ agents. Both lower distribution friction, but neither implies safety review — apply the install checklist above. For governance policy templates, see our open-source agent tool governance roundup.
FAQ
Does validation make a skill safe, how many tokens each skill costs, and how execution differs across harnesses are the three questions teams ask before shipping SKILL.md to production.
Does skills-ref validation mean a skill is safe?
No. The skills-ref validator checks frontmatter validity and naming conventions only, not body quality or security. A passing skill can still contain prompt injection, exfiltration code, or overbroad permissions. Treat validation as a syntax gate, then run content review, sandbox testing, and allowlist approval before production use.
How many tokens does each skill cost at startup?
Each installed skill costs about one hundred tokens at startup for name plus description metadata. The full SKILL.md body loads only on activation and should stay under five thousand tokens, with references, assets, and scripts loading on demand. Keep SKILL.md under five hundred lines to protect context.
Will one SKILL.md run unchanged in Microsoft Agent Framework?
Usually yes for content, but execution differs. The specification uses three disclosure levels, while Microsoft implements a four-stage variant with Advertise, Load, Read resources, and Run scripts using load_skill, read_skill_resource, and run_skill_script tools. Author portable paths and declare compatibility so both models resolve resources correctly.
What is the safest way to try third-party skills?
Pin a reviewed version, install least-privilege, and sandbox by default with explicit tool allowlists and short-lived credentials. Review SKILL.md plus scripts and references for exfiltration, and treat public registries as untrusted after the ToxicSkills findings. For credential hygiene, see our credential-scoped agents guide.
The Bottom Line
Treat every SKILL.md as permissioned infrastructure with progressive disclosure, strict frontmatter, and explicit trust boundaries. Author lean portable bodies, pin reviewed versions, and sandbox third-party code by default. The portability win is real, but no third-party skill should run with your agent’s full ambient authority — apply least privilege whether or not it is audited.
Ship one well-scoped skill with correct name and description, an under-500-line body, declared compatibility, and a pinned release. Measure startup tokens, activation rate, and sandbox denials. Then expand the catalog. Portability plus least privilege is the production pattern — visit the NiteAgent arena to compare implementations.
How This Guide Was Built
This guide synthesizes the Agent Skills specification, Microsoft Agent Framework documentation, client adoption data, GitHub star measurements, and published security research. We prioritized exact limits, token budgets, and implementation differences over anecdotes. Every claim traces to a linked primary source you can verify.
This guide is based on the Agent Skills specification, vendor documentation, and published security research — we did not install or run any third-party skill while writing it. Star counts were measured via the GitHub API on 2026-10-05. Security figures come from Snyk and arXiv sources linked above.
Related reading
- our MCP token-cost audit
- our credential-scoped agents guide
- our open-source agent tool governance roundup
- our prompt-injection defense guide
- our MCP tool-poisoning trust-model breakdown
- our multi-harness plugin marketplace analysis
📖 Related Reads
- Hermes Tutorials — Hermes Agent setup, configuration, and advanced workflows
- NoCode Insider — AI workflow automation with no-code tools, agents, and APIs
Cross-links automatically generated from NiteAgent.
← Back to all posts


