agent-skills-toolkit
Whole-library tier-compliance evaluation against the Advanced Skill Library Standard. Rendered from the one deterministic report object; the verdict is the gate's.
Executive summary
A derived, plain-language read of the deterministic result.
How to read this report
The colored matrix above is the whole verdict in one glance: every check is a chip, color-coded pass / fail / warn / not-applicable, grouped by tier. The tier is decided by a deterministic gate with a real exit code, not by opinion. Everything below expands that picture.
agent-skills-toolkit declares the Gold (Advanced) tier and earns Gold. Of the 35 checks in the spine, 35 do not fail (32 pass, 0 warn, 3 not applicable) and 0 fail. The deterministic gate exits 0.
No requirement blocks the declared tier; the library satisfies its claimed grade outright.
What was evaluated
The subject identity, then the component inventory.
| Component | Type | Version |
|---|---|---|
| askit-backlog | skill | 0.1.0 |
| askit-build-agents-md | skill | 0.1.0 |
| askit-build-chain-contract | skill | 0.1.0 |
| askit-build-command | skill | 0.1.0 |
| askit-build-docs | skill | 0.2.0 |
| askit-build-hook | skill | 0.1.0 |
| askit-build-mcp | skill | 0.1.0 |
| askit-build-output-style | skill | 0.1.0 |
| askit-build-samples | skill | 0.1.0 |
| askit-build-settings | skill | 0.1.0 |
| askit-build-skill | skill | 0.1.3 |
| askit-build-statusline | skill | 0.1.0 |
| askit-build-subagent | skill | 0.2.0 |
| askit-build-workflow | skill | 0.1.0 |
| askit-capability-advisor | skill | 0.1.0 |
| askit-capability-gap-analysis | skill | 0.1.0 |
| askit-capability-whats-new | skill | 0.1.0 |
| askit-decision | skill | 0.1.0 |
| askit-deprecate | skill | 0.1.0 |
| askit-evaluate | skill | 0.1.3 |
| askit-init-marketplace | skill | 0.2.0 |
| askit-init-plugin | skill | 0.1.0 |
| askit-migrate | skill | 0.1.0 |
| askit-release | skill | 0.1.0 |
| askit-standards-watch | skill | 0.1.0 |
| askit-template-manager | skill | 0.1.0 |
| askit-skill-author | subagent | 0.1.2 |
| askit-evaluator | subagent | 0.1.0 |
| askit-explorer | subagent | 0.1.0 |
| askit-file-search | subagent | 0.1.0 |
| askit-file-ops | subagent | 0.1.0 |
| askit-reviewer | subagent | 0.1.1 |
| askit-quality-grader | subagent | 0.1.1 |
| askit-evaluate | command | 0.1.0 |
| askit-build-skill | command | 0.1.0 |
Methodology and scope
Three layers, kept separate so the verdict stays honest. Only the deterministic gate decides the tier.
Decides the tier. The portable Node gate runs every check with a real exit code, no model in the loop. Re-run node scripts/check.mjs . to reproduce it.
Advisory only. A skill is run against its eval set; results are evidence and never decide a gate result.
Advisory only. A qualitative pass over scoping and design; it informs the insights, it does not move the badge.
Status legend
Confidence and limitations
Conformance findings are exact and reproducible: the same gate on the same commit yields the same result. Vacuous passes (U11, G6, G1) mean there was nothing to validate, not that a feature was exercised.
Tier compliance - evidence ledger
One row per requirement, all 35 itemized and grouped by tier. Every non-pass carries a why-it-matters note.
Requirement satisfied; no finding raised. checks/library-json.mjs
Requirement satisfied; no finding raised. checks/anatomy.mjs
Requirement satisfied; no finding raised. checks/frontmatter-valid.mjs
Requirement satisfied; no finding raised. checks/name-matches-dir.mjs
Requirement satisfied; no finding raised. checks/description-score.mjs
Requirement satisfied; no finding raised. checks/reference-links.mjs
Requirement satisfied; no finding raised. checks/instruction-budget.mjs
Requirement satisfied; no finding raised. checks/manifest-drift.mjs
Requirement satisfied; no finding raised. checks/version-match.mjs
Requirement satisfied; no finding raised. checks/mermaid-valid.mjs
Nothing to validate for this subject (vacuous pass). checks/mcp-valid.mjs
Requirement satisfied; no finding raised. checks/skill-registration.mjs
Requirement satisfied; no finding raised. checks/agent-restricted-fields.mjs
Requirement satisfied; no finding raised. checks/agents-dir-registerable.mjs
Requirement satisfied; no finding raised. checks/metadata-placement.mjs
Requirement satisfied; no finding raised. checks/catalogue-manifest-shape.mjs
Requirement satisfied; no finding raised. checks/command-size-cap.mjs
Requirement satisfied; no finding raised. checks/agent-targets.mjs
Requirement satisfied; no finding raised. checks/prefix.mjs
Requirement satisfied; no finding raised. checks/components-index.mjs
Requirement satisfied; no finding raised. checks/components-mirror.mjs
Requirement satisfied; no finding raised. checks/chain-contract.mjs
Requirement satisfied; no finding raised. checks/command-contract.mjs
Requirement satisfied; no finding raised. checks/workflow-skills.mjs
Requirement satisfied; no finding raised. checks/per-target-presence.mjs
Requirement satisfied; no finding raised. checks/library-regression.mjs
Nothing to validate for this subject (vacuous pass). checks/deprecation.mjs
Nothing to validate for this subject (vacuous pass). checks/hook-documentation.mjs
Requirement satisfied; no finding raised. checks/self-hosting.mjs
Requirement satisfied; no finding raised. checks/release-notes.mjs
Requirement satisfied; no finding raised. checks/index-drift.mjs
Requirement satisfied; no finding raised. checks/docs-frontmatter.mjs
Requirement satisfied; no finding raised. checks/folder-readme.mjs
Requirement satisfied; no finding raised. checks/source-doc.mjs
Requirement satisfied; no finding raised. checks/docs-presence.mjs
The climb / burndown
No blockers; the declared tier is satisfied.
No blockers: the declared tier is satisfied. There is nothing to climb.
Improvement path
One card per gap: the issue, priority and effort, and a copy-paste prompt that drives the matching askit builder and re-runs the gate.
No action required; nothing failed or warned.
Insights
None for a deterministic conformance report.
No insights in a conformance report
Insights are produced by the advisory review layer (askit-reviewer), which a deterministic conformance report does not run. Run review mode to populate this section; a conformance report does not fabricate qualitative notes.
Evidence and sources
Citations grounding the findings: the check module, Standard clause, or subject file.
- CHECKscripts/check.mjs - the portable deterministic gate (
node scripts/check.mjs .); its exit code drives the verdict. - CLAUSEStandard v0.16 - the 35-check spine and the Bronze / Silver / Gold tier definitions.
- FILElibrary.json - the subject identity (name, version, tier, agent-targets, prefix).
Report metadata
Provenance for this evaluation, plus the legend.
Status legend
Generated by the askit-evaluate report renderer. The conformance layer is deterministic and reproducible; re-run node scripts/check.mjs . to reproduce every finding. This report adds no judgment and does not change the verdict.
Per-check glossary
What each of the 35 checks verifies, in one line - a plain-language reference for every row above.
| Check | Tier | What it verifies |
|---|---|---|
U1 library-json | Bronze | Without a valid library.json a tool cannot identify the library, its version, or the Standard it pins, so nothing downstream can grade, install, or emit it. |
U2 anatomy | Bronze | The agentskills.io anatomy (a root AGENTS.md and the standard component folders) is how any agent discovers what the library contains; a broken anatomy makes the library unreadable to the tools meant to load it. |
U3 frontmatter-valid | Bronze | A component whose frontmatter does not parse, or lacks a name or description, cannot be loaded or selected by an agent; the parse is the contract between the file and the runtime. |
U4 name-matches-dir | Bronze | When a component's declared name does not match its directory, references and manifests point at the wrong place; a kebab-case directory that equals the name keeps lookups unambiguous across agents. |
U5 description-score | Bronze | A description below the clarity bar makes a skill hard for an agent to select for the right job; it may fail to fire when it should, or fire when it should not. |
U6 reference-links | Bronze | A reference link that does not resolve sends the agent to a missing file mid-task, breaking progressive disclosure exactly when the detail is needed. |
U7 instruction-budget | Bronze | A body over the instruction budget risks the model dropping its later steps, so the exact step that differentiates the skill can be the one lost at runtime. |
U8 manifest-drift | Bronze | When a native manifest disagrees with library.json, the agent loads something other than what the library declares; generated (not hand-edited) manifests are what keep the two in lockstep. |
S1 agent-targets | Silver | Without a declared agent-targets list the library does not say which agents it converges across, so the Convergent guarantees (matching manifests, per-target presence) have nothing to check against. |
S2 prefix | Silver | A consistent prefix is how a multi-skill library avoids name collisions on a shared agent and signals which components belong to it; an unprefixed component is ambiguous once installed beside others. |
S3 components-index | Silver | When the components index does not list what is on disk, tools that read the index (install, emit, grade) miss real components or reference ones that do not exist. |
S8 components-mirror | Silver | A one-way index lets the index and disk drift apart silently; mirroring in both directions catches both an orphan on disk and a phantom in the index. |
S4 chain-contract | Silver | An orphan or phantom chain edge means a declared delegation points at nothing, or a real delegation is undeclared; the chain contract is what makes cross-component handoffs honest and reviewable. |
S7 command-contract | Silver | A command that maps to zero or many skills is ambiguous at invocation; one command to exactly one skill keeps the slash entry point predictable. |
S5 workflow-skills | Silver | A workflow step that references a skill which does not exist breaks the run at that step; validating the references keeps a multi-step workflow executable end to end. |
S6 per-target-presence | Silver | If a declared agent target is missing its native manifest, the library claims to converge on an agent it cannot actually load onto; per-target presence makes the cross-agent claim real. |
U9 version-match | Bronze | When a component version disagrees with library.json, release tooling and consumers cannot tell which version they actually have; the versions must agree to make a release honest. |
U12 mermaid-valid | Bronze | A mermaid diagram that does not parse renders as a broken block in the docs and the docs site, so a diagram meant to explain the library instead signals it is unmaintained. |
U11 mcp-valid | Bronze | An MCP server definition that is malformed or carries an inline secret either fails to connect or leaks a credential into the repository; validating it keeps the integration safe and loadable. |
U13 skill-registration | Bronze | A skill on disk that the manifest does not register ships but is invisible to installers; the catalog must list everything the library delivers, and a registered skill with no directory cannot be delivered at all. |
U14 agent-restricted-fields | Bronze | Claude Code refuses hooks, mcpServers and permissionMode on a plugin-shipped agent for security reasons, so an agent declaring one has configured something the runtime will not honour - and nothing tells the author, which is the whole hazard. |
U15 agents-dir-registerable | Bronze | Claude Code scans agents/ for *.md and registers every file it finds, so a file the plugin excludes from its own registration is still a live subagent - with whatever name and frontmatter it happens to carry - and it escapes every check that reads the registration list. |
U16 metadata-placement | Bronze | The Standard places version, tier, status and the other governance keys under `metadata`; declared at the top level they are read by nothing, so an author who wrote a version has not recorded one and receives no signal either way. |
U17 catalogue-manifest-shape | Bronze | A .claude-plugin/marketplace.json that cannot be parsed, carries no plugins array, or mixes skill-source and plugin-source entries is read by no scope or by only one of two - so entries the author declared are catalogued by nothing, with no signal either way. |
U18 command-size-cap | Bronze | Codex migrates a plugin's commands into skills and SKIPS any whose rendered skill exceeds its size cap - no skill is written and no error is raised, so an oversized command simply does not exist on Codex while the repository still shows it. |
G3 library-regression | Gold | Chain edges are where delegation breaks quietly; a regression eval per edge turns the chain contract from a declaration into a tested guarantee that a refactor cannot sever unnoticed. |
G6 deprecation | Gold | A component removed or replaced without following the deprecation policy breaks consumers who depended on it; an explicit deprecation gives them a status, a reason, and a migration path. |
G1 hook-documentation | Gold | An undocumented hook changes the agent's behavior invisibly; documenting each hook is what lets a reviewer and a consumer see what fires and why before they install it. |
G2 self-hosting | Gold | Without CI running the gate, conformance is a claim, not a proof. Any change can silently regress the library, and a consumer cannot point to a green badge that says the standard held on the latest commit. |
G5 release-notes | Gold | A changelog is for maintainers; release notes are for users. Without a curated RELEASE-NOTES.md, adopters have no human-readable summary of what changed and why they should upgrade. |
G4 index-drift | Gold | A stale INDEX.md misrepresents the library to anyone reading the catalog; generating it and drift-checking in CI keeps the published index honest on every commit. |
G7 docs-frontmatter | Gold | The audience/level/doc-role frontmatter taxonomy is what lets the docs site route a reader to the right page; without it the generated site cannot organize the library's documentation. |
G8 folder-readme | Gold | A meaningful folder with no README leaves a reader guessing what it holds; a short folder guide is the cheapest orientation a contributor or consumer gets. |
G9 source-doc | Gold | A script with no what-it-is / what-it-does / why docblock forces a maintainer to reverse-engineer its purpose; the four-field header keeps the toolkit's own machinery legible. |
G10 docs-presence | Gold | A library without the Diataxis quadrants (tutorials, how-to, reference, explanation) leaves whole classes of reader unserved; their presence is what makes the documentation complete rather than incidental. |