design-skill-evalslisted
Install: claude install-skill SylphxAI/skills
# design-skill-evals
# Design Skill Evals
Produce a **Skill Evaluation Program** that can disprove whether an exact skill
bundle improves realistic work while preserving routing boundaries, safety,
cost, and evidence integrity.
## When to use
- An exact skill bundle needs falsifiable routing/behavior evals with holdouts and replay
- Promotion claims need candidate-bound digests, controls, oracles, and provenance
- Not for the skill procedure itself (`author-skill`) or portfolio decisions (`curate-skill-repository`)
## Atomic boundary
Own the eval contract: candidate digests, routing/behavior tasks, controls,
rubrics/oracles, artifact checks, model/provider matrix, hidden-data handling,
metrics/uncertainty, provenance, replay, attestation, expiry, and regression
triage. Do not own the skill procedure, repository portfolio decision, runtime
capability, or adoption/outcome telemetry.
## Workflow
1. Freeze the question: exact skill name/description/body/references/scripts,
intended jobs/artifacts, nearest neighbours, risks, output budget, supported
runtimes/tools, and claim that the eval may falsify.
2. Read `references/skill-eval-systems.md`. Compute an injection-contract digest over
name+description and behavior digest over the complete ordered bundle. Bind
candidate commit, catalog digest, task/rubric/runner/policy/model-registry
digests, parameters, seed, tool availability, retries, and expiry.
For repeated or continuously monitored evaluations, also