← ClaudeAtlas

create-skraft-evallisted

Use when creating, refreshing, expanding, or reviewing a SKRAFT Vally skill evaluation at tests/skills/<skill>/eval.yaml in the skraft-plugin repository. Covers behavior coverage, baseline-versus-isolated-treatment discrimination, natural prompts, outcome rubrics, non-activation cases, regression guards, fixtures, static Vally validation, trial budgeting for statistical power, staged live spend, and optional paired measurement. Do not use for dotnet/skills evals, generic skill-test scaffolding, agent suites, skill authoring, or debugging an already-running evaluation.
SebastienDegodez/skraft-plugin · ★ 8 · AI & Automation · score 68
Install: claude install-skill SebastienDegodez/skraft-plugin
# Create a SKRAFT Vally Skill Evaluation Design one trustworthy SKRAFT skill evaluation. Ground it, approve it before writing, validate it statically, then freeze it during optional measurement. ## When to use Use for creating, refreshing, expanding, or reviewing `tests/skills/<skill>/eval.yaml`. Do not use for dotnet/skills, real-agent suites, target-skill authoring, live-run debugging, eval-spec tests, or optimization that rewrites the instrument during measurement. ## Non-negotiable rules 1. Read current evidence before proposing scenarios. Do not design from memory. 2. Never name the target skill in a stimulus prompt or copy distinctive wording from its body. 3. Write prompts and rubrics in English. Prompts sound like natural developer requests and describe WHAT must be achieved, never HOW to implement it. Rubrics judge observable outcomes, not methods, commands, labels, or skill vocabulary. 4. Budget trials for statistical power, not for a floor. The verdict is a two-sided sign test on discordant pairs, which cannot reach `p <= 0.05` below six of them: a flawless 5W/0L sweep scores 0.0625. Ties consume pairs, so plan 12-15 total trials (`stimuli x runs`) and never fewer than the repository floor of five. 5. Spend the budget on runs before breadth. `3 stimuli x 5 runs` and `5 stimuli x 3 runs` cost the same, and only the first produces per-scenario cells worth reading. Prefer three or four stimuli chosen by rank over one stimulus per case class. 6. Keep each stimulus