happier-instruction-eval

Featured

Evaluate two or more Happier instruction, constitution, or skill variants with blinded organic tasks, controlled context, behavior-based scoring, privacy-safe evidence, and an advisory synthesis. Use only when the user explicitly asks to compare/evaluate instruction variants or an approved instruction program names an evaluation boundary.

AI & Automation 1,649 stars 144 forks Updated today MIT

Install

View on GitHub

Quality Score: 92/100

Stars 20%
100
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
50
License 10%
100
Description 5%
100

Skill Content

# Happier Instruction Evaluation Evaluate whether instruction wording changes agent behavior without telling the agents they are being evaluated. This workflow is advisory: it produces evidence for a human decision and never edits the canonical instruction owner by itself. ## 1. Establish the decision Name: - the instruction owner and variants being compared, including the current baseline; - the concrete behavior the change is meant to improve or failure it is meant to prevent; - one organic task or a small risk-selected set of tasks that can expose that difference; - a rubric of observable outcomes fixed before any run; - the user decision the evidence will inform. Do not evaluate prose elegance in isolation. A useful task forces the instruction to affect routing, investigation, ownership, implementation shape, validation, stopping, or reporting. Skip evaluation when a source inspection or deterministic check can decide the question directly. ## 2. Control the comparison Hold constant everything except the instruction variant when practical: - use the same task prompt, repository basis, allowed tools, permissions, time/effort budget, and available evidence; - give each run only the ordinary context an agent would receive for that task; - label variants and output locations neutrally so neither runner nor judge sees “baseline,” “preferred,” model identity, or another run's existence; - isolate writes in separate temporary directories or authorized worktrees; never sw...

Details

Author
happier-dev
Repository
happier-dev/happier
Created
8 months ago
Last Updated
today
Language
TypeScript
License
MIT

Similar Skills

Semantically similar based on skill content — not just same category