← ClaudeAtlas

plumbline-measurement-sliceslisted

Commands and hard-won honesty rules for Plumbline measurement, benchmark, council A/B and Council-GUI slices. Load before running the metrics harness (emit_run, process_health, council_review_scorer, arm_a_review_runner, council_measurement_run, council_free_diversity_probe), before the Council GUI composition root, or before publishing any benchmark/catch-rate claim.
DYAI2025/Plumbline · ★ 6 · AI & Automation · score 63
Install: claude install-skill DYAI2025/Plumbline
# Measurement, benchmark and Council-GUI slices Migrated out of the always-loaded root `CLAUDE.md` on 2026-07-31 so it loads only for measurement/benchmark/GUI work. Same rules, same authority. ## Commands ```bash # Benchmark + measurement harness (config/claude/metrics/ → metrics/): python3 config/claude/metrics/emit_run.py --corpus-id <id> --mode <core|full> \ --metrics '{...}' --gate-outcomes '{...}' --human-overrides 0 # append a run to runs.jsonl python3 config/claude/metrics/process_health.py # SPC + drift attribution python3 config/claude/metrics/challenge_token_oracle.py # deterministic catch oracle python3 config/claude/metrics/council_review_scorer.py # catch / cry-wolf scorer python3 config/claude/metrics/arm_a_review_runner.py # single-model arm (Arm A) python3 config/claude/metrics/council_measurement_run.py # A/B council measurement python3 config/claude/metrics/council_free_diversity_probe.py # free-tier probe (EXP-009) # Council GUI (Slice 4) — the real composition root; fails LOUD on a missing precondition: config/claude/bin/plumbline-council-gui --self-check # wiring proof, crosses no boundary config/claude/bin/plumbline-council-gui # serve (live needs COUNCIL_INFERENCE_LIVE=1) ``` ## Benchmark-claim honesty (learned) When publishing benchmark results (README/docs), a claim must carry its own scope and **both** anti-Goodhart metrics. The v0.10 n=6 slice showed c