agent-evaluation
FeaturedRun and score exactly one Benchmark Case run with CLI execution, Trace provenance checks, and private Rubric isolation.
Install
Quality Score: 93/100
Skill Content
Details
- Author
- Prism-Shadow
- Repository
- Prism-Shadow/penguin-harness
- Created
- 1 weeks ago
- Last Updated
- today
- Language
- TypeScript
- License
- Apache-2.0
Similar Skills
Semantically similar based on skill content — not just same category
benchmark-design
Design and calibrate a multi-Case capability Benchmark with repeated independent evaluations and a traceable baseline.
agent-review-benchmark
Generate evidence-linked Guided Review artifacts and deterministic local Agent benchmark summaries. Use when reviewing Agent-produced diffs by intent, recording structured follow-ups, comparing multiple Agent/model/configuration runs against versioned repository assertions, or preparing reproducible review and benchmark evidence for /ship or Mission validation.
agent-benchmark
Self-benchmark: YOU write the code, adversarial reviews it (multi-provider), you fix, you write tests, adversarial reviews tests, you fix. Measures YOUR quality as an agent. Run in different models (Opus, Sonnet, Haiku) and compare results.