agent-evaluation-framework-builder
SolidDesigns an eval suite for an LLM agent or pipeline including success metrics, trajectory scoring, LLM-as-judge setup, and regression test cases.
Install
Quality Score: 83/100
Skill Content
Details
- Author
- Notysoty
- Repository
- Notysoty/openagentskills
- Created
- 6 months ago
- Last Updated
- 3 weeks ago
- Language
- JavaScript
- License
- MIT
Similar Skills
Semantically similar based on skill content — not just same category
agent-evaluation
Measures a skill's practical value with fixed business tasks, paired with-skill and without-skill trials, explicit scoring, and honest uncertainty and cost reporting. Use when the user says "does this skill help", "benchmark our agent workflow", "compare this skill to the baseline", or "prove the new workflow works".
agent-evaluation-rubrics
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and outcome measurement for agent pipelines.
evaluation
Build evaluation frameworks for agent systems. Use when testing agent performance, validating context engineering choices, or measuring improvements over time.