evaluation
SolidThis skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and outcome measurement for agent pipelines.
Install
Quality Score: 81/100
Skill Content
Details
- Author
- docxology
- Repository
- docxology/template
- Created
- 11 months ago
- Last Updated
- today
- Language
- Python
- License
- Apache-2.0
Integrates with
Similar Skills
Semantically similar based on skill content — not just same category
advanced-evaluation
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated quality assessment.
evaluating-skills
Use when testing whether a new skill improves agent behavior, or when validating a change to an existing skill's language.
agent-ux-eval
Evaluate any AI agent experience against the Agent UX Evaluation Framework. Use this skill whenever asked to evaluate, assess, audit, score, or review an AI agent, chatbot, copilot, or agentic workflow. Also trigger when the user mentions "agent evaluation framework", "experience evaluation", "agent UX audit", or "score this agent". Supports screenshot-based, Chrome plugin, and URL-fetch evaluation methods.