eval-agent
FeaturedRun evaluation tests against an agent to assess quality and archetype resistance
Install
Quality Score: 91/100
Skill Content
Details
- Author
- jmagly
- Repository
- jmagly/aiwg
- Created
- 1 years ago
- Last Updated
- today
- Language
- TypeScript
- License
- MIT
Integrates with
Similar Skills
Semantically similar based on skill content — not just same category
agent-evaluation
Designs and runs reproducible evaluations for AI agents, prompts, tools, skills, and model-backed workflows using realistic datasets, isolated baselines, objective assertions, rubric grading, trajectory analysis, cost/latency tracking, and regression comparison. Use when measuring agent quality, optimizing skill triggering, comparing prompts or models, or gating an AI feature release. Not for ordinary deterministic unit tests.
agent-quality
Use when evaluating a coding-agent product, gating a release, or when the user mentions evals, Agent Quality, 评测, 模块测试, 整体测试, trajectory, LLM judge, regression fixture, or independent verification of agent behavior. Use after an implementer claims done. Not for ordinary app unit tests with no agent loop.
agent-evaluate
Define behavioral contracts, run adversarial tests, and detect regressions for AI agents — invariants, edge cases, statistical analysis, and benchmark-production gap detection