← ClaudeAtlas

prompt-evaluation-harnesslisted

Use when building a prompt evaluation harness — LLM-as-judge, G-Eval, faithfulness and answer-relevancy scoring, hallucination rate, and CI/CD regression suites. Triggers on "eval harness", "LLM-as-judge", "DeepEval", "RAGAS", "faithfulness", "hallucination rate", "G-Eval".
noctua84/nescio-ai · ★ 0 · AI & Automation · score 73
Install: claude install-skill noctua84/nescio-ai
# Prompt Evaluation Harness ## Purpose Create a prompt evaluation harness that delivers actionable, measurable results. **Category**: AI & Automation ## Inputs ### Required - **Objective**: What you want to achieve with this deliverable - **Context**: Relevant background information ### Optional - **Constraints**: Any limitations or requirements to consider - **Existing Work**: Previous documents or data to build on ## Context Before starting, read the repo's `CLAUDE.md` and any relevant notes under `memory/` (e.g. `memory/repo/<repo>/`, `memory/feedback/`) for prior decisions and constraints. ## Process ### Step 1: Context & Research - Review any existing prompt evaluation harness documents in the project - Identify key stakeholders and their requirements - Select the most appropriate framework: DeepEval (50+ metrics), RAGAS, HELM (Stanford) ### Step 2: Analysis & Framework Application - Apply the selected framework to structure the prompt evaluation harness - Identify gaps, opportunities, and risks - Define success metrics: Faithfulness Score, Answer Relevancy, Hallucination Rate, Contextual Precision/Recall - Document assumptions and dependencies - Validate approach against industry best practices ### Step 3: Build the Deliverable - Structure the prompt evaluation harness using the output format below - Include specific, actionable recommendations — not generic advice - Add concrete numbers, timelines, and benchmarks where applicable - Cross-reference with exis