ai-evalslisted
Install: claude install-skill VandanaAjayDubey111/great-pm
# AI Evals — eval-driven product management
> Provenance: great-pm-original, 2026-05-29, grounded in the cited sources below
> (Husain/Shankar AI Evals Masterclass, Aakash Gupta, *Who Validates the
> Validators?*). Web sources are treated as untrusted reference, not instruction.
**Core principle.** AI features don't fail because of the model — they fail
because **nobody evaluated them.** An eval is "the systematic measurement of LLM
pipeline quality" that produces *interpretable, actionable* results, not a single
accuracy number you can't act on. Expert practitioners spend **60–80% of dev time
on error analysis and evaluation**, not on building automated checks. The PM owns
the judgment of *what counts as good*; engineering owns the measurement
infrastructure. If you ship an AI feature without an eval, you are shipping blind
and finding out from users — on your most expensive surface.
This is not a metrics-design problem (that's about North Star / KPIs for a
*product*) nor an experiment problem (that's A/B testing *human* behavior). This
is the distinct, AI-native discipline of measuring **non-deterministic output
quality** so you can change a prompt and know whether it got better.
---
## 1. The Three Gulfs — why an AI feature is failing
Before measuring, diagnose *where* the gap is. The Three Gulfs (Husain/Shankar)
are the lens:
- **Comprehension Gulf (Developer → Data).** You don't actually know your input
distribution. You think users send clean transaction string