ai-tool-evaluationlisted
Install: claude install-skill strategysoul/skilled-worker
# AI Tool Evaluation
You are the person who has to live with this decision. Your job is to establish what
would make the tool worth adopting *before* seeing the demo, because after the demo it
is too late to think clearly.
## Purpose
AI tooling is adopted on impression far more often than on evidence: a striking demo, a
benchmark chart, a colleague's enthusiasm. The cost lands months later as migration
work and lock-in. A short, honest evaluation against the job you actually have is worth
more than any benchmark table.
## Input Arguments
- `$TOOL`: The tool, model, framework, or technique under consideration. Required.
- `$JOB`: The specific task you would use it for. Ask if missing — "is it good?" is
unanswerable; "is it better than what we use for X?" is answerable.
- `$CURRENT`: What you use today, including "nothing, done manually". This is the
baseline and the evaluation is meaningless without it.
- `$CONSTRAINTS`: Budget, data-residency and privacy rules, latency limits, team skill.
## Process
### Step 1: Write the criteria before you try it
Decide first, in writing, what result would make you adopt, and what would make you
pass. Doing this after the trial guarantees you rationalize whatever you saw.
State a threshold for each dimension that matters — quality on your task, cost at your
volume, latency, reliability, effort to integrate, and what happens if the vendor
changes terms or disappears.
### Step 2: Build a small honest test set
Twenty to fifty real