agentic-driving-weaker-modelslisted
Install: claude install-skill gidde032/agentic-workflow
# Driving weaker models — the playbook
**The central claim:** a weaker model inside a strong verification harness
approximates a stronger system. Cheaper-tier dry runs passed inside the harness
— skills, gates, review, paired tests — while in the same period the strongest
available model authored three plans that cold review raised 32 findings
against, 4 of them CRITICAL.
The harness is not a patch for weak models; it is the system, and every tier
needs it. Tier buys fewer errors, not zero. What changes by tier is how much
structure the brief must carry and where the human checkpoint goes.
Observations, evidence base, and provenance: `examples.md`.
---
## 1. The gap taxonomy — where cheaper sessions actually differ
Observed, not assumed. Each gap: the evidence, then what it predicts.
**G1 — Trusted docs are treated as ground truth; the verify instinct doesn't
fire.** An agent repeated a stale skill claim verbatim rather than spending one
`git ls-files` to check it, while otherwise executing doctrine perfectly.
Predicts: any error in your skills, briefs, or specs propagates straight into
output. Corollary: skill and spec correctness is load-bearing — review and
dry-run-validate them. Seven skill patches came out of four dry runs.
**G2 — Fluent, unresolved self-contradiction in analysis prose.** A debugging
run delivered a correct fix wrapped in a diagnosis that claimed an impossible
thing, noticed the tension mid-paragraph, and moved on without resolving it.
Predicts: