← All creators

ityaadiii

User

Nine Claude Skills for evaluating AI systems and the skills you install to do it. Every one has a stated refusal.

9 indexed · 0 Featured · 0 stars · avg score 70
Prolific

Categories

Indexed Skills (9)

AI & Automation Listed

automation-break-even

Convert model accuracy into an honest estimate of work saved, cost saved, or headcount impact. Use for AI business cases, ROI estimates, automation proposals, or any claim of the form "the model is X% accurate so it saves X% of the work". Measures the real review costs first and refuses to multiply.

0 Updated 1 weeks ago
ityaadiii
AI & Automation Listed

baseline-before-model

Force a non-model baseline to be built and measured before any LLM is proposed for a task. Use when scoping an AI feature, choosing a model for extraction, lookup, classification or routing, or when someone asks which model to use for a job. Refuses to recommend a model until the boring alternative has a number.

0 Updated 1 weeks ago
ityaadiii
AI & Automation Listed

eval-set-builder

Build an evaluation set from real traffic and check whether it is large enough to support the error budget being claimed. Use when someone wants to measure model accuracy, certify an error rate, decide if a model is safe to deploy, or asks "how many test cases do we need". Refuses to certify a budget the sample cannot support.

0 Updated 1 weeks ago
ityaadiii
AI & Automation Listed

llm-judge-auditor

Audit an LLM-as-judge setup for self-preference, position bias, and lack of human agreement before its scores are trusted. Use when a model grades model output, for pairwise preference evaluations, automated scoring, or when someone reports win rates from an AI judge. Refuses to accept judge scores with no measured human agreement.

0 Updated 1 weeks ago
ityaadiii
AI & Automation Listed

model-bakeoff

Compare two or more models on the same task and decide whether any observed difference is real. Use when choosing between models or providers, running a bakeoff or shootout, checking if a new model beats the current one, or interpreting benchmark results. Uses paired tests and corrects for multiple comparisons. Refuses to name a winner on noise.

0 Updated 1 weeks ago
ityaadiii
AI & Automation Listed

prompt-and-grader-audit

Read an eval's prompt and its grader together as a single artefact and find where they contradict each other. Use when eval scores look wrong, a model is being punished for correct behaviour, scores are suspiciously low, or before trusting any eval result. Refuses to accept a score where the prompt and grader disagree.

0 Updated 1 weeks ago
ityaadiii
AI & Automation Listed

skill-actually-works

Measure whether an installed skill changes the output at all, by running the same tasks with it enabled and disabled. Use when auditing an installed skill library, deciding which skills to keep, or when someone claims a skill improved results. Refuses to credit a skill with an effect it cannot demonstrate.

0 Updated 1 weeks ago
ityaadiii
AI & Automation Listed

skill-degradation-check

Test whether a skill invents data when its inputs, tools or API connections are missing, instead of reporting the gap. Use before trusting a skill on client or production work, when a skill produces suspiciously complete output, or when auditing a skill library for honesty. Refuses to pass a skill that fills gaps silently.

0 Updated 1 weeks ago
ityaadiii
Code & Development Listed

skill-safety-review

Statically review a SKILL.md or agent skill before installing it, flagging credential access, exfiltration, instruction override, and review-evasion patterns. Use before installing any third-party skill, plugin or agent extension, or when asked whether a skill is safe. Read-only. Never executes what it scans.

0 Updated 1 weeks ago
ityaadiii

Bio shown is the top-scored skill's repo description as a fallback — real GitHub bios land in a future update.