← ClaudeAtlas

experiment-auditlisted

Use this skill for scientific and ML-research reasoning work — evaluating experimental claims, auditing training runs or ablations, checking whether a statistical claim holds up, assessing reproducibility, reconciling contradictory results, reviewing a paper's methodology or results section, writing reviewer-style feedback, or producing a structured research report. This is a scientific reasoning discipline, not just a tool wrapper, so it applies even with no live data source — e.g. reviewing a pasted table of results, sanity-checking a claimed effect size, or evaluating an ablation described in prose. Trigger on phrasing like "did I mess up this experiment," "is this result real," "why did my loss/reward do X," "which run is better," "is this ablation confounded," "review this paper's claims," "write reviewer feedback," "is this reproducible," "what should I conclude from this," or "write up these results." When the user's data lives in Weights & Biases, this skill also covers the experiment-audit-mcp integr
SreeDharshan-GJ/experiment-audit · ★ 4 · AI & Automation · score 75
Install: claude install-skill SreeDharshan-GJ/experiment-audit
# Experiment Audit — Scientific Research Reasoning Engine ## What this is Experiment Audit is a **scientific research reasoning engine**: a discipline for evaluating experimental and empirical claims the way a careful reviewer or co-author would — checking what the evidence actually supports, separating measurement from interpretation, naming uncertainty precisely, and never letting a clean-looking number stand in for a checked one. The **MCP server** (`experiment-audit-mcp`, eight tools over a user's Weights & Biases project), the **CLI**, and the **Python package** are integrations of this engine — they are how it pulls structured evidence out of a live experiment-tracking backend when one is available. They are not what this skill is *for*. A huge share of real requests in this domain — reviewing a paper, sanity-checking a table someone pasted in, writing reviewer feedback, reasoning about an ablation described in prose — involve no MCP call at all, and the same reasoning discipline applies to all of them. Think of it as two layers: 1. **The reasoning engine** (this skill, `prompts.md`, `examples.md`) — how to evaluate evidence, phrase findings, hedge accurately, catch contradictions, and write up conclusions, regardless of where the evidence came from. 2. **The MCP integration** (`reference.md`, the tool table below) — the specific, calibrated tools available when the evidence lives in W&B: what each one computes, its exact thresholds, and its document