← All creators

tangxiangru

User

AI handles execution, humans own the direction, and every run becomes an inspectable research artifact on disk.

87 indexed · 0 Featured · 805 stars · avg score 81
Prolific

Categories

Indexed Skills (87)

AI & Automation Solid

astronomy-sample-the-published-table-into-chains

Use at study design and implementation when the source's constraints reach you as a table of best-fit values with 1-sigma errors for two or more models and no posterior samples were released. Covers rebuilding the ensemble that table describes, writing it in the layout this field's posterior tools read, and what to state about the construction so it is evidence rather than decoration.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

astronomy-the-caption-is-the-figure-specification

Use at literature survey the moment you have the source's full text, and again when the plotting code is written, whenever you are reproducing a figure whose rendering you cannot see. Covers mining each numbered caption for the panel order, the series and their colours, the reference model the residuals are taken against and how the data were normalised, and holding those fixed against a later stage that finds a better choice.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

astronomy-the-joint-posterior-is-the-parameter-result

Use at study design when the figure slate is fixed, and again at analysis and writing, when the deliverable is constraints or posterior distributions on parameters for two or more competing models. Covers why one overlaid triangle plot over the source's own parameter list is the exhibit that answers it, which parameters get an axis, and why a row of one-dimensional error bars reads as the figure never having been drawn.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

chemistry-a-cut-variant-takes-its-analyses-with-it

Use at literature survey, at study design, at every descope decision and again at writing, when the source's method is a family — the same module dropped into two or more backbones, or one architecture published in several named variants — and you are about to run only one of them. Covers listing which of the source's downstream analyses were produced from which variant before any of them is cut, shrinking a variant rather than deleting it, and what a saliency map, case study or ablation computed on the surviving variant is and is not evidence for.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

chemistry-accuracy-and-cost-for-every-module-you-swap-in

Use at study design, through experimentation and again at analysis when the method under test is a drop-in replacement for a standard layer — a different basis, kernel, activation family or transform — and the source claims the replacement is both more accurate and cheaper. Covers giving every alternative module a cell in the accuracy column and in the cost column, fixing one matching convention across both, and dividing the runtime by the invariant already sitting in your own results file before you publish a contradiction of the source's ratio.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

chemistry-reproduce-the-scoring-path-before-you-replace-it

Use at implementation, experimentation and analysis when you are reproducing a published benchmark number and the source's scoring path is one you can read — which rows are scored, in what order, how many the loader drops, which epoch is reported, how tasks are pooled, over how many seeds. Covers implementing that path exactly before improving it, the one-row-per-step ladder from the published rule down to your own honest estimate, and why one un-replicated step makes the reproduction gap you report uninterpretable.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

decide-the-input-or-the-deadline-decides-it

Use at study design, and again at every stage boundary after it, when the run's own notes still carry an open question about which file or which system one of the named experiments will run on - the shipped stand-in, the authors' release, or one you generate from the Methods. Covers writing the default outcome beside every open question, ranking the list by that default rather than by difficulty, and the three-route ladder for a system the task did not ship.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-value-you-did-not-measure-still-has-a-source

Use at Stage 06 and Stage 07 when a deliverable the task named cannot be produced by this run at all — a wet-lab measurement, a synthesised material, a proprietary benchmark, hardware you do not have. Covers the difference between fabricating a number and citing one, where the cited value belongs, and why omitting the section is the worst of the three options.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

the-attribution-is-the-deliverable

Use when the task statement names interpretability, explainability, feature importance, saliency or attribution among its outputs or objectives. The graded artifact is then the attribution map itself — per input unit, by the field's standard estimator, drawn as a figure — not a diagnostic about the model's internals and not an argument that the model is uninterpretable.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

train-the-named-architecture

Use at study design, implementation and experimentation when the brief's deliverable is a model you have to build — it names an architecture family (graph network, autoencoder, diffusion module, surrogate net) or a training regime (pre-training, fine-tuning, self-supervised, inverse design). Covers why a cheaper model class scores near zero however well it performs, why a scaled-down run of the named architecture beats a released checkpoint on every architecture criterion, and what to ablate.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

astronomy-error-budget-is-the-audit-trail

Use at analysis and writing when a result rests on a fit or a calibration chain and you are about to quote it with a single uncertainty. Covers itemising the error budget term by term, keeping the fit's own bookkeeping visible, and why the audit trail is the result a referee checks first.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

astronomy-figure-is-the-unit-of-result

Use at study design when choosing the figure list, and again before writing, when a result is about to be reported as a pooled number or a table. Covers why the figure is the unit a result is delivered in here, which panels a paper of this kind is expected to carry, and what a pooled number hides.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

chemistry-canonical-units-thresholds-incumbent

Use at study design and analysis when a chemistry result is about to be reported in the units your code happens to produce, or without the program the field already uses. Covers anchoring to the incumbent, converting to the canonical unit, and turning an error distribution into a threshold success rate.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

chemistry-ranked-entities-and-property-curves

Use at analysis and figure planning when the computation ranks entities — molecules, poses, fragments, atoms — or sweeps a property along a coordinate. Covers printing the named ranked list and the property-versus-coordinate curve, the two artifacts most often computed here and least often reported.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

citation-discipline

Use when adding, verifying or cleaning citations and BibTeX entries in Stage 07 (Writing), when a reference cannot be resolved cleanly from DBLP or CrossRef, when checking that a cited paper actually supports the claim attributed to it, or when filling citation_verification.json.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

close-the-gap-to-the-published-number

Use at Stage 05 and Stage 06 the moment a reproduction lands materially off a number the source study published — a different order of magnitude, an inverted trend, a collapsed estimate. Covers why the gap is a defect in your pipeline until you have shown otherwise, how much of the remaining budget to spend closing it, and what to write when it will not close.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

draw-the-source-figure-panel-for-panel

Use at study design when planning figures for a reproduction, replication or validation task, and again before the report is written. Covers deriving each panel's series list and axis ranges from the source's rendered figure, giving every source result a panel before your own hypotheses claim the slots, and printing the source's named constants as labelled values.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

earth-comparator-set-lives-outside-the-supplied-archive

Use at study design when the supplied archive is about to become both the input and the thing you compare against. Covers where the comparison set has to come from, what a self-comparison cannot establish, and how to build a comparator from outside the shipped data.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

earth-report-the-lattice-and-show-the-field

Use at analysis and figure planning when a geospatial or gridded result is about to be reported only as regional aggregates. Covers reporting the stratified lattice, showing the field the strata came from, and which map a study of this kind is expected to publish.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-model-you-can-audit-is-not-a-model-that-scores

Use at study design and implementation when choosing between a method you can validate quickly and a stronger one you are not sure you can afford. Covers pricing the expensive method with a measurement instead of an impression, the go/no-go that has to be written before the clock is spent, and why the safe choice is only safe on the axes nobody is grading.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-scoreable-file-in-the-first-hour

Use at the first stage of a run whose deliverable is a predictions file, and again at every stage when one still does not exist. Covers why a trivial submission written early dominates a good one written late, what the first version should contain, and how to improve it in place without ever leaving it invalid.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

assume-this-stage-is-the-last-one-you-get

Use at the first three stages of a run under a hard wall clock, when the plan defers the modelling to a later stage. Covers the measured probability that the later stages never execute, why deferring to the stage designed for the work is the most expensive available choice, and what each early stage should leave behind if it turns out to be the last one to run.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

draw-the-system-not-your-study

Use at study design when allocating figure slots, and again at analysis and writing, on tasks where the source's own rendered figures are not available to copy. Covers the four slots reserved for the system before any hypothesis claims one, drawing the loaded arrays instead of their counts, why a panel that reports a shortfall is not the panel carrying the result, and why a deliverable marked covered_by an artifact path is not covered.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

earth-a-verified-answer-key-does-not-change-the-question

Use at literature survey when you find that the supplied archive holds the source study's published output as well as its raw inputs, and again at hypothesis generation before anything is frozen. Covers what to do with the source's headline numbers in the hour you first recompute them, why a confirmed answer pulls a run into auditing the method that produced it, and the ordering rule that keeps the critique behind the delivered product.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

earth-shape-of-the-record-and-share-of-the-budget

Use at analysis and again at writing when the deliverable is a multi-year record - an annual time series, a reconstruction, a trajectory - and you are about to report it as a period mean, a cumulative total and a validation residual against a reference version of itself. Covers the change over the record's own length, the extremes, the per-unit split, and the record's share of the budget it is one term of.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

earth-the-technique-grid-needs-values-in-its-cells

Use at study design when the figure list is chosen, at implementation before any pipeline code is written, and again at analysis, when the supplied archive is stratified by measurement technique or instrument and you are about to describe that stratification. Covers reading the per-technique columns the archive already ships, drawing the technique grid with estimates in it rather than file counts, and why a no-peek rule about the reference product must not reach your figures.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

earth-two-orderings-of-a-regional-decomposition

Use at study design when the figure and table plan is fixed, at analysis, and again at writing, whenever the deliverable splits a global or basin-wide total into per-region parts and you are about to report each part as an absolute rate. Covers the two normalisations every row owes and why their disagreement is the result, where the intensity denominator has to be captured before you need it, and what the spare cell of a small-multiple grid is for.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

earth-window-mismatch-is-an-alignment-problem

Use at study design, and again when planning figures, whenever a comparator you hold - a model projection ensemble, a scenario run, a prior published assessment, a sibling record - is reported over a different period, baseline epoch, initial state or unit than your result, and you are deciding whether the comparison can be made at all. Covers re-baselining onto a common start date, plotting an ensemble that publishes only horizon endpoints, reading the crossing date, expressing prior assessments as revisions, and where a genuine refusal belongs.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-deliverable-is-not-an-instruction

Use at study design when listing the task's deliverables into report_plan.json, and again before writing when checking coverage. Covers how to tell a research deliverable from the harness's own operating instructions, why a padded list is worse than a short one, and what to write when a deliverable is genuinely out of reach.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-detection-score-is-a-claim-about-its-population

Use at study design and at analysis whenever a detection or ranking score (area under a precision-recall or ROC curve, recall at a fixed precision) is about to be compared against another study's number, or when one arm detects items the other misses. Covers publishing the population beside the score, the prevalence ladder to run when the source never states its own, and the characterisation of the extra detections that needs no annotation.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-null-test-bounds-the-instrument-not-the-answer

Use at analysis and again at writing whenever you run a permutation, shuffle, placebo, unforced-control or power test against your own headline result, especially when it comes back saying the result is not distinguishable from noise. Covers giving every condition the task names its own value line and stating the relation across them as a result, keeping the estimate and the bound as two results with two different subjects, and the sentence order that stops a bound replacing the answer.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-refutation-banner-over-a-confirming-panel

Use at hypothesis freeze, and again after the last revision pass, when a run has found that the supplied data or your own reproduction disagrees with the source it names. Covers the branch-name test that stops an agreement being printed as a refutation, and the enumerate-and-search sweep that keeps fidelity verdicts and internal labels out of the title, the headings and the figure banners.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-supplied-parameter-file-is-a-list-of-questions

Use at study design, analysis and writing when the task ships a small file of named constants, ranges, entity tables, case lists or run settings. Covers treating each entry as a question your run must answer in the file's own labels and units — including entries your own audit shows are wrong — and why an agreement count is not an answer.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

chemistry-ablations-and-curves-without-an-accelerator

Use at study design after you have priced a scaled-down training arm and found the machine cannot carry it — no accelerator visible, or no wall clock for one arm. Covers the one-row-per-named-component table with the inference switch that removes each part, why an input ablation does not answer a component criterion, and the ladder of curves that still ships when nothing can be trained.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

chemistry-fill-every-row-of-the-comparator-table

Use at literature stage and study design when the source publishes performance broken out by class of system, and you are about to choose your evaluation panel with a filter written for throughput. Covers transcribing the table as rows, auditing the inclusion filter against those rows before it is frozen, and buying one target per row before a second target for any row.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

chemistry-group-attribution-over-the-split-and-the-baseline-mode

Use at study design, experimentation and analysis when the deliverable includes which substructures, functional groups or motifs drive the model's predictions, once the attribution estimator is already chosen. Covers widening from the one molecule the source drew to the whole evaluation split with per-molecule normalisation, running the identical attribution on the comparator model so a claim of better interpretability becomes measurable, and treating a learned edge or subgraph mask as a first-class output.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

chemistry-interaction-inventory-of-the-modelled-complex

Use at study design, analysis and writing when the result is a modelled or predicted molecular complex — a docked pose, a co-folded assembly, a binding interface — and RMSD, DockQ or lDDT is about to be the whole answer. Covers the reference-versus-prediction contact inventory per interaction class, the pocket-cropped figure with the interactions drawn, and the mechanism sentence.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

claims-before-harness-forensics

Use at hypothesis generation and study design on reproduction and method-evaluation tasks, once close reading of the release has turned up defects, ambiguities or under-specification, and again when ordering the report. Covers labelling every planned experiment as a test of a claim or a test of self-consistency, the count gate that follows, and where reproduction-fidelity statistics belong.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

disclose-by-construction-not-by-absence

Use at analysis when figures are rendered and at writing when they are captioned, and whenever an internal review asks you to disclose something you could not do. Covers why a disclaimer drawn inside a figure's axes destroys the result it annotates, the single location a caveat is stated in and what counts as a second copy, and how to describe the substitute you built instead of the gap you had.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

do-not-grade-your-own-result-down

Use when drafting limitations, the discussion or the abstract, and any time you are about to call your own result unimproved, inconclusive or unverifiable. Covers the hedge that contradicts the run's own decision record, and the check a caveat has to fail before it is published.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-combination-is-not-the-candidate-set

Use once more than one trained artifact exists on disk -- two checkpoints, two seeds, two architectures, a continuation run -- and something is deciding which of them, or which combination of them, writes the predictions file. Covers the ballot that lists every artifact as a submission on its own before any blend, re-running it whenever a training job finishes, the known-bad canary that tests the objective, and persisting a rejected candidate's predictions.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-comparison-you-never-run-defaults-to-your-preference

Use at hypothesis drafting, study design and implementation when the task could plausibly be attacked by more than one family of method - hand-built features fed to a fitted model, a network trained on the raw structure or sequence, a pretrained backbone, retrieval - and your hypothesis list mostly compares variants inside one of them. Covers separating the hypotheses that would change what you build from the ones that would change an argument, requiring code on both sides of a family claim, running the comparison at a budget you actually have and reading each side's slope rather than its level, and demoting a comparison you will not run into a priced assumption.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-cut-point-is-a-fitted-parameter-not-a-setting

Use whenever a column of your submission is decided by comparing a continuous score against a number you chose -- whether to commit an answer or declare the row unanswerable, whether to flag a borderline case, which output to emit when the model is unsure. Covers the two questions that number silently answers, why it gets fitted on the smallest labelled sample in the run and then applied to the largest split, the statistic it should have been swept on, and the commit-rate print-out that tells you it is on the wrong side of the tail.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-priced-bias-is-a-work-item-not-a-caveat

Use at implementation, experimentation and analysis on a task graded by an error metric over a predictions file, when a diagnostic you ran after freezing your design says the numbers you are about to ship are biased - too high, too low, on the wrong scale, in the wrong units - and a pre-registered decision rule is the reason you are recording it rather than fixing it. Covers the one test that separates a forbidden search over candidates from an ordinary bias correction, where the correction may be estimated, the held-out check that has to pass before you apply it, and how to ship the corrected file while still reporting the frozen verdict.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-second-model-family-before-a-fifth-sample

Use at literature survey, study design, implementation and experimentation when the predictions come from running a pretrained checkpoint you picked off the shelf over each row — a language model that reads the text and answers, an encoder, any released artifact — rather than from fitting a model on the training rows, and especially when the next thing you planned is another sample, seed, temperature or voter from the checkpoint you already downloaded. Covers treating the set of checkpoints as an experimental axis with a deadline of its own, the best-single / oracle / best-vote measurement on your own labelled rows that decides whether to buy a better aggregator or a different model, and why a ceiling computed from your own predictions bounds your shortlist rather than the task.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-serial-queue-is-a-measurement-you-did-not-take

Use at study design when the run queue is priced, at implementation when the first launcher script is written, and at every stage boundary while fits, simulations or samples are still running, especially when the brief states a CPU allocation and your queue happens to run one job at a time. Covers why running nproc is not the missing step, a stated allocation versus an enforced cgroup limit, the two-job A/B and the width ladder that settle how many concurrent processes to run, processes versus per-process threads, and re-pricing every arm you declined while the node sat idle.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-throughput-number-belongs-to-the-runtime-not-the-model

Use at the survey, study design and implementation stages when the task needs a pretrained model, solver or library you must download and run on the machine you were given, and the first configuration you try is too slow to cover the split in the time you have. Covers why the seconds-per-item you just measured is a property of the runtime, the workload and the machine as much as of the component, which field to change before demoting it, re-asking the component question after you fix the runtime, and checking that your fallback still has the property you picked the original for.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

add-the-baseline-back-on-the-split-you-cannot-score

Use at study design, implementation and every write thereafter, whenever the value you submit is assembled from parts - a fitted baseline plus a model's residual, a level plus a shape, a de-trended prediction that has to be re-trended, any inverse transform - and the validation arrays and the graded arrays are produced by separate calls. Covers assembling every split through one function, using the baseline you already fitted as a label-free reference vector on the graded split, why row count, header, dtype and finiteness cannot see this class of error, and putting the gate inside the writer rather than in a script somebody has to remember to run.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

an-oracle-ceiling-is-not-headroom-yet

Use at hypothesis generation, study design and implementation once a submission exists and you are choosing where the remaining hours go -- in particular when a lever looks worth building because you worked out what it would pay if its setting were chosen perfectly for every row, or when the distance between your metric value and a published number for this dataset and metric looks like room to improve. Covers computing that best-possible value as an explicit ceiling, inverting it into the correlation a real predictor would need before it clears your ship bar, measuring the correlation your inputs actually carry, and checking that a published number you are chasing was computed the way your submission will be scored.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

calibrate-the-level-on-the-window-you-cannot-score

Use at implementation and afterwards whenever the rows you will be scored on lie outside every window you can check against truth — a forecast horizon that starts where the supplied history ends, a later time period, a different site, batch or cohort, a test split whose label column has been removed — and your only bias check was run on a backtest fold or a random validation split. Covers why "my predictions are unbiased" is a statement about the folds and not about the graded rows, how to measure the overall level of your predictions on the graded rows with no labels at all, why the level ratio is only the alarm and a metric scan is the number, what to do when the two windows disagree, and when a low forecast is correct rather than broken.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

close-the-review-item-on-its-object-not-its-verb

Use whenever a reviewer, critic or supervisor returns numbered suggestions and you are about to record which were executed, deferred, declined or replaced, and whenever you mark one of your own registered hypotheses or planned experiments as done. Covers the substitution that makes a false close-out read as true - same verb, cheaper object, a tenth of the scale - the four close-out fields that make it impossible to perform by accident, and the two searches that show whether the mechanism you claimed is in code and in your runs' recorded configuration, or only in prose.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

combine-is-the-third-outcome-of-an-adoption-test

Use at hypothesis generation and study design when you are about to write a decision rule of the form "adopt the new method only if it beats the current one on held-out data", and at implementation and experimentation when a second scorer, prompt, view, feature set or model has just come in below the one you are already shipping. Covers registering combine as a third outcome beside replace and discard, the two counts that say whether a losing method still holds information, blending scores instead of decisions, and the pre-registration and nested cross-validation that stop a blend search from inventing its own lift.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

copy-the-graders-limits-not-just-its-logic

Use at study design, implementation and experimentation when the task's score is produced by *executing* what you submit - running generated programs against hidden test cases, simulating, decoding, solving, rendering - and where you are building a local copy of the scorer to choose among candidates or to measure a method before committing to it. Covers reading the resource limits out of the shipped scorer's source rather than out of the prose that summarises it, giving every limit a named constant with its source line beside it, why a replica looser than the grader is far worse than one that is stricter, the failure taxonomy your local report must be able to express (a bucket that is empty in every arm is the tell), and the item-by-item calibration that catches a replica whose average already agrees.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

cost-the-rung-you-need-not-the-cheapest-one-in-the-family

Use at literature survey, hypothesis generation and study design of a task that hands you a training split and an unlabelled test split, scores predictions by a fixed error metric, and has published best numbers for that dataset and metric, when you are about to decide that the kind of model behind those numbers does not fit your clock. Covers which rung of the ladder is worth timing at all, replacing "hours to finish the reference schedule" with "epochs until this passes what I already have", and pricing the first member of another family before the Nth member of this one.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

count-exact-rows-before-you-fit-a-correction

Use at implementation and experimentation when the target may be a deterministic function of the inputs - a computed score, a derived column, a simulator or rule-based output - your reconstruction of it is close but not equal, and the next thing you planned was to train a model on the difference. Covers the fraction-reproduced-exactly measure that mean error hides, how to choose its tolerance from the residuals instead of by taste, how to read a residual that takes only a few distinct values, and the gate a learned correction must clear before it goes on top of an analytic base.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

count-your-hypotheses-not-your-arms

Use at study design, implementation, experimentation and analysis when the run has a hard wall clock, models that take hours to train, and more than one configuration you would like to try in parallel - several seeds, several widths, a warm restart, a variant of the variant. Covers reading the machine you actually have instead of the thread count you typed, measuring the contention tax rather than assuming it, the distinct-hypothesis count that decides whether a launch buys anything, computing `converged` from the log instead of declaring it, reporting a blend of arms the clock cut as the repair it is, and the idle-cores failure at the other end of the run.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

the-unit-of-analysis

Use at analysis and figure planning when the brief names the units its data is grouped into — patients, cells, classes, labs, behaviours — and you are about to report one pooled number over all of them. Covers why the pooled number hides the result, which strata a study of this kind is expected to report, and when an aggregate is the right answer after all.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

answer-the-why-not-only-the-what

Use when writing results and discussion, and when a task or a reviewer asks why an effect happens rather than whether it does. Covers the difference between reporting an effect and accounting for it, and what a mechanism claim needs behind it.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

cover-what-the-task-named

Use at study design and again before writing, to check that every deliverable the task statement names has been produced. Covers how to enumerate what was asked for, why partial coverage scores worse than it feels, and what to do when a named deliverable is out of reach.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

information-exhibit-the-intermediate-objects

Use at analysis and writing when a multi-stage pipeline is about to be reported by its end-to-end metric alone. Covers exhibiting each stage's intermediate object, and re-running the source's own demonstrations on the source's own inputs rather than on yours.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

information-fill-the-whole-results-grid

Use at study design and again at writing when the source reports a grid — variants crossed with backbones, datasets or metrics — and you are about to fill part of it. Covers reproducing the whole grid at reduced N where you must, and why a labelled reduced-N cell beats an empty one.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

latex-repair

Use when a LaTeX build fails or produces a broken PDF in Stage 07 (Writing) — undefined control sequences, missing style packages, unresolved citations or references, float placement blowing the page budget, or a build_log.txt full of errors you need to triage.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

life-benchmark-against-the-incumbent

Use at study design when a life-science method result is about to be reported on its own numbers. Covers the head-to-head against the incumbent tool, the cost table that goes with it, and finding an orthogonal truth set the method was not fitted to.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

life-full-study-skeleton-including-the-wet-lab-half

Use at study design and again when laying out the results section, to check every slot of a life-science study is filled. Covers the skeleton a paper of this kind carries, and what to put in the slots this run cannot compute rather than leaving them out.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

material-as-specified-run-and-stage-diagnostics

Use at study design and implementation when a protocol is specified and you have found a reason to deviate, or when a pipeline stage is about to run without its conventional diagnostic. Covers running the protocol as specified as the foreground result, and leaving every stage's default panel behind you.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

material-landmark-scalars-in-physical-units

Use at analysis when a materials result exists as a curve, a distribution or a trajectory and is about to be reported as one. Covers extracting the landmark scalar a reader compares — peak position, transition temperature, barrier height — in the property's physical unit, against a reference value.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

math-canonical-curve-on-the-cost-counter

Use at figure planning when a convergence or performance curve is about to be drawn against wall-clock, or folded into a composite panel. Covers the field's plain two-curve figure, plotting against the algorithm's own cost counter, and why it comes before any richer diagnostic.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

math-equal-effort-baselines-and-knob-sweeps

Use at study design when the source names competing algorithms and they are about to become a related-work paragraph instead of arms. Covers running every named baseline at equal tuning effort, and sweeping the parameter you claim credit for.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

mine-the-papers-you-were-given

Use when the task ships PDFs in related_work/, at literature stage and before the study plan is costed. Covers reading those papers for the named tools, benchmarks, events and metrics the work will be judged against — as a work list rather than as background — and what to record for each one.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

neuroscience-comparator-ladder-and-per-unit-predictions

Use at study design and analysis when a model is about to be compared against one alternative, or a fit reported without a negative control. Covers the two-sided comparator ladder, the control representation panel, and splitting per-unit predictions into the ones a measurement validates and the ones that stay predictions.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

neuroscience-stratify-and-report-detection-metrics

Use at analysis when a detection or classification result is about to be reported as one accuracy over a pooled population. Covers per-group and per-class precision, recall and confusion matrices at a stated threshold, and sweeping the degradations the recording modality actually suffers.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

paper-writing

Use when drafting, structuring or revising the manuscript or report in Stage 07 (Writing) — shaping the contribution into one story, writing the abstract and introduction, fixing prose that reads generic or templated, ordering sentences for clarity, or deciding what Figure 1 should show.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

physics-discriminate-model-families-and-defend-the-fit

Use at analysis when a fit is about to be reported as the answer without a rival model being excluded. Covers naming the competing model families, showing which the data rules out, and treating the fit protocol — range, weighting, priors — as part of the result rather than as a setting.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

physics-two-estimators-propagation-and-a-forward-model

Use at study design and analysis when a physical quantity is about to be reported from one estimator, or an uncertainty quoted without propagation. Covers measuring it a second independent way, propagating the error through the chain, and generating the observable forward from the fitted model to check it.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

publish-what-the-run-already-computed

Use at Stage 06 and again before the report is finalised, when deciding which of the run's results enter the deliverable. Sweeps the run's own outputs for quantities it computed and never published, and covers the three shapes that sweep finds — the diagnostic never persisted, the column requested and dropped, the feasibility measurement discarded — and what to promote out of an appendix.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

reproducibility-check

Use in Stage 08 (Dissemination) when assembling the release or submission bundle — auditing whether the run's code, data, results and figures are actually reproducible by someone else, writing the readiness checklist and threats-to-validity notes, or deciding what has to be disclosed as not verified.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

result-table

Use when turning measured results into a table or figure for the paper — building a LaTeX or markdown results table from workspace/results/*.json, deciding what uncertainty to report, choosing which baselines and ablations belong in the main table, or writing a caption that stands alone.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

run-the-conditions-the-source-ran

Use at study design, before any experiment of your own is costed, on reproduction and method-evaluation tasks. Covers enumerating the systems, scenarios, stress sweeps and case studies the source names, running each one by name, measuring the preconditions the method declares it needs, and what to do when one of them fails.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

the-canonical-figure

Use when planning figures, at study design and again before writing. Covers the figures a paper in this field is expected to contain, why an original figure does not substitute for a standard one, and how to decide what to draw first.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

the-reproduction-is-a-hypothesis

Use at Stage 02 and Stage 03 whenever the task is to reproduce, re-implement or verify a published study and the hypotheses you are drafting are all about something else. Covers how to write the reproduction itself as a falsifiable frozen commitment, why a self-invented question crowds it out, and how to budget between the two.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

the-supplied-item-is-the-graded-unit

Use at study design whenever the task ships a specific named object in data/ — one paper, one structure, one instance — and again before writing. Covers reporting that item's own numbers under its own name, choosing the worked example by the task's pointer rather than by your result, and how to widen scope without dropping it.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

use-the-sources-own-names

Use at Stage 06 and Stage 07 when writing up a reproduction, and any time you have given a reproduced quantity, equation, figure or sequence a name of your own. Covers why a correct reproduction under private names reads as a missing one, which names have to be carried, and where they have to appear.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

venue-checklist

Use when the target venue's submission requirements matter — checking a draft against NeurIPS, ICML or ICLR expectations, deciding which required sections (checklist, broader impact, reproducibility, LLM disclosure) the paper needs, or running the Stage 08 submission-readiness review.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

evidence-not-assertion

Use whenever a number, a comparison or a claim is about to enter a stage summary or the report — at analysis and writing, and any time you are tempted to state a value you have not computed in this run. Covers where a number must come from, what to do when the experiment did not run, and why an honest gap outscores a plausible sentence.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

record-what-you-learned

Use when a run is finished and the report is written, after the report is written, to record one reusable lesson for the next run in this field. Covers what counts as a lesson worth passing on, what must never be passed on, and how to write it.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

reproduce-then-extend

Use when the task is to reproduce, replicate or re-implement a published study, at design time and when reporting results. Covers what a reproduction must report, how to compare against the source study's numbers, and why the reproduction comes before any improvement.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Listed

run-the-requested-analysis

Use when the supplied data looks synthetic, degraded, incomplete or wrong, and whenever you are tempted to reframe the study around what you found about the inputs, the harness or the evaluation. Covers what to do with a real data problem without losing the study.

805 Updated 2 weeks ago
tangxiangru

Bio shown is the top-scored skill's repo description as a fallback — real GitHub bios land in a future update.