count-your-hypotheses-not-your-arms

Solid

Use at study design, implementation, experimentation and analysis when the run has a hard wall clock, models that take hours to train, and more than one configuration you would like to try in parallel - several seeds, several widths, a warm restart, a variant of the variant. Covers reading the machine you actually have instead of the thread count you typed, measuring the contention tax rather than assuming it, the distinct-hypothesis count that decides whether a launch buys anything, computing `converged` from the log instead of declaring it, reporting a blend of arms the clock cut as the repair it is, and the idle-cores failure at the other end of the run.

AI & Automation 805 stars 25 forks Updated 2 weeks ago NOASSERTION

Install

View on GitHub

Quality Score: 82/100

Stars 20%
97
Recency 20%
90
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# Six arms of one architecture is one experiment, run six times Launching a training run costs one backgrounded command. Killing one costs a decision. So arms accumulate: a wider net, a warm restart, another seed, a variant of the variant. Each launch is defensible on its own and nothing in the run ever subtracts. The usual advice — *parallel arms divide your cores* — is a guess, and on a shared node it is often the wrong guess in the wrong direction. Do not guess. A run with several live arms owes three ledgers, and it is almost always the third one that costs the score. ## Ledger 1 — the machine you have, not the one you typed `torch.set_num_threads(3)` and `OMP_NUM_THREADS=3` are your choices. They are not a hardware fact, and the moment you write one it starts being quoted back to you as a constraint. Read the real numbers, into a file, before the first launch and again at every stage boundary: ```bash nproc cat /sys/fs/cgroup/cpu.max 2>/dev/null # the quota, if there is one cat /proc/loadavg ``` > **A resource number that appears in your report must appear in a file on disk > first. If it does not, you have quoted your own configuration flag as a > measurement.** The tell is mechanical: grep your notes for the core count you cite and see whether any artifact contains it. Measured on the run below, its own monitoring script wrote `{"nproc": 44, "load_average_1min": 22.24}` every fifteen minutes for three hours, while its report said the model was *"trainable on...

Details

Author
tangxiangru
Repository
tangxiangru/AutoR
Created
6 months ago
Last Updated
2 weeks ago
Language
Python
License
NOASSERTION

Integrates with

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Solid

a-comparison-you-never-run-defaults-to-your-preference

Use at hypothesis drafting, study design and implementation when the task could plausibly be attacked by more than one family of method - hand-built features fed to a fitted model, a network trained on the raw structure or sequence, a pretrained backbone, retrieval - and your hypothesis list mostly compares variants inside one of them. Covers separating the hypotheses that would change what you build from the ones that would change an argument, requiring code on both sides of a family claim, running the comparison at a budget you actually have and reading each side's slope rather than its level, and demoting a comparison you will not run into a priced assumption.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

a-throughput-number-belongs-to-the-runtime-not-the-model

Use at the survey, study design and implementation stages when the task needs a pretrained model, solver or library you must download and run on the machine you were given, and the first configuration you try is too slow to cover the split in the time you have. Covers why the seconds-per-item you just measured is a property of the runtime, the workload and the machine as much as of the component, which field to change before demoting it, re-asking the component question after you fix the runtime, and checking that your fallback still has the property you picked the original for.

805 Updated 2 weeks ago
tangxiangru
AI & Automation Solid

cost-the-rung-you-need-not-the-cheapest-one-in-the-family

Use at literature survey, hypothesis generation and study design of a task that hands you a training split and an unlabelled test split, scores predictions by a fixed error metric, and has published best numbers for that dataset and metric, when you are about to decide that the kind of model behind those numbers does not fit your clock. Covers which rung of the ladder is worth timing at all, replacing "hours to finish the reference schedule" with "epochs until this passes what I already have", and pricing the first member of another family before the Nth member of this one.

805 Updated 2 weeks ago
tangxiangru