Skill v1.0.0
currentAutomated scan100/100version: "1.0.0" name: reproducibility description: Make computational work reproducible and re-verify reported numbers — pinned environments, fixed seeds, deterministic pipelines, Makefiles/CI, and re-running code to confirm headline results. Use when setting up an experiment repo, pinning dependencies, making runs deterministic, writing a Makefile/CI, packaging results so others can rerun them, or auditing whether a reported number actually reproduces. Trigger on "reproducible", "pin the environment", "set seeds", "deterministic", "Makefile", "rerun and verify", "can we reproduce". Pairs with research-engineer and verify-evidence.
Reproducibility — if it doesn't rerun, it isn't a result
A number nobody can regenerate is a claim, not evidence. This skill makes runs deterministic and reproduces headline numbers from the code before they're trusted.
Pin the environment
python -V > .python-versionpip freeze > requirements.txt # or use uv/poetry/conda lockfiles# prefer a lockfile (uv.lock / poetry.lock / conda env export) over loose pins
Record: OS, Python/R version, key library versions, hardware if GPU/threads matter.
Make it deterministic
import os, random, numpy as npSEED = 0random.seed(SEED); np.random.seed(SEED)os.environ["PYTHONHASHSEED"] = str(SEED)# torch: torch.manual_seed(SEED); torch.use_deterministic_algorithms(True)# sklearn: pass random_state=SEED everywhere
Eliminate hidden nondeterminism: unordered dict/set iteration in outputs, parallelism races, wall-clock or time()-based values, network/data that changes.
One-command rebuild (Makefile)
.PHONY: all data train evalall: evaldata: ; python src/make_data.py --seed 0train: data ; python src/train.py --seed 0 --out artifacts/model.pkleval: train ; python src/eval.py --model artifacts/model.pkl --out results/metrics.json
Anyone runs make all and gets the same results/metrics.json. Add a CI job that runs it on push.
Reproduce-the-number audit (verify-evidence mode)
- Locate the exact code path that produces each headline number.
- Run it fresh from a clean state with the recorded seed/env.
- Compare to the reported value — match within tolerance, or flag the discrepancy.
- Record provenance: commit hash, command, inputs, output file, and the value.
- Tag each number Code-verified (reproduced) or Open (could not reproduce — say why).
Package for handoff
READMEwith exact rerun steps;data/(or a fetch script + checksum);src/; pinned env;
results/ regenerated by make. Seed and commit recorded next to every figure/table.
When to use / when NOT to use
- Use when: setting up an experiment repo, pinning an environment, making a run deterministic, packaging results for handoff, or auditing whether a reported number reproduces.
- Do NOT use when: the question is which statistical test is correct (route to statistics) or whether the evidence supports the claim (route to verify-evidence). This skill only certifies that the number regenerates.
Anti-rationalization table
| If you're tempted to… | Why it fails | Do this instead | |
|---|---|---|---|
| "pip freeze later, the versions are basically stable" | A silent minor-version bump changes a metric and you can't tell when | Pin or lock the environment before the first real run; commit the lockfile | |
| "the seed doesn't matter, the effect is huge" | Unseeded splits and inits make the exact headline number irreproducible and invite p-hacking by rerun | Set and record the seed for data, model, and split; log it next to the result | |
| "I'll just fix the one stale value in results.json by hand" | A hand-edited result is no longer traceable to code; it is Asserted, not Code-verified | Re-run the pipeline and regenerate the whole file | |
| "it reproduces on my machine, ship it" | Author-only reproducibility hides hardcoded paths and nondeterminism | Re-run from a clean clone or CI before tagging Code-verified |
Exit criteria
- [ ] Environment is locked (lockfile committed) and OS and runtime recorded.
- [ ] All randomness seeded; no wall-clock or iteration-order leakage in outputs.
- [ ]
make allregenerates every headline figure and table from raw inputs. - [ ] Each headline number is tagged Code-verified or Open with provenance (commit, command, input, output file, value).
- [ ] A fresh clone or CI run reproduces results within stated tolerance.
Red flags (stop and escalate)
- A number only its author can regenerate. Tag it Asserted; do not ship.
- Any manual, unscripted step in the raw to processed to result chain. Script it or stop.
- Reproduces numerically but the test choice looks wrong. Route to statistics.
- Reproduces but may not support the claim. Route to verify-evidence.
<!-- Anatomy (when-to-use, anti-rationalization, exit criteria, red flags) adapted from agent-skills (MIT). See /ATTRIBUTION.md. -->