<< All versions
Skill v1.0.1
Automated scan100/100unclejobs-ai/second-claude-code/loop
+4 new
──Details
PublishedAugust 17, 2026 at 01:08 PM
Content Hashsha256:f4c06ac9261fbe77...
Git SHA3ea8f7107919
Bump Typepatch
──Files
Files (1 file, 3.5 KB)
SKILL.md3.5 KBactive
SKILL.md · 68 lines · 3.5 KB
version: "1.0.1" name: loop description: "Use when benchmarking and evolving prompt assets through a fixed-suite optimization loop" effort: high disable-model-invocation: true
Iron Law
A loop without a benchmark is meaningless.
Red Flags
- "Let's just run it and see what happens" → STOP, because the Iron Law says a loop without a benchmark is meaningless — a fixed suite with numeric scoring must exist before any run.
- "The new variant feels better" → STOP, because evaluator output must expose a numeric score that can be parsed deterministically — subjective improvement is not a valid promotion signal.
- "Baseline is close enough, let's promote the candidate" → STOP, because winners are promoted only when the score beats baseline by at least
min_delta— close is not enough. - "I'll mutate the docs and test files too" → STOP, because
.mjs,docs/,README*, andtests/are excluded from mutation scope — only allowed targets in the suite allowlist may be changed. - "One more generation might improve it" → STOP, because the loop must stop when budget is exhausted,
min_deltais not met, or scores plateau — open-ended iteration wastes tokens without measurable gain.
Loop
Run a Karpathy-style optimization loop over the plugin's own prompt assets.
When to Use
- Maintainers want to improve
skills/**/SKILL.md,agents/*.md,commands/*.md, ortemplates/*.md - A fixed benchmark suite with numeric scoring already exists
- Candidate variants should compete under the same time budget and the best variant should be promoted to an isolated branch
Subcommands
| Command | Purpose | |
|---|---|---|
list-suites | List bundled loop suites in benchmarks/loop/ | |
show-suite <name> | Inspect the validated suite manifest | |
run <name> | Create an isolated run branch, score a baseline, generate candidates, and promote a winner if min_delta is met | |
resume <run_id> | Reload the saved loop state from .data/state/loop-active.json |
Workflow
- Use
node scripts/loop-runner.mjs show-suite <name>to load the suite and verify the allowed targets. - Use
node scripts/loop-runner.mjs run <name> ...to create the isolatedcodex/loop-...branch and run worktree. - Inside that run worktree, create 3-5 prompt variants only within the selected targets.
- Evaluate every candidate with the suite's fixed case commands. Parse scores from JSON or review-style output only.
- Keep the top 1-2 elite candidates per generation. Stop when budget is exhausted,
min_deltais not met, or the loop plateaus. - Only promote a winner when all hard gates pass and the score beats baseline by at least
min_delta.
Hard Gates
- Selected targets must stay within the suite allowlist and the global v1 mutation scope.
- Contract/runtime smoke checks must still pass when those files exist in the repo.
- Evaluator output must expose a numeric score that can be parsed deterministically.
- Winners are promoted only into the run branch/worktree, never into the current workspace.
State
- Active state:
${CLAUDE_PLUGIN_DATA}/state/loop-active.json - Artifacts:
.captures/loop-<run_id>/ - Bundled suites:
benchmarks/loop/*.json
Gotchas
- Do not add routing for this skill in
prompt-detect.mjsfor v1. It is slash-command only. - Do not mutate
.mjs,docs/,README*, ortests/as loop targets. - Do not declare a winner when baseline remains best or when
min_deltais not met. - Do not write to the main worktree directly. The isolated loop branch is the output surface.