<< All versions

Skill v1.0.0

currentAutomated scan100/100
yacb2/aidex/plan-exec
──Details
PublishedSeptember 28, 2026 at 04:28 AM
Content Hashsha256:9a1f59ec50918ad7...
Git SHA
──Files
Files (1 file, 18.9 KB)
SKILL.md18.9 KBactive
SKILL.md · 290 lines · 18.9 KB

version: "1.0.0" name: plan-exec description: 'Use whenever the user asks to execute, implement, resume or continue a written plan — a .context/plans/ document or any plan with checkboxes/phases — however small the phase looks, even a single-file edit: the plan''s checkboxes and Execution log are updated by this skill, so do not carry out the phase directly from the plan text. Fires on "implement the plan", "execute plan X", "let''s execute the plan", "continue with phase Y", "resume the plan", "run phase 1 of the plan", "run the plan phase by phase". Enforces between-phase discipline: code-review, commit, handoff when context grows. Not for: creating the plan itself (/aidex:plan); one-shot tasks with no phases; bug fixes (/aidex:bugfix); pure refactors with no plan document.' disable-model-invocation: false allowed-tools: Bash Read Write Edit Agent model-policy: per-stage


Plan Execution

Drive the implementation of a written multi-phase plan with consistent between-phase discipline: review the diff, commit, and hand off the session when context grows. This skill centralizes the workflow so the user does not have to repeat it in every prompt.

Default autonomy

On run start, apply Mode A autonomy automatically — do not wait for the user to grant it. Questions live in the initial alignment moment only; after that the run proceeds start-to-finish (deny/pre-authorized/mandated/autonomous — see "Operating mode" below).

Operating mode

Front-loaded, then autonomous start-to-finish. Resolve every question at Orient (phase 0); after that, run all phases without interrupting. Follow the shared autonomy canon (autonomy-conventions.md). The operative rule here:

  • Ask everything up front, at Orient. Surface clarifications and confirm any

publication the plan implies (deploy/publish/release) before phase 1. If the plan did not pre-authorize a publish step, surface it at the end — not mid-run.

  • Evaluate batch-promotion at Orient (mandatory, one line). Before phase 1, decide

whether the plan's afk-impl phases should run as a durable Workflow, and say so in one line. The full rule is in `references/01-unattended-batch-execution.md` § Promotion at Orient. It is a kickoff decision, never a mid-run interruption.

  • Do not re-ask for steps this skill mandates. Invoking plan-exec authorizes you

to code-review the diff, author the commit message, commit per phase, and hand off when context grows. Do them — never stop to ask "should I commit? is the message OK? should I review? should I hand off?"

  • Planned migrations and dependency changes are autonomous. If the plan calls

for a migration or a dep install/update/downgrade, run it — commit, deps, and additive migrations are not gated. A destructive migration (data loss) is the exception: it stays gated (global DB rule).

  • A mid-run bifurcation that is not destructive → do it and document it (in the

plan doc / final summary) so you can review it afterward. Don't stop for a doubt that breaks nothing; verify the assumption (investigate, don't guess).

  • Only stop for: a deny-class destructive action (skip + document), an

un-pre-authorized publish (surface at the end), or a genuine hard blocker you cannot resolve (missing credentials, truly unknowable intended behavior).

  • **On an ambiguous fork you cannot cleanly classify — consult the

durability-arbiter before stopping. Read [`../conventions/agents/durability-arbiter.md`](../conventions/agents/durability-arbiter.md) and pass it to the Agent tool as the prompt (`model: sonnet`, `effort: high`, read-only — `model-policy: per-stage`, so the gate's depth is pinned rather than inherited from the run it is judging), with the situation + the run's autonomy surface + the phase's proof (verification output, commit SHA). Follow its `CONTINUE` / `ASK` / `STOP` verdict; batch any `ASK` to the end. If it errors or returns nothing, apply the rule above and proceed — never block on the arbiter** (it is a forcing function, not a gate).

Otherwise: proceed. The user will redirect if needed.

Unattended / batch execution (opt-in, gated)

The default path above is interactive. For unattended/batch runs ("execute the whole plan while I'm away"), this skill can launch the plan as a durable Workflow — each phase a fresh bounded agent, a two-stage gate per phase, crash-resumable via the journal. Promote only when the work is decomposable + machine-verifiable + unattended and each phase's real work dwarfs the per-agent floor; the mandatory Orient evaluation is the opt-in.

Read `${CLAUDE_PLUGIN_ROOT}/skills/plan-exec/references/01-unattended-batch-execution.md` before promoting anything — promotion threshold and its measured ~22k/agent cost floor, the three shipped workflow forms and how to pick one, how to derive args from the plan, the phase tier map, and what happens when a phase fails its gate.

Workflow

0. Orient

  1. Read the plan document fully (path is in the user prompt or in

.context/plans/). If multi-file, read 00-index.md plus the current phase file. You may skip Execution log entries for already-completed phases (canon §Execution log) — they are proof journaling, not spec.

  1. Identify: total phases, current phase (first unchecked checkbox), success

criteria per phase, verification step.

  1. Check whether this plan is bug work. Plans carry no type field; the

back-link runs the other way, so resolve it by grep: grep -rl "escalated_to: plan/<slug>" .context/backlog/. If the originating item carries type: bug, every behavior-changing phase is bound by RED→GREEN (bugfix): the test is written and fails for the right reason before the fix, and the GREEN output is that phase's proof. No matching item → carry on normally.

  1. Check the prior phase's review evidence. If a previous phase completed

this session or an earlier one, confirm its Execution-log entry in 00-index.md carries a review: <verdict> · <n> findings line. A missing entry means the between-phase code-review was skipped — run it now, on the prior phase's diff, before starting the current phase; do not proceed silently on an unreviewed phase.

  1. Honor the plan's Isolation surface if it declares one. If the plan already

recorded an Isolation note (from plan's Step 5, at plan-creation time), act on it directly: run the recorded worktree.sh new command before phase 1 (--no-infra only when the plan says code-only); if the project has no worktree setup, EnterWorktree and note it. If the plan predates this and has no Isolation note, run worktree bootstrap if .context/worktrees/00-index.md does not exist yet, then worktree.sh new here at Orient, before phase 1. Enter the worktree only if the plan/user authorized it — do not auto-enter one that was not approved. Before creating any worktree/branch, resolve and state its base branch and require explicit confirmation if it is not the repo's default (worktree's branch-base rule) — never fork off the ambient checkout silently. No declared surface and no plan-recorded parallelism → run in place. The plan doc stays source-of-truth in the main tree: a fresh worktree has only committed files, so update the plan and record proof_links at its main-tree path (a gitignored/uncommitted .context/ plan is absent from the worktree) — see the canon's Lifecycle note.

  1. Probe for concurrent work before touching anything. The user runs

parallel sessions and worktrees on the same project, and a session blind to them will take decisions another session owns. Two commands, seconds: git worktree list and git log --all --since="24 hours ago" --oneline. If another live line of work shows — a worktree you did not create, fresh commits this session did not make — name it in your first status message, keep hands off its files and branches, and route any decision that belongs to it back to the user instead of taking it here.

  1. Set the run's spend before phase 1. Each phase's tier says how hard its work

is; the canon's table (plan-conventions.md §"Optional phase metadata") maps that to a default model and effort level. Those cells are defaults, and this is the one moment to override them — remaining quota, model availability, or a cheaper cell the user prefers. State the override and its reason in one line and log it to the Execution log; a phase whose cell you changed is not a phase whose plan you edited. Ask nothing: an unstated tier is standard, and no override is the default. A tier row whose cell restricts `tools:` also needs a registered agent definition, since agent() has no tools option: the toolset travels only as agentType, the name of an agent file. One definition per tier the plan uses (never per phase) — model/effort from the canon row, tools: exactly the row's list, user-invocable: false — living in the project's .claude/agents/<name>.md, or ~/.claude/agents/<name>.md when several projects share it; those are the two locations an agent name resolves from, project first. Pass its name as the batch phase's agentType. No definition → the phase runs unrestricted and pays ~49k of prefix per call instead of ~14k. The definition must be on disk before this session started. The agent registry is read at session start: a stub written mid-session is invisible to the session that wrote it, and a batch naming it dies at once with agent type '<name>' not found. So check for it here, at Orient — if it is missing, write it and hand off; the next session sees it. Two things measured on a real batched run (2026-09-20): the restricted implementer's first call read 15,826 cache_creation against 50,841 for the same run's unrestricted verifier, and the phase's model cell won over the definition's model: line — agentType carries the toolset, nothing else.

  1. Create a TaskList mirroring the plan's phases so progress is visible.
  2. Front-load the work-list for chained multi-item runs. A single plan's phases

are already an ordered queue (walk them). But when this session chains multiple plans/items (close several plans, then clear backlog), fix the cross-item order once here — via the AskUserQuestion survey → a durable .context/worklists/ work-list (see worklist-conventions.md). Then walk it with worklist-advance.sh between items instead of pausing to ask "what next?". Emergent work (class b) is appended (--append) and continued, not asked; only a class-(c) fork or the publication gate interrupts. No interactive channel (claude -p, cron): skip the survey, walk the items in the order they were given, and record the defaulting in the run's final summary — autonomy-conventions.md § When there is no interactive channel.

1. Execute each phase

For each phase in order:

  1. Implement the tasks in the phase. **Plan code is a sketch, not a paste

source: any code block or line reference in the plan was frozen at plan-write time — before applying one, read the current file, confirm the surrounding code still matches, and check for sibling call-sites/branches the plan did not enumerate. The phase's acceptance criteria and gate are the contract; the plan's code is illustrative except inside a Contract** block (exact signatures/shapes/DDL), which is binding.

  1. Run the verification step the plan declares (tests, type-check, build,

manual check). If none is declared, run the minimum that proves the change works (relevant test suite + type-check). Iterate on the selection, not the whole suite: ${CLAUDE_PLUGIN_ROOT}/skills/audit/scripts/affected-tests.sh --command prints one runnable command for the tests covering the phase's diff; exit 3 means no selection is available, so run everything and say so. The between-phase checkpoint commits on the selection, stated rather than silent — say which subset ran and that the full suite has not. The full suite gates the INTEGRATION boundary: the merge, the push, or the end of the run (decision/2026-08-24-full-suite-gate-moves-from-commit-to-integration). An # INCOMPLETE selection is the exception — unmapped scope forces the full suite in-phase. It also names changed files that measurably break and have no E2E — write that spec in-phase.

  1. If verification fails: fix root cause. After 3 failed attempts on the same

approach, stop and ask the user.

  1. Mark the phase's checkboxes as done in the plan file. **Record the phase's

proof** — the verification output, the commit SHA, a request/response payload, or a screenshot of the flow — in the plan front-matter proof_links (or under .context/proofs/<slug>/ for larger captures) per conventions (00-global.md §7.1). Don't mark a phase done you can't show works.

  1. **When execution departs from the plan, the plan edit ships in the commit that

departs** — not in the close-out commit. Otherwise the phase diff is reviewed against a stale plan, and the reviewer cannot tell a deliberate change of course from an omission. Same coupling plan-conventions.md already applies to tests. Where the plan is not committable this collapses to editing it before the commit, in the same turn: aidex's own repo gitignores .context/, and it is the exception — 1,135 plan files are tracked across the six fleet repos, 0 here (census 2026-09-07).

Scoped plans carry a file contract. When the plan's front-matter says
mode: scoped, its **Files:** list is the declared blast radius, written on
deliberately incomplete investigation — so it will sometimes be wrong. The contract makes
that visible, not impossible: (1) log any file you touch outside the list in the
Execution log, one line; (2) re-apply the five triage signals (plan-conventions.md
§The five signals) to that file — if any flips to full, stop and re-triage the
whole plan with plan; (3) independently, if the file list has doubled, stop
and re-triage. Six extra trivial files trip no signal but mean the contract misread the
change — a failure rule (2) cannot see.
Loop (opt-in, per phase only): if a single phase is mechanical and its verification is a
pure machine gate (e.g. "make all <suite> pass" / "type-check clean"), that one phase may be
spec'd as a loop via loop and run by /goal/ralph-loop — mirroring the plan →
loop pointer at the phase level. Do not loop the executor itself: the between-phase
checkpoint (review/commit/handoff) is judgment work, and irreversible steps
(push/release/deploy) stay outside any auto-loop and human-gated. (commit is
not irreversible — it is part of the checkpoint, not a gated step.)

2. Between-phase checkpoint (MANDATORY)

After each phase passes verification, before starting the next phase, run the shared checkpoint — read ${CLAUDE_PLUGIN_ROOT}/skills/conventions/references/checkpoint-conventions.md and follow its four moves (scoped review with its recorded anchor and findings addressed by remedy · commit · register-don't-discuss · auto-handoff without asking). It is one canon with two consumers (this skill and the backlog sweep) and is not restated here; test_checkpoint_lockstep.sh fails this file if it grows its own copy. What is plan-specific:

  • Scope. ${CLAUDE_PLUGIN_ROOT}/skills/conventions/scripts/resolve-review-scope.sh --files working-diff,

or --base <phase-start-sha> branch-vs-main for a phase that spans commits — and ${CLAUDE_PLUGIN_ROOT}/skills/conventions/references/review-scope-conventions.md owns which reviewer covers which scope. Exit 3 is an empty scope, never a passing review.

  • Where the evidence goes. The Execution log in the plan's 00-index.md takes the

review: <verdict> · <n> findings · scope=<scope> anchor=<anchor> line before the commit.

  • Deferrals use register-item.sh --origin plan --plan <this plan>

(`references/03-deferring-emergent-work.md`).

  • Which model runs which step — orchestrate, implement, and do the mechanical work

with different models (`references/04-model-tiering.md`).

  • The seed's `slug:` line is the PLAN's name plus the phase — the name existed before

the first handoff and does not move. CHARTER comes from the plan's name and goal.

3. Final phase

After the last phase:

  1. Run the full verification the plan declares (or the project's standard

pre-deploy check: tests + build + lint).

  1. Code-review and commit as above.
  2. If the plan implies a release (user-facing changes, feature complete):

surface the project's release command as an option (detect it — many projects expose a /release-style command). Do not run it without explicit user approval — releases are deploy-coupled.

  1. Update the plan document: mark all phases complete, add a closing note

with the final commit SHAs if useful.

  1. Close out the run: tear down isolation if a worktree was entered at Orient, log the

worktree usage line, suggest a coverage sweep if the plan touched mapped src paths, run guided human verification — it emits a proof artifact, and a plan with nothing human-visible skips it by recording the reason, never by omission — reconcile deferrals to a BL-NNN or a CLOSE, and fire the notifier. Read ${CLAUDE_PLUGIN_ROOT}/skills/plan-exec/references/02-close-out.md and follow it step by step — each step has a guard and an ordering that matter, and doing them from memory is how a worktree survives its plan.

Per-project adjustments

This skill ships stack-agnostic defaults. Projects often override them — detect the project's own conventions, don't assume:

  • Review/commit/release commands. Use the project's own slash commands or

helpers (look in .claude/, available commands, or CLAUDE.md). Do not assume a specific command name exists.

  • Stricter project rules. Read the project's CLAUDE.md (and any project

memory) before the first phase — it may define test runners, commit style, version-bump coupling, or release gates this skill cannot know about.

If the project's CLAUDE.md or memory contradicts this skill, the project wins.

What this skill does NOT do

  • It does not create plans (use plan).
  • It does not skip verification to move faster — every phase is verified.
  • It does not deploy or release without explicit user approval.
  • It does not run E2E tests against dev environments — use the project's

isolated test runner if E2E is required by a phase.

All versions