Skill v1.0.0
currentLLM-judged scan95/100version: "1.0.0" name: verify description: | Evidence protocol for completion claims (the Seal Test): name the proving command before running it, show the check can fail, bind the result to a tree state, and label every claim executed / read / unverified. Use when: about to say done, fixed, passes, or complete; before a commit; before handing work to the review chain. NOT for: deciding whether the work is worth doing, or reviewing someone else's diff (code-review).
verify — the smith's own gate before handoff
A completion claim is a mark, and a mark must be earned. The community baseline ("no completion claims without fresh verification evidence") stops at freshness — it accepts any green command as proof. This protocol demands more: a check that cannot fail proves nothing, and evidence that outlives its tree state proves less than nothing.
A claim is sealed only when all four conditions hold:
The four conditions
1. Named before run
State the proving command before executing it, next to the claim it proves. Choosing the command after seeing what passes is how a test suite's green becomes "feature works" — the claim-to-command mapping is the verification, the run is just its execution.
2. Able to fail
Show the check CAN go red. For a bug fix: the reproduction failed before the change (or fails when the change is reverted). For a new test: it was observed failing against the pre-change code. A check that has never been seen red is a counterfeit seal — this is the condition the generic gates skip, and the one that catches assertion-free tests, mocked-away behavior, and wrong-file test runs.
An absence claim is only as good as the search behind it
"There are no callers", "the original has no such branch", "nothing reads this column" — these are claims about the whole tree, and their entire evidence is that one search did not find something. An existence claim comes with a coordinate anyone can open; an absence claim comes with nothing to open, so the cost of checking it is asymmetric and it tends to pass unread.
Condition 2 applies to the search itself: run it against a case you know is present. A pattern that finds nothing there was broken, not the tree empty. Report the command and the scope it ran over, so the next session re-runs it instead of re-inventing it.
An absence you cannot demonstrate that way is unverified and reads "not found", which is a different sentence from "not there". Never promote one to a document — a policy or a ledger row built on a search nobody could repeat is the same defect as a read reported as passing, with a longer fuse.
arm-check asks condition 2 of a whole module, one arm at a time
Condition 2 is asked of the case in front of you. A module has arms nobody ever asked it of, and counting them by hand is how the count rots: #262's own table says 33 arms where hooks/review-history-guard.py now has 31, because the file changed twice after the count was taken.
arm-check hooks/review-history-guard.pyarm-check hooks/review-history-guard.py --tests "bin/test tests/test_chain_hooks.py -q"
With no --tests it lists the arms and mutates nothing. With --tests it makes each arm wrong in turn, runs that command, and names the arms nothing kills — restoring the module from held bytes and comparing the sha256 after every mutation, never with git checkout, which reaches the uncommitted work in the rest of the tree.
`--timeout` bounds the wait, and not the work. Each operator's command is waited for at most 900 seconds by default (--timeout 0 removes the bound), and an arm asks two operators, so one arm can take twice that. When the bound is reached, only the command's own process is killed. A wrapper command is the one that leaks: where --tests names a runner that starts pytest as a child of its own, as the example above does, a timed-out pair leaves that suite running, unbounded and unreported, beside every arm after it. Against a module whose suite can approach the bound, name the pytest command in --tests directly rather than a wrapper. Whether the bound should reach the whole process group is #313.
There are two ways to be wrong and the counts differ by a lot, so the report keeps them apart and the number you quote has to say which one it is:
| Operator | Asks | Measured 2026-09-09 on hooks/review-history-guard.py | |
|---|---|---|---|
invert | would a case notice this test being backwards | 0 of 29 survived | |
remove | would a case notice this arm being absent | 12 of 32 survived |
That third column is a measurement and not a property of the module: it moves when either the module or tests/test_chain_hooks.py changes, which is the rot this checker exists to end. Re-take it with the command above rather than reading it as current — a number in a document is exactly what #262 says goes stale.
The two denominators differ because a pair both operators answer identically is asked once. A handler with one type left is aimed at what nothing raises by either operator, so invert refuses it and the report names the pair; remove keeps it, because #262's table is a removal count. Three of this module's arms are that shape.
An arm counts watched when either is noticed, because does any case depend on this arm is answered by one yes. #262's own table of unwatched arms is a remove count — every sentence in it is about taking something out — so that is the row to compare it with, and the combined total is not.
An arm is what #262's rule says it is: an ExceptHandler counting each member of an except tuple separately, or an If, While or IfExp counting each top-level member of the boolean test separately, plus a match_case's alternatives and guard and a comprehension's if guards. The walk is derived from the grammar rather than from a list, and it refuses an AST node type it does not recognise instead of skipping it — a walk that skips silently is the rotted hand count with a shebang on it.
Two things it does not claim. A survivor is not automatically a defect: an arm that cannot be constructed, or one whose removal preserves behaviour, belongs in the report and is not work anybody owes. And it is report-only, exit 0 either way — whether an unwatched arm should fail a run is an open decision, not an omission.
3. Bound to the tree
Evidence attaches to a tree state, not to a session. Note the state the proof ran against (commit + dirty files); any edit after the run breaks your seal — re-run, don't re-tell. Same drift logic the evidence ledger applies to spec coordinates, applied to your own claims.
4. Executed, read, or unverified — labeled
Every claim carries one label. executed (ran it, read full output, exit code checked) · read (inferred from reading code — a judgment, not proof) · unverified (say who or what can answer). Reporting a read as passing is the lie this whole skill exists to prevent.
Procedure
- List the claims this response is about to make.
- For each, name the proving command (condition 1). No command exists →
label read or unverified, never "should work".
- Run fresh; read the FULL output; record exit code and the line that
proves or refutes (conditions 2–3).
- Any claim red or unproven → report the actual state. Your seal is
withheld, not negotiated.
Every agent seals what it verified, and one of them is final
Every agent seals what it verified, and the one seal over the whole project is the sealer's. A smith's proof block is that smith's seal over its own slice and is legitimate; a warden's report seals what its review looked at. So the word is not overloaded by having many instances. What was missing is that one of them is final, and nothing anywhere said so.
Two properties already in the design tell the final one apart, which is why they are the two to name rather than some new mark invented for it.
- Scope. Every other seal covers what that agent touched. The sealer's
covers a tree nobody is still editing — which is why nothing edits between the broad seal and the PR is a rule about that one seal and about no other.
- Form. Every other seal is text. The sealer's is the only one drawn, and
only over a run that earned it: broad-gate draws the stamp, or writes the values it is drawn from, when every check passed and the cell was written, and on no other path. It draws on a person's terminal itself; anywhere else the Stop hook draws those values at the end of the turn of the session that spawned the sealer. So seeing the drawing means the last seal was earned.
A bare "the seal" is ambiguous the moment more than one exists, so every reference to an INSTANCE names whose — the smith's seal, the sealer's seal, the warden's review mark. The concept and its formats stay bare: the Seal Test, a seal block, a counterfeit seal, SpecSeal itself. The shape to watch for is a sentence that names one party and leaves the seal anonymous — the warden's audit of the seal names an auditor and not what is audited, and the answer there is the smith's.
Scope — cheap and often, broad and once
Condition 3 decides where a broad check belongs. Evidence binds to a tree state, so any run followed by an edit was spent, not banked. Measured on one work item here: twenty broad runs, every one of them followed by more edits — 121 after the first, 1 after the last. Not one of those seals survived to the handoff, and they cost more than half the session's tool time.
Review rounds are edits you have already scheduled. Sealing the whole suite before them seals a tree that is going to change.
| When | What runs | Seal | |
|---|---|---|---|
| Each slice | the tests for what you just wrote | executed, narrow | |
| Phase boundary | your module and the ones it touches, in parallel | executed, scoped | |
| Handoff to review | nothing broad | the suite is unverified, and says so | |
| Review rounds 1..n | what the fix touches | executed, narrow | |
| After the rounds settle | full suite, lint, typecheck — once | the one broad seal |
Narrowing is not permission to go quiet. A suite you did not run is unverified with who answers it, never omitted — condition 4 is what makes the smaller scope honest rather than a loophole.
The two ways that loophole reopens are below. Both were measured, and both looked like a sealed claim at the time.
The narrow command still has to be able to fail
Condition 2 does not relax when the scope shrinks. A smaller command is still a command being chosen, and it is easier to choose wrong — it can come back green because it examined nothing.
Measured: a documentation-only change was sealed with a Python linter aimed at the docs directory. That directory holds no Python.
$ ruff check docs/warning: No Python files found under the given path(s)All checks passed! (exit 0)
Exit 0, nothing read, and a green line in the seal block. Before a narrow command counts, say which files it opened. If the answer is none, the honest row is unverified, not a passing one.
What can fail depends on what changed, which is the axis worth picking the command from:
| What changed | What can fail | |
|---|---|---|
| Prose only — docs, comments, docstrings | Whatever actually parses those files. Often nothing does, and then there is no seal to earn here. | |
| A comment or docstring inside source | That the module still imports. | |
| Logic | The tests covering that file. | |
| A merge, or a dependency bump | The broad gate. What breaks is outside what you edited, so a scope drawn around your own edit cannot see it. |
The last row is the one that gets narrowed by mistake. A merge is the moment the tree stops being the one you tested.
The answerer has to exist
unverified; CI covers it satisfies condition 4 on its face and voids it underneath. The row names an answerer; nothing ever checked there was one.
Measured: a session narrowed to a single domain's tests and deferred the rest to CI. That repository's workflows assigned reviewers, deployed on push to the default branch, and validated a migration graph. None of them ran the suite, the default branch had no protection, and the pre-commit hooks were lint and typecheck — on the committer's machine. The deferred suite had no answerer at all, and the smith's seal read as though it did.
Resolve the answerer before writing the row:
deferral-check . # is the test suite answered here?deferral-check . --kind all # tests, lint, typecheck
It separates four outcomes that a reader otherwise collapses into one: something answers on pull requests · something answers, but only after the point being deferred · a local hook answers on the committer's machine only · nothing answers. Exit 1 means the row you were about to write is not true.
It reads command text rather than a YAML graph, so a runner reached through a composite action, a reusable workflow, or a script it cannot see reads as absent. Disagreement with what CI really does is a finding about this tool — not permission to keep the row.
And something has to read the row afterwards
An answerer that exists still answers nothing if nobody returns to the row. One here said the gates had never been seen rendering in a TUI and named the user; months later the user hit exactly that and asked why every gate is yes/no.
unverified-check . # what is still open, and whereunverified-check --baseline origin/main seal/specs/ # and nothing was deleted from it
It never fails for an open item — punishing an honest row is how sessions learn to write none. It fails for a section it cannot read, because a tolerant reader reports zero there and zero reads as "everything has been closed". Close an item by marking it ✅ with what closed it; the row stays.
`--baseline` names the branch you merge into, and what it reads is `git merge-base <ref> HEAD` — where you forked from it on a branch checkout, and the base's tip in CI, whose checkout of a pull request is your head already merged into the base. Either way a work item squashed into that branch after your branch was cut is not your removal (#272). Before that, squashing one item turned every sibling red, and each one paid a merge, a re-run broad gate and a re-pushed pull request to clear it.
The broad gate — after the rounds, then compare against the base
It fires after the review rounds settle, never before them. Nothing broad runs at the handoff to review; the proof block carries the suite as unverified, and that label is what keeps the narrow scope honest. §Scope holds the reason and the measurement — a seal taken before the rounds is spent by the first finding. Findings are the expected case rather than the exception: the review chain runs up to three rounds, and five while a 🔴 is open (docs/review-chain-spec.md).
"After the rounds settle" is a row, not a moment, and this is the row: the last `rounds/round-N.md`'s `Pass` box is checked. Nothing in that record's verdict table is still open. The phrase alone names no particular rounds — a work item has build phases with a progression of their own and a review chain with its own — so a reader who reaches for the moment has to guess, and one who reaches for the box does not.
Needs a fix: no is the ordinary way a run arrives at a checked box, and it is the reviewer's own answer rather than a reading of the table. The two part on one case: a run that ends at the round cap closes its last finding deferred <home>, which is a closing word, so the box is checked while the reviewer's row keeps the yes it had while the round was running. Nothing rewrites that row afterwards, and nothing should — it is what the reviewer concluded. So the box is what says the run ended, and round_record.py seal — skills/code-review/scripts/round_record.py, the generator the review orchestrator types as round-record — refuses on the box for that reason.
It belongs to the `sealer`, and the four conditions above are its whole procedure. The rule used to say when the gate fires and which agents may not take it, and named nobody who may — so it was assembled from these sentences by whichever session remembered them, differently each time. agents/sealer.md is the agent, broad-gate --base <base> --record <item> is the command, and skills/code-review/orchestration.md §The last record's `Broad gate` cell is read at a READY pull request owns when it is spawned. What the sealer returns is a report; it judges no failure and fixes none.
An expensive suite argues for this placement, not against it. A run that takes fifteen minutes is a finding about the run and deserves its own ticket. Moving it ahead of the rounds means paying it twice.
Which base, and it is the one the merge is judged by. The gate resolves --base once, before anything runs, to the ref CI will read: the base's upstream where the checkout declares one, else origin/<base>, else the ref as given. Every check it runs takes that one commit — the survivor range, the two baselines, the scratch worktree the base comparison checks out, and the Broad gate cell. A plain branch name is a LOCAL ref and a runner has no local branches, so a checkout one commit behind its remote used to seal green over a question nobody was asking while CI refused the same commit (#423). Where resolving moves the answer the gate prints one line naming both refs, both commits and the distance, and runs anyway; where the two agree it prints nothing. It never fetches, so a remote-tracking ref is only as fresh as the last fetch — which is why the stamp's panel names the ref beside the commit rather than the commit alone.
What the sealer's seal covers is declared rather than remembered. The arms exist so that its one run says what CI will say, and for three releases the list was kept in step with .github/workflows/hygiene.yml by whoever remembered. #424 added a step to that workflow's release job, nobody added the arm, and no case in the suite went red — so from that merge onward a green seal covered a shorter list than the merge was judged by. skills/verify/scripts/broad_gate.py's PARTITION is the declaration that ends it: every step of that job is mirrored by a named arm or excluded with a written reason, there is no third state, and a case holds the table against the workflow from both sides. A step added to the workflow fails the suite until somebody classifies it, and a row naming a step that was renamed away fails it too.
A seal says what it did not answer. Where the repository being gated has that workflow, the panel carries a workflow row — <n> of <total> not answered — and the names of those steps go to stderr beside the line that names the repository's own command. The count is on the panel because a panel value is 23 columns and a step name is a sentence; the names are printed because a number alone sends the reader back to the two files this declaration exists to stop them opening. A repository with no such workflow sees neither, and nothing else about its run changes.
What the count does not say is whether a mirrored arm asks the same question its step asks. The partition says a step is on the list; two readers of one question can still disagree about what they are checking. #473 is the work item about that class, and it opened with one live instance in the repository this plugin is developed in. The gate ran the survivors and corrections arms on a main base, where SpecSeal's own workflow skips both steps. That instance is closed: where the base names main and the gated repository's workflow carries those steps, the gate leaves both arms out and says so (broad_gate.py#SKIPPED_AT_MAIN). A case holds that list against the workflow's guards. The case reads a guard on the base and no other kind of condition, so the class stays open for the next kind.
The arms the plugin ships are the arms the gate can run. Four steps of SpecSeal's own release job have a local answer and no arm: three run a script under .github/scripts/, which no plugin ships, and one is shell written inline in the workflow with no script either side can share. A repository that wants checks of its own sealed names them in the Broad gate row of seal/config.md — the row this gate already runs first — rather than in the arm list of a script that ships to everybody.
A full suite carries failures that were already there. Attributing them to this work is a wrong finding; passing over them is a silent pass. Both are avoided the same way: read the run against the base commit.
That comparison is reactive. A baseline answers whose failure this is, so it is taken once an unexplained failure appears, and only for the tests that failed — git stash, that one file, git stash pop. Measured: a session ran a 19,000-test suite serially for a baseline before any failure existed, and was about to run it a second time after the change.
- Failing on base too — not this work's finding. Name it, record it as a
follow-up, and it does not block.
- New — this work broke it. Back to the implement/review loop, and the
broad gate runs again afterwards.
Measured on one repository: a full suite showed ten failures, all ten reproducing on the base commit and none of them in the domain being changed. Without the comparison, every one of them is triage the author cannot act on and is not allowed to fix.
Three returns and stop. This counts trips through the gate, not the review chain's rounds. A fourth trip is not another bug; it means the narrow scope is missing a class of breakage, which is an architecture question for the user — the same reading as the 3+ Fix Rule.
It does not take the review rounds' exception, and the difference is worth stating because the numbers look alike. docs/review-chain-spec.md raises the round cap to five while a 🔴 is open, on the grounds that a round catching a regression the last fix made — coordinate named, patch already run — is not the shape the cap was written for. Every return through THIS gate is that shape: a broad gate goes red because something the fix touched broke. Raising the cap for the regression case would raise it for the normal case, which is the one thing it exists to bound.
Nothing edits between the broad seal and the PR. The tree that passed is the tree that ships. One more small fix after the run means the run is stale, and the gate is paid for again.
The cost of the run you repeat
Scope decides how often the expensive command fires. It does nothing about what the command costs, and the narrow run is the one that multiplies — every slice, every round.
Measure before accepting it. The suspect is usually wrong: one session blamed SQL echo logging for a slow suite, turned it off, and measured 131s → 137s. What it actually found was the suite running serially while a parallel runner sat installed and unused, its recipe written in a pyproject.toml comment with the measured numbers beside it (225s → 83s). A recipe in a comment is not a recipe. Nobody runs it, and an agent least of all.
So before paying a check's cost repeatedly, spend a minute on where the time goes, and look for the fast path the project already has: a runner in the lockfile, a target in the Makefile that is not the default, a line in a README or a comment. Repairing one is cheaper than paying its absence once per round.
Capture once, filter locally. A command re-run to see its output a different way returns nothing you did not already have. Measured in one session: the same test scope ran at 194s piped through tail, then again at 203s piped through grep "^FAILED" — three and a half minutes for a different view of a result already produced. Redirect to a file and read the file:
out=<scratchpad>/<work-item-id>/run.txt; mkdir -p "$(dirname "$out")"uv run pytest <scope> > "$out" 2>&1; tail -8 "$out"grep '^FAILED' "$out" # same run, second question
The tell is a second invocation whose only difference is after the pipe. The file is under the scratchpad and carries the work item id, never a name every session on the machine would also pick: agents of one session share one scratchpad, and a generic name was overwritten mid-round by a sibling work item's three times in one run (#544).
session-cost reads a finished transcript and reports the split — command time, model time between calls, the repeats, and how many tools went out per turn. It is what fills the cost row, since none of it is visible from inside:
session-cost --latest # newest transcript for this reposession-cost <transcript.jsonl>
Seal block
End with this block. Values that cannot be filled honestly stay none — <reason> (an unfillable row is a finding, not an embarrassment):
🔏 verify sealed @ <commit-ish>[+dirty: <files>]· <claim> — <command> → <key output line> (exit <n>) [executed]· <claim> — <where read, file:line> [read]· <claim> — unverified; <who/what answers> [unverified]· broad gate: <not yet — due when the last round record's `Pass` is checked | ran at <sha> vs base <sha>[; earlier run: <sha> vs base <sha>]>· cost: <n> check runs, <m> minutes of command time· red proven: <how the check was seen failing, or none — <reason>>
The cost row is there because nobody notices this from inside. One session spent 22 of its 26 command-minutes in the test runner across fourteen runs, and that only surfaced when someone parsed the transcript afterwards — session-cost is that parse, made repeatable. A number in the report puts it in front of the person who can decide the suite is worth fixing — the same reason none — <reason> is written rather than omitted.
The broad gate row is what makes "once" survive a handoff. A session picking up round 3 watched neither round before it, and no command leaves a trace in the code it ran against — so it either repeats a run that is already sealed or ships assuming somebody else made it. It belongs in round-N.md too, where the next session actually looks.
The block feeds forward: the smith ends reports with it, the warden audits the smith's seal instead of re-deriving it, and the round records carry it across sessions. A seal the warden cannot audit from the block alone was not a seal.
Measure the segment, and feed the flow log
After every smith or warden segment ends, measure its transcript with skills/verify/scripts/session_cost.py (the session-cost command above is the same script) and post the numbers, and what they say, as a comment on one of this repository's two measurement logs. Which of the two is not a filing detail — one of them is deleted at every release, so a reading put in the wrong one is a reading thrown away, and the paragraph below is what decides it.
There are two logs, and the difference is kind rather than scope. A reading about the segment that just ended — its span, its calls, its tools per turn, and what those say — belongs to the rolling flow-measurement log, which opens at a release, accumulates until the next version ships, and is discarded by the release that ships it. Readings that span versions — a rate held against an earlier version's baseline, or an observation that a later measurement answers — belong to the durable flow-baseline log, which is maintained rather than accumulated. Posting one of those to the rolling log schedules it for deletion at the next release.
A repository declares each of them the same way, with a label on an issue, and each keeps the same invariant: exactly one open. Neither is named by number, because a number goes stale the moment its issue closes.
One command does the lookup, the four readings below, and the post.
session-cost --segments <the run's transcript> --post --says <file|->
It resolves the label rather than a number, --label flow-baseline reaches the durable log with the same code, and it opens no issue in any state. It refuses without --says, because the numbers are the script's and what they say is yours: a command that invented the sentence would be posting a judgment nobody made. The readings below are what its exits mean, and they are still worth knowing, because a session that meets a refusal has to know which one it hit.
The command does not make anybody run it. What was measured is that this meter sat unreferenced through a full day of measurements nobody took — the measurement was not taken, not that the posting failed. --post removes the procedure a session otherwise reconstructs from this section every time; it does not remove the remembering, and nothing here does.
By hand it is two lookups. Find the rolling log first — gh issue list --label flow-measurement --state open — rather than assuming a number, and the durable one the same way with --label flow-baseline. Where more than one is open, treat it the way any broken invariant is treated: name it rather than guessing which one is current.
A reading of zero open is two different facts, and one call tells them apart. gh issue list --label flow-measurement --state all says whether the label has any history at all:
- No issue has ever carried it. This whole section is a no-op: nothing
is measured, nothing is posted, nothing fails, and nothing asks. Most installed repositories never create the label, so most segments end here — that is the expected case, not a missed one.
- The label has a history and nothing is open. The log stopped and
nobody reopened it. Name that in the segment's own handover and stop there: opening one is not a session's act. Two sessions finishing segments at the same moment both read zero and both create, and the next release then fails on two or more — the same invariant broken from the other side. The release-time script that rolls the log (.github/scripts/roll_flow_measurement_issue.py, where a repository has one) is where that invariant and its reasoning are written.
`flow-baseline` splits the same way, and for the same reason. Run the --state all lookup against it too. No issue has ever carried the label and this repository simply has no durable ledger: post the segment's own numbers to the rolling log, and leave the cross-version reading in the record that would have cited it. The label has a history and nothing is open, and somebody closed the durable ledger — name that, in the same handover and on the same grounds, and open nothing. Folding the second into the first is what sends a cross-version reading nowhere while everything looks fine.
What a label answers that a milestone cannot, and why tidying one of these issues by hand breaks the next release, are the tracker's own conventions rather than this skill's. docs/issues-and-milestones.md holds them where a repository keeps one.
Where a log is open, do the two steps as part of the segment that just finished, not as a follow-up someone might do later:
- Run
session_cost.py --segmentsagainst the run's transcript. It
walks every segment beside it — a subagent's transcript lives under ~/.claude/projects/<project-dir>/<session-id>/subagents/agent-*.jsonl — joins each to the spawn whose result it opened at, and prints one row per segment: the agent, its own span, calls, tools per turn, mean gap and tokens. One command, where this step used to be one invocation per transcript.
A segment's own span is the number no other mode has. A cycle row's delegated is the spawn call's own interval, and where a harness writes that result on acceptance it reads seconds for an agent that ran twenty minutes. The row here is read from the agent's own file instead.
A resumed agent is one row per slice, and the mode takes the split that used to be done by eye. Its transcript holds several stretches of work in one file, and the idle gap between two of them belongs to neither: read whole, a segment that worked for thirty-seven seconds reports a span over two hours. The cut is the coordinator's own message row. Where a file has an idle gap and no such row, the report prints one row and names the gap rather than leaving it inside a span. The split is the same given the agent's own file: --segments <agent-*.jsonl> prints that file's slices, named by the file, because the spawn that would name them is in the run's transcript and not in view. That is the path a harness's task output hands the orchestrator, so it is the file a session holding a fix pass's result already has.
A §6 line means an agent spawned another agent, which skills/agent-contract/SKILL.md §6 withholds from every agent whatever its own definition says. The mode notices and stops nothing — the evidence is a transcript on the machine that ran the agent, in no commit and on no CI runner, so a report somebody already runs at this boundary is the only place it surfaces. Carry the line into the handover; it is about the agent's conduct rather than its cost, and no other row says it.
Read the counts above the table before the table. The mode prints how many transcripts it walked, how many spawns it found, the tolerance it joined within, and how many went unmatched on each side — even when they agree. A join that silently matched nothing reads exactly like a run that spawned nothing, which is why the numbers are there to be seen holding. Given a resumed agent's own file, nothing is walked or joined, and a header naming the file and its coordinator messages stands in their place.
One segment measured on its own is still `session_cost.py <transcript>` with no mode flag, unless the coordinator restarted it. The plain reading keeps its numbers, and for a file those messages cut into two stretches of work or more it adds one line saying how many coordinator messages the file holds: its span then covers every stretch of work and the waits between them, and --segments <transcript> is the reading that splits it. A row of the per-segment table is that plain reading of another file, or of one stretch of one.
A segment is one agent's own stretch of a chain, and a spawn cycle is not one. Every segment has a transcript of its own: a smith's, a warden's, and the orchestrator's, which is one file however many agents the run spawned. The spawn cycles inside that file are bands over the orchestrator's own minutes, never segments in their own right — and a cycle counted as a segment is how three segment kinds came to have bands a later run can be read against while the most expensive one had none. The two modes follow the distinction: --spawns slices this transcript into those bands, and --segments opens the transcripts of the agents this run spawned, one row each, or reads one agent's own resumed file one stretch of work at a time.
The orchestrator's own boundary is not a user line — `session_cost.py --spawns` is what takes it. Every other segment is measured whole, because its transcript holds nothing but itself; the orchestrator's holds every cycle of the run, so the whole file was the only row it ever had. A spawn cycle is not the review chain's cycle, which docs/review-chain-spec.md owns: it ends when a spawn call's result arrives and begins where the row before it ended, so the head is the framing before the first spawn, cycle N runs from spawn N-1's result to spawn N's, and the tail is the closing work after the last one.
Post a cycle row as a band, and never as an attribution. Inside one window the orchestrator waits on the previous agent, verifies the report it hands over, and frames the next prompt, and no transcript field marks where any of those ends. Cycle 1 is the one row without that window: the run's own start is a boundary a script can take, so its framing goes to the head row instead, and the two are read together.
Read `delegated` before quoting a cycle's model time, because what a spawn's result MEANS is the harness's and not this skill's. That column is the spawn call's own tool_use-to-tool_result span. Where a harness writes the result when the agent FINISHES, the column is the delegated wall clock and it is out of the row's other columns, which is the double count gone. Where a harness writes it when the spawn is ACCEPTED, the column reads seconds and the agent's own wall clock is in none of the row's columns and none of any other row's. It is the gap between one row's last call and the next row's first, and a row's span starts at its own first call while its model time never counts the gap before it. So the rows partition the run's calls and not its wall clock — measured at 12 to 31 per cent of a run — and the agent's own transcript is where its number is, either way.
A span runs from its window's first call to the last call to END, and that is worth knowing because it used to end at the last call to BEGIN (#300). The old rule read the end of the last element of a list sorted by start, so a long-lived call — a background command, a suite spanning the window — ended after the window counting it, and a reading could report a span shorter than a single call inside it. Two consequences for anyone comparing readings across that change: a span taken before it is the shorter of the two wherever a call outlived its window, and identical everywhere else; and command could exceed 100 per cent of the span, which the new rule narrows and does not close, because command time is a sum over calls and calls can run concurrently. Where you see a share above 100, read it as command seconds against wall-clock seconds with calls running at once, never as a broken number. Batching is the ordinary way in and a background command is the rarer one: every tool call in one assistant message carries that message's timestamp as its start, so a batch of three overlaps by construction. Measured over the transcripts on one machine, every second of overlap above a second came from calls batched into one message and none of it from a call that crossed a turn.
A `git` call counts as `git` wherever it runs on the line, and that is worth knowing because it used to count only at the start (#377). The release that carries #377 is the one CHANGELOG.md lists it under, and --segments prints the same warning on the page. The family was read by position, so cd /x && git status — the shape nearly every worktree session writes — was charged to other, and so was a gh call inside a loop or after a leading assignment. It is now read by command word: the first word after any separator, newline or reserved word, never a word inside quotes or inside $( … ). The family also reads the command with its newlines kept and removes a heredoc's body up to its closing line, where it used to drop everything after the operator, so a gh issue create after cat > body.md <<'EOF' and a bin/test after a python3 - <<'EOF' script are read. What a reader comparing readings across that change must know: in a reading taken before it, git and test read low and other reads high by the same calls, the other note may name a command that was a git run, and the two repeats figures, which keep only test, lint/type and build, may read low. Span, command, model, idle, tokens, tools per turn and the slowest list do not move.
A call that only reads counts as `read`, and that is worth knowing because it used to count as `other` (#642). The release that carries #642 is the one CHANGELOG.md lists it under, and --segments prints the same warning on the page. A call is read when every command on it is a read word (sed, grep, rg, cat, head, tail, ls, find, wc, awk, nl, sort, diff) or a word that touches no file (for example cd, echo, test, [[ and the words that close a loop or an if; session_cost.py's NEUTRAL_WORDS is the whole list), at least one reads, nothing writes and nothing is hidden: sed -i, sort -o, find -delete, a redirection into a file, a redirection before a command's first word, a heredoc or a $( … ) keeps it other, and so does any command not on either list. What a sed or awk program writes or runs from inside its quotes is not seen, and such a call is read. It is judged after the other four, so grep -rn pytest stays test and ls && git status stays git. What a reader comparing readings across that change must know: in a reading taken before it there is no read row, other holds those calls, and the other note may fire where it would not now and name a read command. Every other row, the two repeats figures, span, command, model, idle, tokens, tools per turn and the slowest list do not move.
Where the report gives no between-the-rows figure and prints the rows' spans against the run's own, a call outlived the cut its row ends at. Assigning a call by its start is what makes the calls partition, and it leaves a long-running one — a background command, a suite spanning a cut — in the row it began in while the next row has already started, so two rows' spans cover the same seconds. The head row's cut is the first spawn's START and every other row's is a spawn's RESULT, which is why the line names the cut and not the result: a head call can outlive its own row's cut and still end before any spawn's result arrives.
The two figures it prints are not the rows' overlap, so do not subtract them and post the difference as one. A row's span ends at its own last call to end and so does the run's, so every row's interval sits inside the run's and the difference is the gaps between the rows minus their overlap. Quote that run's span and its rows' columns, and leave the between-the-rows share out of the reading rather than substituting either sum.
So a long cycle span is not a long agent run. Where a row's span exceeds its own parts by an hour, that hour is a gap INSIDE the row, above the fifteen minutes model time stops counting at — the orchestrator issuing nothing between two of its own calls. Read it as idle time in the orchestrator, never as the agent it spawned.
- Post what the numbers say.
--postdoes it: your sentence first, the
reading fenced beneath it, on the issue the label resolved to — the rolling log's for the segment's own numbers, the durable one's for a reading that spans versions. By hand it is gh issue comment <n> --body-file <file>, where <n> is the issue number the lookup above returned, and that is the command --post composes.
And write down what ran the segment, in that segment's own record. rounds/round-N.md and phases/phase-N.md both carry a | Ran by | row: the agent and the model, joined by the word on — specseal:smith on <the model it was spawned with>. It goes in beside the numbers because it is the half of the reading the numbers cannot carry. A measurement saying a segment took nineteen minutes and ninety-nine calls answers nothing on its own; the same reading beside what produced it is the first comparison anybody can make.
It is yours to fill and not the segment's, and that is the reason the row exists on this side of the handover rather than in the agent's own report. An agent is told what it is, so a value it writes about itself is the value it was told, and the model is a spawn-time argument you chose — agents/*.md pins none. You are the only party that knows.
Whose row it is and whose keystrokes fill it are different questions, and for phases/phase-N.md they come apart: that record is written by the segment that just ran, not by you. So either hand the value over in the spawn prompt — and the segment transcribes what it was given rather than a value it decided for itself — or fill the row afterwards, the way Fixes checked by is filled. For rounds/round-N.md the question does not arise, because you write that file yourself.
Where you genuinely cannot name a model — a segment spawned through another harness, say — unknown — <why> is the answer and a bare unknown is not, in the shape nobody — <why> already has. An honest blank with a reason beside it is worth more than a confident guess, because a reading nobody can trust reads exactly like one nobody took.
Do this before the next segment spawns, or at the latest together with the round or phase record that closes this one — never later, and never as a question. The destination issue is a log, not a decision: whoever is watching a segment finish posts to it without asking whether to, the same way the rest of this skill's completion claims are not asked about before they are made.
When the run ends, one more reading is taken, and it is the run's. Everything above measures one segment — its span, its calls, its tools per turn. A run is every segment plus the rounds and the commits between them, and no segment's numbers say what the run cost. So the report a finished run hands over — the pull request body's chain section, and the comment that carries the numbers — holds one table, this run beside the last run measured, with the same rows in the same order every time. That is what lets two runs be read side by side without re-deriving either. It is not a third destination: the table goes to the rolling log named above, beside the segment readings it gathers up.
| Row | Taken from | |
|---|---|---|
| Rounds — finding and verifying, and reopenings | the round records | |
| Wall clock, routing commit to last record, build and chain apart | git log of the branch | |
| Commits, by kind | git log --format=%s | |
| Findings by severity | the records' verdict tables | |
Findings by Location: record · code or tests · docs | the records' Location column. A record is anything under seal/specs/, seal/ledger/, seal/releases/ or seal/ledger.md — the work item's own documents, the ledger, its release files and its fragments alike — because a finding in any of them is about the run's own paperwork rather than about the tool. Those four paths, not the whole of seal/: the root also holds config.md, follow-up.md and README.md, which are the repository's own and belong under docs | |
| Records' share of the diff | git diff --numstat against the base, the record paths above counted apart | |
| Model turns · output tokens · cache write · cache read | session_cost.py's token line over the run's main transcript | |
| Segments: count, minutes and tokens per kind | the agent completion notices, or the subagent transcripts | |
| Broad gate: whether it has run, how many times, and at what SHA | the last round record's Broad gate cell, which holds not yet or one entry per full-suite run, newest first, each a SHA with the base it was compared against — the count is the number of entries, read off the cell rather than remembered (#174) |
The tokens are counted, not estimated, and counted the same way every time. session_cost.py <the run's main transcript> prints the token line, which sums output_tokens, cache_creation_input_tokens and cache_read_input_tokens over every usage block in that transcript and in every file under its <session-id>/subagents/ directory. One command, so the row cannot be summed one way this run and another way the next. The line also says how many transcripts it covered, and that count is what a reader checks the number against.
The run's main transcript is the `*.jsonl` sitting directly under the project directory, never one under a <session-id>/subagents/. --latest takes the newest file anywhere beneath the project directory, so on a run that spawned segments it usually lands on a segment rather than on the run — the path it prints as its first line is what says which, and a token line reading 1 transcript for a run that spawned six is the same thing said twice. Give the main transcript by path when that happens.
A comparison against a run whose transcript covered only part of its branch says so beside the number, in the prose under the table rather than as a column. A resumed session, a transcript split at the user lines, a segment spawned through another harness — each of those leaves the token row covering less than the branch, and a smaller number with nothing said about it reads as a run that cost less.
The table carries no verdict. It makes two runs comparable and does nothing else. What a row meant on this branch — why the rounds ran long, why one kind of segment cost what it did — goes in the comment beside the table, never in a column of it.
Counterfeits (stop on sight)
- Output quoted from an earlier run — freshness is per tree state, not per
conversation.
- "Tests pass" for a claim no test covers — green suite, unproven claim.
- Satisfaction vocabulary before the run: should, probably, seems, likely.
- A new test that passed on first run and was never seen red (condition 2).
- Partial evidence generalized — one endpoint checked, "API works" claimed.
- A broad run reported as the sealer's seal when edits followed it — including the
one small fix made after it.
- A pre-existing failure counted as this work's, or waved past without
being named. Both need the base comparison; neither survives it.
- A narrow command that was green because it opened no files — a linter
pointed at a directory holding none of its language, a test path that collected zero tests. Condition 2 does not relax with the scope.
- A check that stopped before reaching the code it was aimed at — a config
error, a collection error, two modules under one name. The run ends early and cheaply, and "it was fast" reads as "it was clean". It is neither a pass nor somebody else's problem: that axis now has no answer, and the honest row is unverified with the cause named. Measured: a mypy . that died in four seconds on duplicate modules, checked neither the source nor the tests, and was the direct reason eleven regressions went out.
- A failure count with no base state behind it. "1 failed" is a claim about
your change only if you know what the tree failed before you touched it.
- A deferral to an answerer nobody resolved. "CI covers it" in a repository
where nothing runs it is the unverified label doing the opposite of its job: it reads as handled.