Skip to content

corps: stand up the adjutant and the fixture keeper - #2

Open
inquerium wants to merge 1 commit into
corps-doctrinefrom
corps-first-wave
Open

inquerium wants to merge 1 commit into
corps-doctrinefrom
corps-first-wave

Conversation

@inquerium

Copy link
Copy Markdown
Owner

Stand-up steps 1 and 2 from docs/CORPS.md. Stacked on #1, which is stacked on VedSoni-dev#7. Also depends on VedSoni-dev#9 for the corpus itself: npm run corps:doctor correctly reports test/fixtures/personas.ts missing on this branch, and that resolves when VedSoni-dev#9 lands.

The adjutant goes in before the roster grows. Building the filter after ten lanes means learning why it was needed by living through a week without it.

The gate was wrong, and production runs proved it

I ran the shape gate against the twelve real runs this pipeline has produced. Four were classified badly, and one of those was serious.

Two exec denials on 2026-07-30 came back "declined to act, a complete outcome." Both finished status: ok and delivered. The lane had done exactly the right thing: it refused to fabricate a pull request it could not open, and said "that is a pipeline breakage, not an empty week." The gate read the phrase "No PR opened" and filed it as a quiet week.

That is the precise silent failure this design exists to prevent, sitting inside the detector meant to prevent it. BLOCKED is now checked before every completion shape and carries the phrasings lanes actually use: hard-denied, allowlist miss, requires exec approval, pipeline breakage, is not honoring it.

A complete attack run came back malformed. It resolved seven DOIs, checked each against the source abstract, and posted its findings. The attacker lanes hold no write tool by design, so their deliverable is a comment and never a branch — and there was no shape for that. Added posted.

Seven runs lost to one offline stretch were charged to the lane's conduct. classify now takes the run record. A run that never got far enough to have an opinion is failed, and a lane is judged only on the runs that answered. A lane failing half its runs has no conduct to assess, it has a machine problem, and the verdict says so.

All twelve now classify correctly. Every one of these is a test case, verbatim.

What the scan says about the live pipeline

 !  world-attacker   11 runs  fail 0.73  esc 0     the machine, not the lane
 !  world-research    4 runs  fail 0     esc 0.5   the task is too big or the capability is missing

Both diagnoses are correct. The attacker's failures are the documented offline stretch. The research lane's escalations are the exec denials, which were fixed the same day — tools.exec.mode is now full and the 08-02 attacker run posted a comment through gh without trouble.

Watchers are tested against stubs

A forced openclaw cron run executes the payload without evaluating the trigger, so it cannot verify a watcher. Both watchers are exercised through an AsyncFunction harness with tools/trigger/json supplied: cold start, quiet, something to say, already said, transient failure, sustained failure, recovery, and wrong-shape response.

That harness earned its place immediately. The adjutant watcher's seen-set was built only from the scanner's rows, so any escalation reported outside that array would have been briefed every morning until someone muted the channel — the exact failure the adjutant exists to prevent. A forced run would never have found it.

An invariant refused the first fixture watcher

test/invariants.test.ts forbids anything under automation/ from naming the record's modules. The first version of fixture-drift.js listed src/db/ and src/record/ so it could tell when the model had moved, and the invariant rejected it.

The rule is blunt on purpose and AGENTS.md says the invariant wins, so the list moved to test/fixtures/GOVERNS and the watcher reads it at evaluation time. This is better than what it replaced: the corpus is the thing making the claim, so the scope of that claim belongs beside it and now moves when fixture-keeper changes what it covers.

Checks

npm test passes, 219 tests. npm run corps:doctor passes everything except the corpus, which lives on VedSoni-dev#9.

CI cannot run upstream: the account is billing-locked.

🤖 Generated with Claude Code

Stand-up steps 1 and 2 from docs/CORPS.md, in that order and for the stated
reason: the filter goes in before the roster grows, because building it after
ten lanes means learning why it was needed by living through a week without it.

Both watchers are tested against stubs rather than by forcing runs. A forced
`openclaw cron run` executes the payload without ever evaluating the trigger,
so it cannot verify a watcher at all. The stub harness caught a real bug in the
adjutant watcher on the first pass: the seen-set was built only from the
scanner's rows, so any escalation reported outside that array would have been
briefed every morning until someone muted the channel.

scripts/adjutant-scan.mjs reads run history, names the shape of every
deliverable, and reports per-lane conduct. The watcher shells out to it rather
than classifying inline, because watcher scripts cannot import and a second
copy of the shape gate would drift from the first inside a month, silently,
with both still returning plausible answers.

The gate itself was wrong, and real run history is what proved it. Twelve
production runs were classified, and four were classified badly:

Two exec denials on 2026-07-30 came back "declined to act, a complete outcome".
Both finished status ok and delivered. The lane had done exactly the right
thing, refusing to fabricate a pull request it could not open and saying "that
is a pipeline breakage, not an empty week". The gate read the phrase "No PR
opened" and filed it as a quiet week. That is the precise silent failure this
whole design exists to prevent, and it was in the detector.

A complete attack run that resolved seven DOIs and posted its findings came
back malformed, because the attacker lanes hold no write tool by design and
their deliverable is a comment, which had no shape.

Seven runs lost to one offline stretch were charged to the lane's conduct
rather than to the machine. classify now takes the run record, so a run that
never got far enough to have an opinion is `failed`, and a lane is judged only
on the runs that answered. A lane failing half its runs has no conduct to
assess, it has a machine problem, and the verdict now says so.

test/invariants.test.ts refused the first version of the fixture watcher, which
listed the record's module paths so it could tell when the model had moved.
The rule is blunt on purpose and the charter says the invariant wins, so the
list moved to test/fixtures/GOVERNS and the watcher reads it at evaluation
time. The corpus is the thing making the claim, so the scope of that claim
belongs beside it, and now moves when fixture-keeper changes what it covers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant