Skip to content

release: propose maintained v0.1.3 stack - #10

Merged
inquerium merged 8 commits into
VedSoni-dev:masterfrom
inquerium:agent/upstream-release-v013
Aug 5, 2026
Merged

inquerium merged 8 commits into
VedSoni-dev:masterfrom
inquerium:agent/upstream-release-v013

Conversation

@inquerium

Copy link
Copy Markdown
Collaborator

Summary

This proposes the maintained seven-commit stack on current VedSoni-dev/primer master.

It includes the macOS v0.1.3 packaging fixes, the corps and accommodation work, and a fixture assertion reconciled with the enforced accommodation validator.

No release tag or GitHub Release has been created.

Validation

  • npm test (264 passing)
  • macOS arm64 standalone package smoke test with official Node 24

This is a proposal for maintainer review.

inquerium and others added 8 commits August 5, 2026 16:48
…inst

The pipeline was four agents announcing to one chat. That does not survive
being scaled up. Every lane added lengthens the review queue, because agents
never merge and a named human admits every change, so more proposers without
a filter is more work for the one scarce resource rather than less.

Three things, then:

A second class of lane. World lanes still write only under world/. Engineering
lanes examine the machinery that reads that content and may write src/, under
a harder gate than the World lanes carry rather than a softer one: never
src/curriculum/, never src/mcp/tools.ts, never test/invariants.test.ts, full
suite green, smallest diff that fixes the thing, and every behavior change
pinned by a test that fails without it. An agent editing the test that
constrains agents is the whole failure in one diff.

An adjutant. One agent, no lane, and the only route to the maintainer. It
holds no write and no edit, cannot open or merge a pull request, and produces
one brief a day. Every other lane hands its deliverable to the adjutant and
decides nothing about whether a person sees it. Fifteen lanes announcing to
one chat produces a muted chat, and a muted channel is a pipeline that has
stopped working without saying so.

A roster that is a file. automation/corps.mjs states which agents exist, what
each may not touch, and what wakes them. scripts/corps.mjs reconciles it
against the running gateway and refuses to create agents or set tool policy,
because those grant authority and authority is granted by a person who knows
they are doing it. Run against the live gateway it immediately found both
existing jobs registered without a timeout, which docs/OPENCLAW.md requires
and which is the difference between a run that fails and one that hangs.

automation/deliverable.mjs is the part of the escalation doctrine that is not
advisory. Prose in a charter is a prompt, and a prompt is advisory. The gate
names what a run handed back, so a reply that asks the operator to decide
something is bounced rather than forwarded, and so a lane that does it
habitually shows up in the numbers. It deliberately does not flag a bare
question mark: the charter asks agents to say what would make them wrong, and
punishing that would train the reports to be less candid.

The rate is the point. A lane blocked in more than a third of its runs has
been given a task it cannot finish with the authority it holds. The fix is to
change the task.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Stand-up steps 1 and 2 from docs/CORPS.md, in that order and for the stated
reason: the filter goes in before the roster grows, because building it after
ten lanes means learning why it was needed by living through a week without it.

Both watchers are tested against stubs rather than by forcing runs. A forced
`openclaw cron run` executes the payload without ever evaluating the trigger,
so it cannot verify a watcher at all. The stub harness caught a real bug in the
adjutant watcher on the first pass: the seen-set was built only from the
scanner's rows, so any escalation reported outside that array would have been
briefed every morning until someone muted the channel.

scripts/adjutant-scan.mjs reads run history, names the shape of every
deliverable, and reports per-lane conduct. The watcher shells out to it rather
than classifying inline, because watcher scripts cannot import and a second
copy of the shape gate would drift from the first inside a month, silently,
with both still returning plausible answers.

The gate itself was wrong, and real run history is what proved it. Twelve
production runs were classified, and four were classified badly:

Two exec denials on 2026-07-30 came back "declined to act, a complete outcome".
Both finished status ok and delivered. The lane had done exactly the right
thing, refusing to fabricate a pull request it could not open and saying "that
is a pipeline breakage, not an empty week". The gate read the phrase "No PR
opened" and filed it as a quiet week. That is the precise silent failure this
whole design exists to prevent, and it was in the detector.

A complete attack run that resolved seven DOIs and posted its findings came
back malformed, because the attacker lanes hold no write tool by design and
their deliverable is a comment, which had no shape.

Seven runs lost to one offline stretch were charged to the lane's conduct
rather than to the machine. classify now takes the run record, so a run that
never got far enough to have an opinion is `failed`, and a lane is judged only
on the runs that answered. A lane failing half its runs has no conduct to
assess, it has a machine problem, and the verdict now says so.

test/invariants.test.ts refused the first version of the fixture watcher, which
listed the record's module paths so it could tell when the model had moved.
The rule is blunt on purpose and the charter says the invariant wins, so the
list moved to test/fixtures/GOVERNS and the watcher reads it at evaluation
time. The corpus is the thing making the claim, so the scope of that claim
belongs beside it, and now moves when fixture-keeper changes what it covers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Stand-up step 3. Both lanes work against the corpus, need no browser and no
network, and produce findings checkable in seconds. Both get an instrument
first, because a lane woken with "find where the model lies" and no tooling
will flail, and a flailing lane produces plausible prose instead of findings.

scripts/model-audit.mjs checks ten properties of the mastery model against
twenty thousand generated histories each. The suite checks the model against
examples someone thought of; this checks it against two hundred thousand nobody
did. One seeded generator, no wall clock, no Math.random, so every
counterexample replays exactly from the seed reported with it. A finding that
cannot be reproduced has given the lane nothing.

The first run produced two counterexamples and both were the properties, not
the model. One compared a raw prior against a value bktUpdate deliberately
clamps, which measures the clamp rather than the model, and the clamp is what
keeps BKT numerically stable at the ends. The other judged review timing by
absolute retention when nextDue targets retrievability, so it flagged a skill
at p_known 0.24 whose decay factor was 0.828, which is exactly right. Both were
fixed as properties. All ten now hold, and the record-level checks hold across
every persona when the corpus is present.

That experience is written into model-auditor's charter, because it is the
lane's most likely failure: a wrong property looks precisely like a defect, and
correcting the property is a real deliverable rather than a lesser one.

src/domain/misconceptions.ts reads the evidence for stable wrong models, which
SPEC.md calls the single highest-value thing in the record and which nothing
had ever read. One class, honestly: a recurring single-character substitution
between what was expected and what the child produced. Narrow, and it covers
the case that matters most in early reading, where a child maps one grapheme
onto another sound consistently and every downstream skill inherits it.

Against the corpus it scores precision 1 and recall 1. It finds the e-for-i
substitution in vowel-confuser unaided, eleven occurrences across ten distinct
items and two skills, from a record carrying no misconception row. It stays
silent on all seven other personas, including stuck-blending, which is the
precision test that matters: a detector that reads failure as confusion will
find a pattern in any struggling child.

The detector never writes to the record and nothing calls it from the tutor's
loop. Whether a candidate reaches the tutor, a parent, or the misconception
table is a human-in-the-loop decision with a boundary in it. A detector that
quietly writes its guesses into a child's permanent record is a different and
worse thing than a detector.

Writing a convincing negative for a pattern finder is harder than it looks. The
first version of the scattered-wrongness test read as random and was not: three
of its six misses were the same substitution, and the detector correctly said
so.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SPEC.md has always said every generated interface must honor every active
accommodation, and called it a hard constraint rather than a hint. It was a
hint. validateInterface took a string of HTML and never learned who the
activity was for, so the only thing between a child who cannot be timed and a
countdown was the tutor remembering. test/fixtures.test.ts pinned that gap by
asserting the function took one argument.

It now takes the child's standing instructions, and a violation refuses the
save the same way a syntax error does: the tutor is told what it broke and
retries, and a child never meets the version that ignored them. Four kinds are
checked mechanically. Countdowns, typography, whether anything on the page can
be heard, and unguarded motion.

The rule that governs the rest of it: a check that cannot be performed does not
pass. It reports that it could not be performed. Silent success on an
accommodation nobody wrote a check for is worse than having no checker at all,
because a reviewer reading a clean report reasonably concludes the activity was
verified. Every unknown kind, and every kind that cannot be settled from source
text, comes back unverifiable and appears in the parent's review under a
heading that says so.

Contrast is the honest example. It needs the page rendered and the computed
colours compared, so it is never claimed either way, and inventing a proxy for
it would be worse than leaving it to a person.

The audit caught a defect in the checker on its first run, which is the reason
that instrument exists. Accommodation kinds are free text written by an adult
about their own child, and "no_read_aloud_pressure" matched the audio check,
which then refused every page that did not speak. That is the precise inversion
of what the parent asked for. A check that can invert an instruction is worse
than no check, so that phrasing now falls through to unverifiable, and the rule
is written into the lane's charter.

Coverage is five of ten kinds and is reported as five of ten. The other five
rest on an adult reading the activity, which is the truth and is the marshal's
backlog rather than a number to round up.

This touches src/mcp/tools.ts, which the engineering-lane charter forbids its
agents from editing. The rule is about widening what the unattended tutor can
reach. This narrows it: save_interface now refuses work it previously accepted.
Worth naming rather than letting a reader discover the exception on their own.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Stand-up step 5, the attacker half. child-sim needs to execute a page and is
not in this change; see below.

Rule 5 in CLAUDE.md forbids streaks, coins, leaderboards and countdown timers
for children. test/invariants.test.ts pinned that both prompts say so, and its
own comment describes prompt text as "the soft layer above the structural
enforcement". For engagement mechanics there was no structural layer under it.
The tutor was asked not to, and that was the whole of it. A streak counter now
refuses the save.

The second thing this reads is quieter and costs more. An activity that records
correct without recording what the child actually said makes every misconception
in that session permanently undiscoverable. src/domain/misconceptions.ts reads
response against expected; with correctness alone the evidence is a column of
zeroes. SPEC.md's central claim is that you can always recompute with a better
model, and that claim is false for any session recorded this way, because a
better model still has nothing to read.

That one does not block, and the reason is the most useful thing in this diff.

A refusal is a demand made of a model that will satisfy it the cheapest way it
can find. For a streak counter the cheapest fix is deleting the streak counter,
which is what rule 5 wants, so refusing works. For a missing response the
cheapest fix is passing response: "x". That satisfies the check and is far worse
than the gap it closes: an empty column can be recognised as empty, while a
column of plausible fabrications will be mined for misconceptions and will yield
them. A gate that can be cheaply faked does not protect the data. It corrupts it
and reports success.

So the line between refusing and flagging is not severity and not confidence. It
is whether the cheapest way to satisfy the check is also the correct fix. That
rule is now in the lane's charter and every new check has to answer it.

Four of the nine corpus samples must come back clean and they are the ones that
took the care: a streak named only in a comment, "point to the picture", a
decorative star, and a good activity that adapts on two misses and stops on
three. Precision 1, recall 1. A reviewer nobody has proved will pass good work
is a reviewer the tutor learns to route around, and once a model learns a gate
is noise, every later check inherits the distrust.

One existing assertion moved. test/validate.test.ts checked that an aliased
runtime produces no warnings at all, which was a proxy for "no spurious
instrumentation warnings" written when instrumentation was the only thing
warned about. That fixture is six lines and legitimately neither adapts nor
stops early. The assertion now says what it means.

child-sim is not here. It has to run a page rather than read one, and doing that
honestly needs a browser or a DOM faithful enough that "the activity works" is a
claim rather than a hope. A shim that half-executes generated HTML would report
passes it did not earn, which is the failure mode this session has already found
three times in its own harnesses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@inquerium
inquerium marked this pull request as ready for review August 5, 2026 21:58
@inquerium
inquerium merged commit a2df6ed into VedSoni-dev:master Aug 5, 2026
1 of 11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant