release: propose maintained v0.1.3 stack - #10
Merged
inquerium merged 8 commits intoAug 5, 2026
Merged
Conversation
…inst The pipeline was four agents announcing to one chat. That does not survive being scaled up. Every lane added lengthens the review queue, because agents never merge and a named human admits every change, so more proposers without a filter is more work for the one scarce resource rather than less. Three things, then: A second class of lane. World lanes still write only under world/. Engineering lanes examine the machinery that reads that content and may write src/, under a harder gate than the World lanes carry rather than a softer one: never src/curriculum/, never src/mcp/tools.ts, never test/invariants.test.ts, full suite green, smallest diff that fixes the thing, and every behavior change pinned by a test that fails without it. An agent editing the test that constrains agents is the whole failure in one diff. An adjutant. One agent, no lane, and the only route to the maintainer. It holds no write and no edit, cannot open or merge a pull request, and produces one brief a day. Every other lane hands its deliverable to the adjutant and decides nothing about whether a person sees it. Fifteen lanes announcing to one chat produces a muted chat, and a muted channel is a pipeline that has stopped working without saying so. A roster that is a file. automation/corps.mjs states which agents exist, what each may not touch, and what wakes them. scripts/corps.mjs reconciles it against the running gateway and refuses to create agents or set tool policy, because those grant authority and authority is granted by a person who knows they are doing it. Run against the live gateway it immediately found both existing jobs registered without a timeout, which docs/OPENCLAW.md requires and which is the difference between a run that fails and one that hangs. automation/deliverable.mjs is the part of the escalation doctrine that is not advisory. Prose in a charter is a prompt, and a prompt is advisory. The gate names what a run handed back, so a reply that asks the operator to decide something is bounced rather than forwarded, and so a lane that does it habitually shows up in the numbers. It deliberately does not flag a bare question mark: the charter asks agents to say what would make them wrong, and punishing that would train the reports to be less candid. The rate is the point. A lane blocked in more than a third of its runs has been given a task it cannot finish with the authority it holds. The fix is to change the task. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Stand-up steps 1 and 2 from docs/CORPS.md, in that order and for the stated reason: the filter goes in before the roster grows, because building it after ten lanes means learning why it was needed by living through a week without it. Both watchers are tested against stubs rather than by forcing runs. A forced `openclaw cron run` executes the payload without ever evaluating the trigger, so it cannot verify a watcher at all. The stub harness caught a real bug in the adjutant watcher on the first pass: the seen-set was built only from the scanner's rows, so any escalation reported outside that array would have been briefed every morning until someone muted the channel. scripts/adjutant-scan.mjs reads run history, names the shape of every deliverable, and reports per-lane conduct. The watcher shells out to it rather than classifying inline, because watcher scripts cannot import and a second copy of the shape gate would drift from the first inside a month, silently, with both still returning plausible answers. The gate itself was wrong, and real run history is what proved it. Twelve production runs were classified, and four were classified badly: Two exec denials on 2026-07-30 came back "declined to act, a complete outcome". Both finished status ok and delivered. The lane had done exactly the right thing, refusing to fabricate a pull request it could not open and saying "that is a pipeline breakage, not an empty week". The gate read the phrase "No PR opened" and filed it as a quiet week. That is the precise silent failure this whole design exists to prevent, and it was in the detector. A complete attack run that resolved seven DOIs and posted its findings came back malformed, because the attacker lanes hold no write tool by design and their deliverable is a comment, which had no shape. Seven runs lost to one offline stretch were charged to the lane's conduct rather than to the machine. classify now takes the run record, so a run that never got far enough to have an opinion is `failed`, and a lane is judged only on the runs that answered. A lane failing half its runs has no conduct to assess, it has a machine problem, and the verdict now says so. test/invariants.test.ts refused the first version of the fixture watcher, which listed the record's module paths so it could tell when the model had moved. The rule is blunt on purpose and the charter says the invariant wins, so the list moved to test/fixtures/GOVERNS and the watcher reads it at evaluation time. The corpus is the thing making the claim, so the scope of that claim belongs beside it, and now moves when fixture-keeper changes what it covers. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Stand-up step 3. Both lanes work against the corpus, need no browser and no network, and produce findings checkable in seconds. Both get an instrument first, because a lane woken with "find where the model lies" and no tooling will flail, and a flailing lane produces plausible prose instead of findings. scripts/model-audit.mjs checks ten properties of the mastery model against twenty thousand generated histories each. The suite checks the model against examples someone thought of; this checks it against two hundred thousand nobody did. One seeded generator, no wall clock, no Math.random, so every counterexample replays exactly from the seed reported with it. A finding that cannot be reproduced has given the lane nothing. The first run produced two counterexamples and both were the properties, not the model. One compared a raw prior against a value bktUpdate deliberately clamps, which measures the clamp rather than the model, and the clamp is what keeps BKT numerically stable at the ends. The other judged review timing by absolute retention when nextDue targets retrievability, so it flagged a skill at p_known 0.24 whose decay factor was 0.828, which is exactly right. Both were fixed as properties. All ten now hold, and the record-level checks hold across every persona when the corpus is present. That experience is written into model-auditor's charter, because it is the lane's most likely failure: a wrong property looks precisely like a defect, and correcting the property is a real deliverable rather than a lesser one. src/domain/misconceptions.ts reads the evidence for stable wrong models, which SPEC.md calls the single highest-value thing in the record and which nothing had ever read. One class, honestly: a recurring single-character substitution between what was expected and what the child produced. Narrow, and it covers the case that matters most in early reading, where a child maps one grapheme onto another sound consistently and every downstream skill inherits it. Against the corpus it scores precision 1 and recall 1. It finds the e-for-i substitution in vowel-confuser unaided, eleven occurrences across ten distinct items and two skills, from a record carrying no misconception row. It stays silent on all seven other personas, including stuck-blending, which is the precision test that matters: a detector that reads failure as confusion will find a pattern in any struggling child. The detector never writes to the record and nothing calls it from the tutor's loop. Whether a candidate reaches the tutor, a parent, or the misconception table is a human-in-the-loop decision with a boundary in it. A detector that quietly writes its guesses into a child's permanent record is a different and worse thing than a detector. Writing a convincing negative for a pattern finder is harder than it looks. The first version of the scattered-wrongness test read as random and was not: three of its six misses were the same substitution, and the detector correctly said so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SPEC.md has always said every generated interface must honor every active accommodation, and called it a hard constraint rather than a hint. It was a hint. validateInterface took a string of HTML and never learned who the activity was for, so the only thing between a child who cannot be timed and a countdown was the tutor remembering. test/fixtures.test.ts pinned that gap by asserting the function took one argument. It now takes the child's standing instructions, and a violation refuses the save the same way a syntax error does: the tutor is told what it broke and retries, and a child never meets the version that ignored them. Four kinds are checked mechanically. Countdowns, typography, whether anything on the page can be heard, and unguarded motion. The rule that governs the rest of it: a check that cannot be performed does not pass. It reports that it could not be performed. Silent success on an accommodation nobody wrote a check for is worse than having no checker at all, because a reviewer reading a clean report reasonably concludes the activity was verified. Every unknown kind, and every kind that cannot be settled from source text, comes back unverifiable and appears in the parent's review under a heading that says so. Contrast is the honest example. It needs the page rendered and the computed colours compared, so it is never claimed either way, and inventing a proxy for it would be worse than leaving it to a person. The audit caught a defect in the checker on its first run, which is the reason that instrument exists. Accommodation kinds are free text written by an adult about their own child, and "no_read_aloud_pressure" matched the audio check, which then refused every page that did not speak. That is the precise inversion of what the parent asked for. A check that can invert an instruction is worse than no check, so that phrasing now falls through to unverifiable, and the rule is written into the lane's charter. Coverage is five of ten kinds and is reported as five of ten. The other five rest on an adult reading the activity, which is the truth and is the marshal's backlog rather than a number to round up. This touches src/mcp/tools.ts, which the engineering-lane charter forbids its agents from editing. The rule is about widening what the unattended tutor can reach. This narrows it: save_interface now refuses work it previously accepted. Worth naming rather than letting a reader discover the exception on their own. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Stand-up step 5, the attacker half. child-sim needs to execute a page and is not in this change; see below. Rule 5 in CLAUDE.md forbids streaks, coins, leaderboards and countdown timers for children. test/invariants.test.ts pinned that both prompts say so, and its own comment describes prompt text as "the soft layer above the structural enforcement". For engagement mechanics there was no structural layer under it. The tutor was asked not to, and that was the whole of it. A streak counter now refuses the save. The second thing this reads is quieter and costs more. An activity that records correct without recording what the child actually said makes every misconception in that session permanently undiscoverable. src/domain/misconceptions.ts reads response against expected; with correctness alone the evidence is a column of zeroes. SPEC.md's central claim is that you can always recompute with a better model, and that claim is false for any session recorded this way, because a better model still has nothing to read. That one does not block, and the reason is the most useful thing in this diff. A refusal is a demand made of a model that will satisfy it the cheapest way it can find. For a streak counter the cheapest fix is deleting the streak counter, which is what rule 5 wants, so refusing works. For a missing response the cheapest fix is passing response: "x". That satisfies the check and is far worse than the gap it closes: an empty column can be recognised as empty, while a column of plausible fabrications will be mined for misconceptions and will yield them. A gate that can be cheaply faked does not protect the data. It corrupts it and reports success. So the line between refusing and flagging is not severity and not confidence. It is whether the cheapest way to satisfy the check is also the correct fix. That rule is now in the lane's charter and every new check has to answer it. Four of the nine corpus samples must come back clean and they are the ones that took the care: a streak named only in a comment, "point to the picture", a decorative star, and a good activity that adapts on two misses and stops on three. Precision 1, recall 1. A reviewer nobody has proved will pass good work is a reviewer the tutor learns to route around, and once a model learns a gate is noise, every later check inherits the distrust. One existing assertion moved. test/validate.test.ts checked that an aliased runtime produces no warnings at all, which was a proxy for "no spurious instrumentation warnings" written when instrumentation was the only thing warned about. That fixture is six lines and legitimately neither adapts nor stops early. The assertion now says what it means. child-sim is not here. It has to run a page rather than read one, and doing that honestly needs a browser or a DOM faithful enough that "the activity works" is a claim rather than a hope. A shim that half-executes generated HTML would report passes it did not earn, which is the failure mode this session has already found three times in its own harnesses. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
inquerium
marked this pull request as ready for review
August 5, 2026 21:58
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This proposes the maintained seven-commit stack on current VedSoni-dev/primer master.
It includes the macOS v0.1.3 packaging fixes, the corps and accommodation work, and a fixture assertion reconciled with the enforced accommodation validator.
No release tag or GitHub Release has been created.
Validation
This is a proposal for maintainer review.