Repository navigation
Conversation
Stand-up step 3. Both lanes work against the corpus, need no browser and no network, and produce findings checkable in seconds. Both get an instrument first, because a lane woken with "find where the model lies" and no tooling will flail, and a flailing lane produces plausible prose instead of findings. scripts/model-audit.mjs checks ten properties of the mastery model against twenty thousand generated histories each. The suite checks the model against examples someone thought of; this checks it against two hundred thousand nobody did. One seeded generator, no wall clock, no Math.random, so every counterexample replays exactly from the seed reported with it. A finding that cannot be reproduced has given the lane nothing. The first run produced two counterexamples and both were the properties, not the model. One compared a raw prior against a value bktUpdate deliberately clamps, which measures the clamp rather than the model, and the clamp is what keeps BKT numerically stable at the ends. The other judged review timing by absolute retention when nextDue targets retrievability, so it flagged a skill at p_known 0.24 whose decay factor was 0.828, which is exactly right. Both were fixed as properties. All ten now hold, and the record-level checks hold across every persona when the corpus is present. That experience is written into model-auditor's charter, because it is the lane's most likely failure: a wrong property looks precisely like a defect, and correcting the property is a real deliverable rather than a lesser one. src/domain/misconceptions.ts reads the evidence for stable wrong models, which SPEC.md calls the single highest-value thing in the record and which nothing had ever read. One class, honestly: a recurring single-character substitution between what was expected and what the child produced. Narrow, and it covers the case that matters most in early reading, where a child maps one grapheme onto another sound consistently and every downstream skill inherits it. Against the corpus it scores precision 1 and recall 1. It finds the e-for-i substitution in vowel-confuser unaided, eleven occurrences across ten distinct items and two skills, from a record carrying no misconception row. It stays silent on all seven other personas, including stuck-blending, which is the precision test that matters: a detector that reads failure as confusion will find a pattern in any struggling child. The detector never writes to the record and nothing calls it from the tutor's loop. Whether a candidate reaches the tutor, a parent, or the misconception table is a human-in-the-loop decision with a boundary in it. A detector that quietly writes its guesses into a child's permanent record is a different and worse thing than a detector. Writing a convincing negative for a pattern finder is harder than it looks. The first version of the scattered-wrongness test read as random and was not: three of its six misses were the same substitution, and the detector correctly said so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stand-up step 3. Stacked on #2. Also depends on VedSoni-dev#9 for the corpus, which both harnesses degrade honestly without.
Both lanes get an instrument first. A lane woken with "find where the model lies" and no tooling will flail, and a flailing lane produces plausible prose instead of findings.
npm run audit— property checks over the mastery modelTen properties, 20,000 generated histories each. The suite checks the model against examples someone thought of; this checks it against 200,000 nobody did. One seeded generator, no wall clock, no
Math.random, so every counterexample replays exactly from the seed reported with it.The first run produced two counterexamples and both were the properties, not the model.
bktUpdatedeliberately clamps to[0.001, 0.999]. That measures the clamp, and the clamp is what keeps BKT numerically stable at the ends.nextDuetargets retrievability. It flagged a skill atp_known0.24 whose decay factor was 0.828 — exactly right.Both fixed as properties. All ten now hold, and the record-level checks hold across every persona when the corpus is present.
That experience is written into
model-auditor's charter, because it is the lane's most likely failure mode: a wrong property looks precisely like a defect, and correcting the property is a real deliverable rather than a lesser one.npm run mine— a misconception detector, and its scoreSPEC.mdcallsmisconceptionthe single highest-value thing in the record: what a great tutor carries in their head and what every worksheet app throws away. Nothing had ever read the evidence for one.src/domain/misconceptions.tsdetects one class, honestly: a recurring single-character substitution between expected and produced. Narrow, and it covers the case that matters most in early reading, where a child maps one grapheme onto another sound consistently and every downstream skill inherits it.It found the wrong model unaided, from a record carrying no
misconceptionrow — the corpus is deliberately unlabelled. And it stayed silent on all seven other personas includingstuck-blending, which is the precision test that matters: a detector that reads failure as confusion will find a pattern in any struggling child.The asymmetry is the design. A missed misconception leaves the tutor working slightly blind, which is where it already was. A fabricated one has the tutor teaching against a model the child does not hold, and a parent told something confident and false about their kid. One false positive fails the run; low recall fires more quietly.
Two constraints in the charter
misconceptiontable is a human-in-the-loop decision with a boundary in it. A detector that quietly writes its guesses into a child's permanent record is a different and worse thing than a detector. Wiring it up is your call, not the lane's.Checks
npm testpasses, 229 tests (10 new for the detector). Both harnesses skip their corpus-dependent halves and say so whentest/fixtures/personas.tsis absent, rather than reporting a pass they did not earn.Writing a convincing negative for a pattern finder is harder than it looks: the first scattered-wrongness test case read as random and was not — three of its six misses were the same substitution, and the detector correctly called it.
CI cannot run upstream: the account is billing-locked.
🤖 Generated with Claude Code