Skip to content

corps: stand up the model auditor and the misconception miner - #3

Open
inquerium wants to merge 1 commit into
corps-first-wavefrom
corps-step-3
Open

inquerium wants to merge 1 commit into
corps-first-wavefrom
corps-step-3

Conversation

@inquerium

Copy link
Copy Markdown
Owner

Stand-up step 3. Stacked on #2. Also depends on VedSoni-dev#9 for the corpus, which both harnesses degrade honestly without.

Both lanes get an instrument first. A lane woken with "find where the model lies" and no tooling will flail, and a flailing lane produces plausible prose instead of findings.

npm run audit — property checks over the mastery model

Ten properties, 20,000 generated histories each. The suite checks the model against examples someone thought of; this checks it against 200,000 nobody did. One seeded generator, no wall clock, no Math.random, so every counterexample replays exactly from the seed reported with it.

The first run produced two counterexamples and both were the properties, not the model.

  • One compared a raw prior against a value bktUpdate deliberately clamps to [0.001, 0.999]. That measures the clamp, and the clamp is what keeps BKT numerically stable at the ends.
  • One judged review timing by absolute retention when nextDue targets retrievability. It flagged a skill at p_known 0.24 whose decay factor was 0.828 — exactly right.

Both fixed as properties. All ten now hold, and the record-level checks hold across every persona when the corpus is present.

That experience is written into model-auditor's charter, because it is the lane's most likely failure mode: a wrong property looks precisely like a defect, and correcting the property is a real deliverable rather than a lesser one.

npm run mine — a misconception detector, and its score

SPEC.md calls misconception the single highest-value thing in the record: what a great tutor carries in their head and what every worksheet app throws away. Nothing had ever read the evidence for one.

src/domain/misconceptions.ts detects one class, honestly: a recurring single-character substitution between expected and produced. Narrow, and it covers the case that matters most in early reading, where a child maps one grapheme onto another sound consistently and every downstream skill inherits it.

precision 1   recall 1   (1 found, 0 fabricated, 0 missed)

ok    vowel-confuser: Reads "e" as "i" (bed read as bid, pen read as pin, den read as din)
      11 times, 10 distinct items, across pa_phoneme_isolate_medial + ph_cvc_short_e
      explains 100% of readable errors, confidence 0.9

It found the wrong model unaided, from a record carrying no misconception row — the corpus is deliberately unlabelled. And it stayed silent on all seven other personas including stuck-blending, which is the precision test that matters: a detector that reads failure as confusion will find a pattern in any struggling child.

The asymmetry is the design. A missed misconception leaves the tutor working slightly blind, which is where it already was. A fabricated one has the tutor teaching against a model the child does not hold, and a parent told something confident and false about their kid. One false positive fails the run; low recall fires more quietly.

Two constraints in the charter

  • The detector never writes to the record, and nothing calls it from the tutor's loop. Whether a candidate reaches the tutor, a parent, or the misconception table is a human-in-the-loop decision with a boundary in it. A detector that quietly writes its guesses into a child's permanent record is a different and worse thing than a detector. Wiring it up is your call, not the lane's.
  • The corpus must never learn what is hunting it. Ground truth lives in the scoring harness, not in the personas.

Checks

npm test passes, 229 tests (10 new for the detector). Both harnesses skip their corpus-dependent halves and say so when test/fixtures/personas.ts is absent, rather than reporting a pass they did not earn.

Writing a convincing negative for a pattern finder is harder than it looks: the first scattered-wrongness test case read as random and was not — three of its six misses were the same substitution, and the detector correctly called it.

CI cannot run upstream: the account is billing-locked.

🤖 Generated with Claude Code

Stand-up step 3. Both lanes work against the corpus, need no browser and no
network, and produce findings checkable in seconds. Both get an instrument
first, because a lane woken with "find where the model lies" and no tooling
will flail, and a flailing lane produces plausible prose instead of findings.

scripts/model-audit.mjs checks ten properties of the mastery model against
twenty thousand generated histories each. The suite checks the model against
examples someone thought of; this checks it against two hundred thousand nobody
did. One seeded generator, no wall clock, no Math.random, so every
counterexample replays exactly from the seed reported with it. A finding that
cannot be reproduced has given the lane nothing.

The first run produced two counterexamples and both were the properties, not
the model. One compared a raw prior against a value bktUpdate deliberately
clamps, which measures the clamp rather than the model, and the clamp is what
keeps BKT numerically stable at the ends. The other judged review timing by
absolute retention when nextDue targets retrievability, so it flagged a skill
at p_known 0.24 whose decay factor was 0.828, which is exactly right. Both were
fixed as properties. All ten now hold, and the record-level checks hold across
every persona when the corpus is present.

That experience is written into model-auditor's charter, because it is the
lane's most likely failure: a wrong property looks precisely like a defect, and
correcting the property is a real deliverable rather than a lesser one.

src/domain/misconceptions.ts reads the evidence for stable wrong models, which
SPEC.md calls the single highest-value thing in the record and which nothing
had ever read. One class, honestly: a recurring single-character substitution
between what was expected and what the child produced. Narrow, and it covers
the case that matters most in early reading, where a child maps one grapheme
onto another sound consistently and every downstream skill inherits it.

Against the corpus it scores precision 1 and recall 1. It finds the e-for-i
substitution in vowel-confuser unaided, eleven occurrences across ten distinct
items and two skills, from a record carrying no misconception row. It stays
silent on all seven other personas, including stuck-blending, which is the
precision test that matters: a detector that reads failure as confusion will
find a pattern in any struggling child.

The detector never writes to the record and nothing calls it from the tutor's
loop. Whether a candidate reaches the tutor, a parent, or the misconception
table is a human-in-the-loop decision with a boundary in it. A detector that
quietly writes its guesses into a child's permanent record is a different and
worse thing than a detector.

Writing a convincing negative for a pattern finder is harder than it looks. The
first version of the scattered-wrongness test read as random and was not: three
of its six misses were the same substitution, and the detector correctly said
so.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant