Skip to content

activities: enforce rule 5, and read what a page does to a child - #5

Open
inquerium wants to merge 1 commit into
corps-step-4from
corps-step-5
Open

inquerium wants to merge 1 commit into
corps-step-4from
corps-step-5

Conversation

@inquerium

Copy link
Copy Markdown
Owner

Stand-up step 5, the activity-attacker half. Stacked on #4. child-sim is deliberately not here — see the last section.

Rule 5 was enforced by prompt text only

CLAUDE.md rule 5 forbids streaks, coins, leaderboards and countdown timers for children. test/invariants.test.ts pinned that both prompts say so — and its own comment describes prompt text as "the soft layer above the structural enforcement". For engagement mechanics there was no structural layer underneath. The tutor was asked not to, and that was the whole of it.

A streak counter now refuses the save.

The quieter one costs more

An activity that records correct without recording what the child actually said makes every misconception in that session permanently undiscoverable. src/domain/misconceptions.ts reads response against expected; with correctness alone the evidence is a column of zeroes.

SPEC.md's central claim is that you can always recompute with a better model. That claim is false for any session recorded this way, because a better model still has nothing to read.

Why that one flags instead of blocking

This is the most useful thing in the diff, and it came out of the tests failing.

A refusal is a demand made of a model that will satisfy it the cheapest way it can find. For a streak counter, the cheapest fix is deleting the streak counter — exactly what rule 5 wants. Refusing works.

For a missing response, the cheapest fix is response: "x". That satisfies the check and is far worse than the gap it closes: an empty column can be recognised as empty, while a column of plausible fabrications will be mined for misconceptions and will yield them.

A gate that can be cheaply faked does not protect the data. It corrupts it and reports success.

So the line between refusing and flagging is not severity, and not confidence. It is whether the cheapest way to satisfy the check is also the correct fix. That rule is now in the lane's charter, and every new check has to answer it.

Scored against declared ground truth

precision 1   recall 1   (6 found, 0 spurious, 0 missed)

ok    good                     (clean)
ok    streak-counter           engagement_mechanics
ok    coins                    engagement_mechanics
ok    no-response-text         no_response_recorded
ok    no-adaptation            no_visible_adaptation, runs_to_a_fixed_length
ok    punishing                punishing_language
ok    commented-out-streak     (clean)
ok    points-to-the-picture    (clean)
ok    gold-star-decoration     (clean)

Four of the nine must come back clean, and those are the ones that took the care: a streak named only in a comment, "point to the picture" (not a points total), a decorative star (awards nothing), and a good activity that adapts on two misses and stops on three.

A reviewer nobody has proved will pass good work is a reviewer the tutor learns to route around — and once a model learns a gate is noise, every later check inherits the distrust. The watcher fires hardest on a spurious refusal, not on a miss.

One existing assertion moved

test/validate.test.ts checked that an aliased runtime produces no warnings at all. That was a proxy for "no spurious instrumentation warnings", written when instrumentation was the only thing warned about. The fixture is six lines and legitimately neither adapts nor stops early. The assertion now states its actual intent rather than the proxy.

Why child-sim is not in this PR

It has to run a page rather than read one, and doing that honestly needs a real browser or a DOM faithful enough that "the activity works" is a claim rather than a hope. A shim that half-executes generated HTML would report passes it did not earn.

That is the failure mode this session has already caught three times in its own harnesses — the shape gate on real run history, two false counterexamples in the model audit, and an inverted accommodation matcher. Shipping a fourth deliberately would be the wrong call. It needs a dependency decision (playwright as a devDep, in a repo with one runtime dep) that is yours to make.

Checks

npm test passes, 253 tests (8 new). npm run activities is the lane's instrument.

CI cannot run upstream: the account is billing-locked.

🤖 Generated with Claude Code

Stand-up step 5, the attacker half. child-sim needs to execute a page and is
not in this change; see below.

Rule 5 in CLAUDE.md forbids streaks, coins, leaderboards and countdown timers
for children. test/invariants.test.ts pinned that both prompts say so, and its
own comment describes prompt text as "the soft layer above the structural
enforcement". For engagement mechanics there was no structural layer under it.
The tutor was asked not to, and that was the whole of it. A streak counter now
refuses the save.

The second thing this reads is quieter and costs more. An activity that records
correct without recording what the child actually said makes every misconception
in that session permanently undiscoverable. src/domain/misconceptions.ts reads
response against expected; with correctness alone the evidence is a column of
zeroes. SPEC.md's central claim is that you can always recompute with a better
model, and that claim is false for any session recorded this way, because a
better model still has nothing to read.

That one does not block, and the reason is the most useful thing in this diff.

A refusal is a demand made of a model that will satisfy it the cheapest way it
can find. For a streak counter the cheapest fix is deleting the streak counter,
which is what rule 5 wants, so refusing works. For a missing response the
cheapest fix is passing response: "x". That satisfies the check and is far worse
than the gap it closes: an empty column can be recognised as empty, while a
column of plausible fabrications will be mined for misconceptions and will yield
them. A gate that can be cheaply faked does not protect the data. It corrupts it
and reports success.

So the line between refusing and flagging is not severity and not confidence. It
is whether the cheapest way to satisfy the check is also the correct fix. That
rule is now in the lane's charter and every new check has to answer it.

Four of the nine corpus samples must come back clean and they are the ones that
took the care: a streak named only in a comment, "point to the picture", a
decorative star, and a good activity that adapts on two misses and stops on
three. Precision 1, recall 1. A reviewer nobody has proved will pass good work
is a reviewer the tutor learns to route around, and once a model learns a gate
is noise, every later check inherits the distrust.

One existing assertion moved. test/validate.test.ts checked that an aliased
runtime produces no warnings at all, which was a proxy for "no spurious
instrumentation warnings" written when instrumentation was the only thing
warned about. That fixture is six lines and legitimately neither adapts nor
stops early. The assertion now says what it means.

child-sim is not here. It has to run a page rather than read one, and doing that
honestly needs a browser or a DOM faithful enough that "the activity works" is a
claim rather than a hope. A shim that half-executes generated HTML would report
passes it did not earn, which is the failure mode this session has already found
three times in its own harnesses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant