cross-model-review §3 + op-rigor §4: a scope qualifier is not a repair; a verdict line needs sufficient observations - #235
Conversation
|
Probe of both rules (15 runs, pre-registered, 2026-09-05) Ran a bare-vs-ruled adjudication drill on five fixtures, three arms (no rules / §3 only / §4 only), two model families (A, B). A first attempt was retracted: the rule text quoted a fixture's exact qualifier and both rules were prepended in one arm, so no shift attributed to either rule. This is the corrected run — cues stripped, rules isolated, key and decision rules committed before any run.
§3 — partial support, with a confound. Family A accepted the unbounded qualifier as a repair in every bare run, using the exact rationalization §3 names ("limits the claim to the specific failure mode, correcting the over-claim"). §3 alone flips it 2/2. But §4 alone flips it 2/2 as well, in §4's own vocabulary — so what is measured is that some claim-vs-evidence rule is load-bearing on this item, not that §3's content specifically is. §3 keeps its Worth flagging: §4 alone made family B worse on the §3 item (3/3 → 1/3). Both misses read §4's "the line is weakened to what was observed" as endorsing a scope qualifier — the move §3 exists to forbid. n=3, so a signal rather than a finding, but the two rules' remedy language may read as pulling in opposite directions, and §3 lands after §4 in a reader's path through the pack. §4 — redundant for adjudication at both tiers (10/10 bare on the two negative items; the positive control passed 15/15). Expected, and not an argument to drop it: this measures judging a harness someone hands you, not catching the defect while writing your own — and the incident §4 came from was the latter. §4's detection value is untested by this design. Controls clean (15/15 both), no over-strictness observed. Both |
…r; a verdict line needs sufficient observations Two rules distilled from PR F-e-u-e-r#233's own review rounds. Three review rounds (codex luna + grok; then codex sol + Fable 5.1 fresh-context ×2), capped; every round's fixes introduced a defect the next round caught. Trail in reviews/2026-09-02-narrowing-is-not-repair/.
eb8fd22 to
afc1b82
Compare
|
Superseded by owner rebuild #241 (merged as 74e97a8); not merged here.
Closing as replaced/landed. |
Two rules distilled from #233's own review rounds (this fork's out-of-tree-cache PR):
qualified claim is no wider than what the evidence establishes, AND every
reproduced counterexample the finding cites falls outside it. Round 1 of op-rigor §2 path (a): locate the competing artifact — in-tree removal can no-op #233
recorded a qualifier ("for this failure mode") as
fixed; round 2 showed twocounterexamples still inside it.
sufficient for every entity AND relation it names, never by proxies merely
consistent with it (three obligations: sufficient observation per entity+relation;
forced fixture conditions named in the line and any claim the probe supports; an
asserted diagnostic result needs a captured invocation/observation, not just a
claim). op-rigor §2 path (a): locate the competing artifact — in-tree removal can no-op #233's probe printed "ran from an out-of-tree cache" from two consistent-
but-unmeasured proxies.
Both ship
unprobedper the covenant.Review — three rounds, capped
gpt-5.6-luna(inlined, read-only); grok-4.6 (isolated HOME, staged file, read-only tools)gpt-5.6-sol; a fresh-context same-family subagent (packet-only, no repo access)Every finding was reproduced (derivation from the inlined packet, or execution against
the cited #233 trail) before being acted on. All fixes are authored, not pasted — see
ADJUDICATION-r{1,2,3}.mdfor each finding's disposition and rationale, including tworejected-with-reasonitems.Family diversity: round 1 used two model families, neither the author's. Rounds 2–3
used the author's own family (fresh-context) plus one other (OpenAI codex sol) — this is
disclosed as a gate on the author's own family for those rounds, not a fully independent
cross-family pair for the whole trail.
What the rounds show, on the rule's own subject: every round's fixes introduced at
least one new defect the next round caught — the heading contradicted its own body (r1),
the replacement "every noun measured" missed the relation the incident actually got
wrong (r2), the replacement "measured" contradicted the body's own non-empirical branch
(r3). The post-round-3 wording is unreviewed — nine single-clause edits, each authored
against the full paragraph and recorded in
ADJUDICATION-r3.md, but not sent to anotherround. Recording this as the shipped gap rather than opening an uncapped round 4.
Full trail:
reviews/2026-09-02-narrowing-is-not-repair/(packets, pre-registeredexpectations written before each round's verdicts were read, all six verdicts,
three adjudication tables, README).