Skip to content

cross-model-review §3 + op-rigor §4: a scope qualifier is not a repair; a verdict line needs sufficient observations - #235

Closed
firaen22 wants to merge 1 commit into
F-e-u-e-r:mainfrom
firaen22:lesson/narrowing-is-not-correcting
Closed

firaen22 wants to merge 1 commit into
F-e-u-e-r:mainfrom
firaen22:lesson/narrowing-is-not-correcting

Conversation

@firaen22

@firaen22 firaen22 commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Two rules distilled from #233's own review rounds (this fork's out-of-tree-cache PR):

  1. cross-model-review §3 — a scope qualifier repairs an over-claim only when the
    qualified claim is no wider than what the evidence establishes, AND every
    reproduced counterexample the finding cites falls outside it. Round 1 of op-rigor §2 path (a): locate the competing artifact — in-tree removal can no-op #233
    recorded a qualifier ("for this failure mode") as fixed; round 2 showed two
    counterexamples still inside it.
  2. operational-rigor §4 — a probe's verdict line is established by observations
    sufficient for every entity AND relation it names, never by proxies merely
    consistent with it (three obligations: sufficient observation per entity+relation;
    forced fixture conditions named in the line and any claim the probe supports; an
    asserted diagnostic result needs a captured invocation/observation, not just a
    claim). op-rigor §2 path (a): locate the competing artifact — in-tree removal can no-op #233's probe printed "ran from an out-of-tree cache" from two consistent-
    but-unmeasured proxies.

Both ship unprobed per the covenant.

Review — three rounds, capped

Round Reviewers Verdicts
1 codex gpt-5.6-luna (inlined, read-only); grok-4.6 (isolated HOME, staged file, read-only tools) luna: body FIX / last line PROCEED (self-contradicting, treated as FIX per cross-model-review §5); grok: FIX ×6
2 (gate) codex gpt-5.6-sol; a fresh-context same-family subagent (packet-only, no repo access) FIX ×7 / FIX ×7
3 (gate, cap) same pair FIX ×3 + 1 Low / FIX ×1 + 6 Low

Every finding was reproduced (derivation from the inlined packet, or execution against
the cited #233 trail) before being acted on. All fixes are authored, not pasted — see
ADJUDICATION-r{1,2,3}.md for each finding's disposition and rationale, including two
rejected-with-reason items.

Family diversity: round 1 used two model families, neither the author's. Rounds 2–3
used the author's own family (fresh-context) plus one other (OpenAI codex sol) — this is
disclosed as a gate on the author's own family for those rounds, not a fully independent
cross-family pair for the whole trail.

What the rounds show, on the rule's own subject: every round's fixes introduced at
least one new defect the next round caught — the heading contradicted its own body (r1),
the replacement "every noun measured" missed the relation the incident actually got
wrong (r2), the replacement "measured" contradicted the body's own non-empirical branch
(r3). The post-round-3 wording is unreviewed — nine single-clause edits, each authored
against the full paragraph and recorded in ADJUDICATION-r3.md, but not sent to another
round. Recording this as the shipped gap rather than opening an uncapped round 4.

Full trail: reviews/2026-09-02-narrowing-is-not-repair/ (packets, pre-registered
expectations written before each round's verdicts were read, all six verdicts,
three adjudication tables, README).

@firaen22

firaen22 commented Sep 5, 2026

Copy link
Copy Markdown
Contributor Author

Probe of both rules (15 runs, pre-registered, 2026-09-05)

Ran a bare-vs-ruled adjudication drill on five fixtures, three arms (no rules / §3 only / §4 only), two model families (A, B). A first attempt was retracted: the rule text quoted a fixture's exact qualifier and both rules were prepended in one arm, so no shift attributed to either rule. This is the corrected run — cues stripped, rules isolated, key and decision rules committed before any run.

item key bare A §3 A §4 A bare B §3 B §4 B
unbounded qualifier recorded fixed INVALID 0/2 2/2 2/2 3/3 3/3 1/3
bounded qualifier (control) VALID 2/2 2/2 2/2 3/3 3/3 3/3
verdict line from proxies INVALID 2/2 2/2 2/2 3/3 3/3 3/3
verdict line with sufficient observations (control) VALID 2/2 2/2 2/2 3/3 3/3 3/3
verified on uncaptured diagnostics INVALID 2/2 2/2 2/2 3/3 3/3 3/3

§3 — partial support, with a confound. Family A accepted the unbounded qualifier as a repair in every bare run, using the exact rationalization §3 names ("limits the claim to the specific failure mode, correcting the over-claim"). §3 alone flips it 2/2. But §4 alone flips it 2/2 as well, in §4's own vocabulary — so what is measured is that some claim-vs-evidence rule is load-bearing on this item, not that §3's content specifically is. §3 keeps its unprobed marker: one tier, n=2, content-specificity unidentified. A placebo-rule arm (an unrelated rule of similar length) would separate "any strictness prompt" from "this rule".

Worth flagging: §4 alone made family B worse on the §3 item (3/3 → 1/3). Both misses read §4's "the line is weakened to what was observed" as endorsing a scope qualifier — the move §3 exists to forbid. n=3, so a signal rather than a finding, but the two rules' remedy language may read as pulling in opposite directions, and §3 lands after §4 in a reader's path through the pack.

§4 — redundant for adjudication at both tiers (10/10 bare on the two negative items; the positive control passed 15/15). Expected, and not an argument to drop it: this measures judging a harness someone hands you, not catching the defect while writing your own — and the incident §4 came from was the latter. §4's detection value is untested by this design.

Controls clean (15/15 both), no over-strictness observed. Both unprobed markers stay as shipped.

…r; a verdict line needs sufficient observations

Two rules distilled from PR F-e-u-e-r#233's own review rounds. Three review rounds
(codex luna + grok; then codex sol + Fable 5.1 fresh-context ×2), capped;
every round's fixes introduced a defect the next round caught. Trail in
reviews/2026-09-02-narrowing-is-not-repair/.
@F-e-u-e-r

Copy link
Copy Markdown
Owner

Superseded by owner rebuild #241 (merged as 74e97a8); not merged here.

  • Condensed R1 landed via cross-model-review §3: narrowing is not repair by itself #241 — cross-model-review §3 "narrowing is not repair by itself": a scope qualifier closes an over-claim finding only when the revised claim is supported by the evidence and every reproduced counterexample the finding cites falls outside the revised scope.
  • R2 (operational-rigor §4 sufficient-observation verdict-line) was rejected as duplicate — its founding case is already covered by op-rigor §4 (a check's name is not its coverage + exit-0-is-evidence + data-path inability).
  • This PR's historical derivation evidence (reviews/2026-09-02-narrowing-is-not-repair/) remains available here and was not recopied into cross-model-review §3: narrowing is not repair by itself #241.
  • cross-model-review §3: narrowing is not repair by itself #241 was independently gated on the exact rebuilt revision (Luna Max + Sol Max, PROCEED/PROCEED) and post-merge verified (tree(C) == reviewed tree; default-branch CI 3/3).

Closing as replaced/landed.

@F-e-u-e-r F-e-u-e-r closed this Sep 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants