Repository navigation
Reference numbering, and a disputed citation label withheld rather than guessed - #8
Merged
Merged
Conversation
evals/DESIGN.md described _attribute_label and OCCURRENCE_MIN_RATIO, which coverage/3 deleted; grep finds neither in src/. Corrected to the ctx-id lookup it actually uses, keeping the reading-order-zipping rejection. Spec and Plan A cover the wrong-paper bug found on a real manuscript: 22 printed references parsed as 18, every claim citing [6] or above judged against a different paper.
… format The parity guard exempted claim-level keys through a hand-written set that held a dead key and omitted three real ones, and no check covered the markdown or editor templates at all. A new claim-level disclosure would have reached the viewer and silently vanished from the other three.
_title_tokens was [a-z]{5,} over lowercased text, so Späth and Müller
contributed nothing and Cristóbal became crist. That set is what
titles_match, _same_work and _title_check_text all compare. _fold existed
in refs.py and was not used here; NFKD alone leaves ß, ø, æ, đ and ł.
Docling reads a hanging-indent numeral column as a table and emits GFM. _parse_bulleted treated the pipe rows as wrapped continuations, gluing five references onto one entry — and _entry then scraped the NEXT reference's DOI onto it, which _title_check passed because the raw string held both papers. On the manuscript that surfaced this, 22 references parsed as 18 and every claim citing [6] or above was judged against a different paper.
…mbered _usable_printed_numerals required seen[0] == 1, so a list whose converter stripped the early numerals and kept the later ones was refused wholesale and then numbered by position anyway. Its docstring licensed that for lists which 'genuinely carry no numerals'; here the list carried ten, and position-vs-printed agreement was available and unread.
…ontested The note reported a difference of two totals, which understated a real run sixfold — '[1]-[19] vs 18 references' where five references were dropped and one numeral duplicated. _covers already computed the gap set and discarded it. Reconciliation.contested was declared with a comment saying that burying it was how a compensating parse error passed unmentioned, and was then never persisted and never rendered.
Three of parse_references' four return paths number by position, not by a printed numeral, and reconcile cannot tell the two apart. The failure-branch note said "carrying N distinct printed numerals" regardless, which is a positive false statement on exactly the shape Task 4 exists to handle: the converter stripped every numeral and the labels are 1..N by construction. Say "distinct labels" instead, and when neither a duplicate nor a gap is found, say so is not evidence of "none present" — only the printed numerals that survived the converter can show a duplicate or a gap at all.
_title_tokens folds, and the score is a substring test, so a page that is only lowercased matches none of the diacritic-bearing tokens. The old broken tokeniser hid it by accident — it produced "stner", and "stner" is in "küstner". Measured on a first page carrying the reference's own title verbatim, 7/7 verified became 2/9 mismatch, and _accept answers a mismatch by unlinking the downloaded PDF and telling the reader the reference carries a wrong or mistyped DOI. A correct retrieval destroyed and the manuscript blamed, on the German, Scandinavian, Polish, Turkish, Spanish and Portuguese references folding was introduced for. Every other _fold call site compares folded against folded; _fold now says that is a requirement, and its lowercase-before-translate ordering — which _TRANSLITERATE's lowercase-only keys make load-bearing — is commented and asserted.
…ot a contest Three findings in reconcile, all of them the report asserting more than the code knew. A refused entry keeps its positional label — it needs one to be a slug and a download path — so _covers' subset test is satisfied by exactly the labels in doubt and only its length test can bite, which a refusal does not change. A five-item list refused from entry 4, against a body citing [1]-[5], printed "numbering confirmed", fired no numbering disclosure and caveated no claim, in the same run that wrote "printed numbering contradicts position from entry 4 on" onto two of its entries. _refusals_unconfirm closes that on both refusal routes, the duplicate label included. contested was "a deposit exists and failed _covers", an extent test requiring len == max(body) exactly — so a deposit identical for every cited label and longer by two references nobody cites was reported as another reading naming different papers. It is now _first_divergence landing at or below the highest cited label; the disclosure says what that means. The failure note was four em-dashes in one sentence with CROSSREF_NO_DOI — which carries an em-dash and a full stop of its own — spliced into the middle of it. Sentences now, crossref last, and one duplicated label appears rather than appear. numbering_ledger's schema description names whose list it describes and documents its seven keys; nothing is required, and the round-trip test validates a ledger reconcile actually wrote.
…gate CHANGELOG gains a 0.7.0 section with the real numbers: 22 printed references parsed as 18, every claim citing [6] or above judged against a different paper, and the title check passing on the mis-attributed DOI because the glued raw string held both papers. Plus this wave's regression, the folding asymmetry that deleted correct downloads. CLAUDE.md's refs.py section gains the two rules the spec drafted: a table row in a bibliography is a reference and not a continuation, and positional numbering is never restored after a printed numeral contradicts it — nor confirmed by an extent check that counted exactly the refused labels.
… returns The case-folder prompt was the only path answer that skipped `clean_path`, so a dragged path — quoted by Finder whenever it holds a space — became a relative name starting with a literal quote. The audit landed under the cwd while every printed line named an absolute folder that did not exist. The other three prompts were immune only because they check `.exists()`; a folder about to be created has nothing to contradict it. A relative answer stays legitimate and is resolved, with the resolution printed unwrapped: that line is where a mangled path shows up before the money is spent. Whitespace arrives as `.` and is refused, on `default_case`'s own grounds that a case folder must not be the directory the user was standing in.
Plan A of two. A docling-rendered bibliography table read 22 references as 18 and judged every claim citing [6] or above against a different paper. Five changes: the claim-level disclosure parity guard, a diacritic fold used by `_title_tokens`, the table-row pre-pass, printed numerals checked against position rather than renumbered, and a numbering ledger that says which labels went missing. Two entries the converter splits mid-title still shift [13]+ on this manuscript. That state is disclosed, not silent, and Plan B is where per-label agreement withholds those verdicts.
Plan B of two, from the spec approved on 2026-09-16. Nine tasks: extract the model seam so refs may call it, a flat-text second reading of the same bibliography, per-label agreement, withholding a disputed label's verdict, the model's reading verified field by field, the wiring, the show-then-resolve escalation, the invariant test, and the docs. Three rulings worth finding here rather than rediscovering: reconcile keeps its signature and corroboration is a second axis; the model's reading can only subtract verdicts unless a human chose otherwise and it is recorded; and a truncated DOI passes both DOI_RE and the verbatim check, so it needs a shape test the spec does not have. The pairing eval task is deliberately excluded and left to a sibling plan.
`refs` is about to read the bibliography with a model, so "check.py is the only module that calls a model" becomes "one file shells out, and a test says which". The judging model was a single global overwritten by every call. With two stages calling the seam, a run whose judging made zero calls would print the reference-list model as the judge of verdicts it never saw. The site travels out of band in a context manager so `_ask`'s signature stays frozen — sixty-two test sites monkeypatch it with a two-argument lambda. The retry now reads ASK_ATTEMPTS instead of hardcoding a second attempt, and extraction gets the retry it never had. The wizard's advertised ceiling moves with it. tests/test_coverage.py is included in this commit because its _parse_json_object import moved with the function, though the task-1 brief's own file list omitted it. 🤖 Generated with Claude Code Co-Authored-By: Claude <noreply@anthropic.com>
Three review findings. `_ask_judge` is also extraction's only call site, so the name and docstring undersold it — a reader of check.py alone could not tell that extraction now retries, which is what moved the wizard's advertised ceiling. The seam test matched the literal `subprocess.run(`, so a future file shelling out via Popen, check_output or os.system would have passed a test whose name promises it cannot. It now matches every spelling a person would write, and its docstring says what a grep can and cannot prove. A zero attempt budget reached a branch whose message called itself unreachable. It raises the type the rest of the seam raises, and says which setting did it.
…rites nothing reconcile is about to arbitrate four readings instead of two. The cheapest new one is a plain pymupdf parse of the PDF the run already has open — free on a docling run, and a second opinion whose failure modes do not overlap docling's own (a mis-rendered bibliography table versus a scrambled multi-column reading order). references_span_flat is a READING, not an ingest: no out_dir, no write_outputs, no model, no network. --parse-only already carries the scar of a read that blurred that line and risked the case's own source_map.json describing a paper the run never finished auditing. Nothing calls the new function yet — Task 6 wires it into reconcile.
label_agreement joins readings of the bibliography on the printed numeral, not position — _first_divergence zips by position and would compare the wrong entries once a reading has split or merged one. A boundary_ambiguous entry and a reading carrying its own label twice both speak for nobody, and any disagreement among the voters that remain is disputed outright: no majority, matching the "a non-unique match is refused, never ranked" rule this codebase already applies to provided files. Nothing calls this yet. Reconciliation gains six fields and RefManifest eight (plus RefEntry.seen_in), all defaulted and additive, so the wire format exists before Task 4 withholds a verdict with it and Task 6 wires the two new candidate readings that make it worth calling.
…dence `label_agreement` compared voters with `_same_work`, which answers the opposite question. `_same_work` asks "is there evidence these differ?" and returns True when there is nothing to compare — correct for `_first_divergence`, which must not manufacture a divergence out of silence. A per-label vote asks "is there evidence these agree?", and the same silence must not manufacture agreement. Measured: the benefit of the doubt applied when EITHER side yielded no title tokens, so a garbled DOI-less entry paired with a legible reference came back agreed — two readings naming unrelated papers corroborating each other, and `agreed` is the one state that lets a verdict through untouched. `_comparably_same` breaks the tie the other way for this caller only. `_same_work` is untouched, because `reconcile` derives `contested` through it and needs the old direction; a test now pins both directions at once. Reverting the comparator fails exactly the two new laundering tests and leaves the DOI-less-but-same-title case passing, so the tightening is not too broad.
…ntically An identical pair exercised the overlap arithmetic trivially and would still have passed if the 0.34 floor were raised almost to 1.0 — guarding the structural case and nothing else. The two readings now differ the way two converters' readings of one entry actually differ, and share two tokens of five and seven.
A label two or more bibliography readings disagree about is dropped from the set check_claims will judge, after the retrieval filter so withholding only ever describes a source that really was fetched. The claim lands unchecked, never not_retrieved — that would falsely say the source could not be obtained when it was — and the withheld labels are named in the note. withheld_refs is a new field, not a reuse of unjudged_refs: CLAUDE.md is explicit that an entry in unjudged_refs means nobody read the source, and a withheld source was read. Landing it there would erase the retrieval gap that does exist and invent one that does not. claim_pairing is a new claim-level disclosure with one state implemented (withheld) and three more Tasks 6 and 7 will add as further branches in the same producer. The reverse-direction disclosure parity test — every key a producer emits must be declared in CLAIM_KEYS, not just the other way round — is new in this task and would have caught claim_pairing being added without it.
…ted set Three review findings. The targeted-subtraction test asserted the surviving verdict and the judged slugs, both computed from `avail` — so an edit that only populated `withheld_refs` in the nothing-survives branch would have passed the whole suite while silently dropping the claim_pairing disclosure for every co-cited claim that kept a real verdict. Simulating that edit now fails the test. `disputed` is copied into a set rather than passed through. A generator is truthy with no length, and each membership test consumes it, so the first claim would withhold and every claim after it would be judged normally — a partial withholding with nothing to notice it. The field comment borrowed "numbering gap" for something numbering does not mean here: `numbering_verified` is about reference positions, this is about two readings naming different papers under one label.
Every value the model proposes is searched for, verbatim, in one of the two extractions it was shown. A value not found is discarded; a title not found discards the whole reading. A DOI is never repaired -- DOI_RE accepts a trailing hyphen, so a line-broken 10.1038/s41591- passes both the regex and the verbatim check, and only the last-character test tells a whole DOI from half of one. Normalisation is two layers and both sides of every comparison get the same one. A numeral proposed twice is recorded and never renumbered. The reading letter a model answers with is matched exactly against A/B/AB, never as a substring of its reply -- a reply like "whichever one was clearer" contains a stray "A" that a substring test would have mistaken for reading A. A model agreeing with a parse is a second reading, not confirmation: it read the same document, so a reference the layout destroyed is one it may also have missed. ask.py, check.py, models.py and the refs manifest schema already carried the parser move and RefEntry.seen_in from earlier work on this branch, so this commit adds only the new module and its tests.
Three review findings.
A reply of `{"num": "9", "reading": "A"}` became a voter carrying no claim about
any work. The whole-candidate discard only fires when a title was PROPOSED and
failed verification, so a reply that omits the identity field entirely never
reached it. An entry with neither a verified title nor a verified DOI is now
dropped with a recorded reason.
The review diagnosed this as a wrong-verdict risk via `_same_work`'s trivial
agreement. That is not the live path: `label_agreement` compares through
`_comparably_same`, which refuses to call two uncomparable entries agreed, so
the phantom made the label `disputed`, not `agreed`. The hazard is the other
direction — a degenerate reply withholding verdicts it never said anything
against. Measured both ways before fixing.
Numbering findings move to their own list. A discarded field is "this value was
not printed"; a numeral duplicate or gap is "this reading's labels do not add
up", which the spec gives a different consequence — so one list reported a
count of discarded fields including things no field ever lost. No `covers` flag
comes with it, and the comment says why: `label_agreement` already enforces
"not covering", and a flag would be a published field with no consumer.
A `reading` value of the wrong type is now reported as a wrong shape rather than
folded into "the model said nothing".
Task 5's re-review found that Task 6 was written against the old provenance shape: it read only `fields_discarded`, so every duplicate numeral and gap `reflist.propose` reports would have been dropped before any reader saw it. `_llm_reference_reading` now returns the provenance object, which is what the task's own Interfaces block already claimed and its code contradicted — a 3-tuple of (entries, notes, model) can carry only one of the two kinds of finding. The findings get their own manifest field, their own schema description, their own console line and their own clause in the reflist disclosure, because "3 values discarded" and "it numbered one entry twice" are different facts. Step 4 also claimed five new manifest fields that Task 3 already landed. It now says so, tells the implementer to verify rather than rewrite, and adds only the one field that is genuinely new.
The flat-text reading and the model's reading are voters. Neither reaches reconcile's arguments or resolve_all, so no model output can cause a source to be resolved, downloaded or judged — only a verdict to be withheld. A run whose model reading names two different papers produces byte-identical entries to --no-llm-refs, and that equality is a test. The pymupdf reading is skipped when the run backend already is pymupdf: a reading agreeing with itself is not corroboration. --parse-only keeps it and loses the model call, because "no network" is what the flag promises. numbering_corroborated is a second axis and numbering_verified keeps its exact meaning. Corroboration is what lets the numbering warning stop being the only thing a clean reference list gets told about. claude absent, or the call failing, leaves a stated reason in the manifest and on the console — never that the model agreed.
… happened Review round 1 on the four-readings task. Branch on whether the model was actually asked (ReflistProvenance.attempted / RefManifest.reflist_attempted), never on whether the reply named itself — a call that verifies cleanly and votes can still leave reflist_model empty, and reading that emptiness as "no call was made" is the cardinal rule's own failure mode. corroborating_readings now names only readings that carried a cited label, not every reading `others` happened to hold — a reading proposing entries for labels the body never cites abstains everywhere in label_agreement and must not be credited as having corroborated anything. The pymupdf-backend skip is decided where the backend is known, and reports itself honestly instead of blaming the PDF for "only one of two extractions had text" when both did and the second was never taken on purpose. The wizard's own cost estimate and the --llm-refs help text now say the same conditional thing. The load-bearing safety test now captures what actually reaches resolve_all, closing the half of Ruling 2 the entries-equality assertion could not see on its own. Plus: an IndexError on an empty converter, a disclosure level that disagreed with its own console line, a triplicated sort key, and a duplicated PDF fixture.
Task 7's menu option 1 took a label out of labels_disputed — so check judges it again — and returned the entry list unchanged. The resolved entries sit in Resolution.resolved and were computed and thrown away, leaving the verdict to be printed against the very paper the resolution ruled against, under a report line saying the user accepted it. The original wrong-paper bug, reached through the one path a human consented to. The Interfaces block said the list is changed only by option 3, which is what made the omission read as intentional. Both now substitute, option 1 per resolved label, and a test pins it because nothing else in the plan would notice.
Review round 2. attempted: bool could not tell "never called" apart from "called and failed" -- a timeout's own note didn't start with "reading discarded", so it fell into the success branch and printed "1 value it proposed was not found in the printed text" for a call that proposed nothing. reflist_outcome is now one of not_attempted/failed/read, published as REFLIST_OUTCOMES; reflist_failure carries why, kept apart from fields_discarded so a reason can never be miscounted as a discarded value. The claude-absent test drove needs_flat=False, which returns before ask.claude_available is ever consulted -- the exact thing it claimed to cover went untested, and the "not attempted" substring it matched was the backend-skip reason. Fixed to reach the guard and assert on the reason's content; the two reasons no longer share a prefix. corroborating_readings now intersects the per-label voters label_agreement itself computes, rather than approximating "carried a cited label" -- which named a reading that skipped an entry, or whose only carrier was boundary_ambiguous, as having agreed on all of them. Walking every outcome by hand caught two more: a plain --no-llm-refs run started printing a broken disclosure once outcome was always a real string, and the console line for a reading discarded whole read like an ordinary success. Both fixed with regression tests.
The escalation. A run that found two readings naming different papers at a cited label writes every reading of every disputed label to case/out/reference_disagreement.md, with the verbatim text each was read from, and prints the path — before any question is asked. Nobody is asked to choose blind, and the test that pins it asserts the file is on disk at the moment Prompt.ask is called, because a test that checks afterwards passes just as well on an implementation that asks first. Four choices: resolve the disputed labels with a model, withhold them, adopt one reading whole, or abort. Option 1 substitutes each resolved entry into the list it returns — un-disputing a label while leaving the entry the resolution ruled against would judge the label against that paper, under a report line saying a person accepted it. A non-interactive run does not prompt: it withholds the disputed labels and keeps every other verdict, recording choice=withheld and chosen_by=default rather than a choice nobody made. chosen_by is what makes this honest. It is the only place a model reading can change what gets resolved and judged, and it is legal because a person is recorded as having asked for it. numbering_verified is untouched by every branch. A window onto the printed list now opens ABOVE the disputed label, not at it: when a reading splits one entry, the damage started in the entry before the first disputed one, so a window beginning at the label showed the symptom and hid the cause. The constant's own comment already claimed the entries either side were visible. And a run no longer prints "numbering confirmed" beside a line withholding a label's verdicts. numbering_verified means extent; a disputed label is content. Widening the flag to cover both — the first attempt — broke two of Plan A's merged tests, correctly: four report formats read that boolean. The fix belonged in the sentence, not the field.
Review round 3. reflist_entries_proposed, corroborating_readings and the
console's reply-of-[] line all asked the same empty value to carry two
meanings -- exactly the class table_warnings was already documented against
("a list, [], and null -- nobody watched").
reflist_entries_proposed is now int | None: None means never recorded, 0
means measured and zero. A round-1 manifest that recorded discarded fields
but never this count no longer reads as "proposed no entries at all", which
contradicted its own short text in the repo's own fixture.
corroborating_readings can legitimately name 0 or 1 readings even when
numbering_corroborated is True (it's a per-reading fact, corroborated is
per-label) -- round 2's intersection made that state real. Fixed the
sentence, not the field: the console and the disclosure now gate the "N
readings agree" claim on the list actually having two names, under a
different token when it doesn't, rather than redefine what
numbering_corroborated means (the thing the hard constraint this round
added says not to do).
An unrecognised outcome now fails closed instead of reading as success; the
propose guard names itself instead of blaming a model that was never asked;
the report strips the same prefix the console already does; numbering
findings stop appending to a call that never returned.
Walking the four-axis table (live/loaded, written-by-an-earlier-commit)
past the rows the review named surfaced one more: a round-1-shaped "not
obtained" note, loaded from a case folder, was being read as a discarded
value by this round's own fix for the silence bug -- reopening N1 on the
load path. Split the normalization to tell a real discarded value apart
from round 1's own failure prefixes.
Review round on the escalation (c9fa042): Critical — resolve_disputed built its RefEntry by hand and never set slug, so resolve_all downloaded every resolved label to sources_resolved/None.pdf and check.py silently refused to judge any of them, while the manifest said they were settled. Every test that reached menu option 1 stubbed resolve_disputed with refs._entry, which sets the one field the real function forgot, and so hid the bug it looked like it was testing. Fixed by building the resolved entry through refs._entry and overlaying the verified fields, and by re-running _unique_slugs over the whole substituted list. Wrote the zero-coverage tests/test_reflist.py suite this needed (17 tests against the real resolve_disputed, mocking only ask._ask), plus one end-to-end test through the real function via _refs_pipeline. Also from the same review: seen_in was translated through the wrong dict (names now passed explicitly, matching propose's own label_a/label_b); option 3 carried a discarded reading's extent claim forward (verified/ source/ledger now cleared, not widened); option 1's console asserted abstention over a malformed reply or an unprinted title alike (now named per the real reason resolve_disputed already records); option 3 could adopt a reading that itself duplicated the disputed label, laundering the coin-flip label_agreement refuses; the menu's default offer of "3" could raise IndexError with nothing to offer; a pre-escalation line promised labels "will be reported unchecked" moments before options 1/3 checked them; and the escalation had no run-level disclosure, so an 11-of-14-resolved run said so only on claims that cite one of the 11. Eleven further minors from the same review, addressed or explicitly declined with reasons in task-7-report.md.
…laces a convention ADR 0002 pre-specified this remedy — a third candidate reading behind an opt-in flag, not a replacement of the parser — and explicitly declined to authorise it. ADR 0003 is the authorisation, with the trigger now evidenced: 22 printed references read as 18, six substantive verdicts about papers the manuscript never cited, every one of them carrying title_check verified. It records the four rejected alternatives so they are not re-proposed. CLAUDE.md's "check.py is the only module that calls a model" becomes "one file in src/ shells out, and a test says which" — the rule was protecting the seam, not the module. The refs.py section gains per-label agreement, why the join is the printed numeral and not the position, and why single prints. README gains a does bullet and a does-not bullet stating the ceiling: corroborated is not verified, every reading read the same document, and nothing a model returns or a user answers sets numbering_verified. Line 216 is untouched and nothing is retracted. No accuracy figure anywhere — none has been measured. evals/DESIGN.md gets a pointer to the sibling eval plan, not the plan.
Most of this design's safety lives in one flag. Parametrised over eight model reply shapes crossed with two extents (a fixture that legitimately verifies and one that legitimately does not) — including the well-formed reply that agrees with the parse, which is the tempting case, not the malformed one — plus all four menu choices crossed with both starting values of the flag, plus a resolver call that raises. The brief prescribing this test had two defects of its own: its reply-shape fixture asserted verified=False against a fixture that legitimately verifies True, and its monkeypatch target (reflist_mod._ask) is never read by the real code, which imports the whole ask module and calls ask._ask qualified — as written it would have let every reply-shape test fall through to a live subprocess call. Both fixed; see task-8-report.md for detail, including a mutation from the brief's own verification step that turned out to be a no-op against the real code. A companion asserts numbering_corroborated CAN be True in the same run: the two axes are independent, and a second axis that moves only with the first would leave the 22-vs-19 false alarm firing. Plus a grep guard over src/, whose docstring says what a grep cannot prove and which has a self-test so it has been observed failing.
…nnot corroborate `_comparably_same` returns `True | False | None`, the same three-valued vocabulary `models.titles_match` publishes — the two are siblings answering opposite questions. A two-valued `False` for "nothing to compare" made a Crossref deposit of bare DOIs dispute every label it voted on while both parses agreed perfectly, and a run then printed an audit with no verdicts in it. `_label_state` is the one place the rule lives now: a cannot-tell voter is excluded from the comparison rather than counted as dissent, which can leave one reading standing and print `single`. A label NO voter says anything comparable about stays `disputed` — it is carried, and nothing establishes what it names. `DERIVED_READINGS` excludes the model's reading from crediting an agreement. `reflist.propose` may only copy from the two extractions that vote here in their own right, so its agreement is one text read twice. It still votes and can still dispute. `stamp_seen_in` matches through the same comparator, so `seen_in` and `labels_disputed` can no longer contradict each other on one manifest; it credits a chosen entry to the reading it came from by identity and merges rather than overwrites, so `[]` keeps one meaning.
… reach disputed `c.unjudged_refs` is decided above every branch of the claim loop, so the withheld path can no longer `continue` past it: a claim citing a withheld [3] and a paywalled [4] reported [4] in none of the three accounts a reader has. It stays empty in exactly one state — a claim where nothing was judged, withheld or provided — where the `not_retrieved` verdict is already the whole report. `_claim_pairing` splits a disputed label on the manifest entry's own status, not on the emptiness of `withheld_refs`: a label nobody retrieved gets its own token and its own sentence instead of "the source fetched under that label", and a `results.json` written before `withheld_refs` existed still reads as the fetched case. `label_agreement` returns `disputed` for two causes and no field tells them apart, so no surface asserts the contradicting one: the note, both disclosures, the console, the disagreement file and both schema descriptions name the disjunction.
…s who made it `resolve_disputed` is passed `disputing` — what each reading says at each disputed label, including the deposit and the model's own reading, which have no printed span and were therefore invisible to a call adjudicating between them. It is context: `_found` still verifies every returned value against the printed extraction alone. The prompt renders one block per extraction it actually has, so a pymupdf-backend run no longer shows an empty `READING B`. Five `resolution_*` manifest fields carry that call's own provenance. A run whose backend skips the reading and then escalates recorded `reflist_outcome: not_attempted` and nothing else, so the report said no model read the reference list on a run where one was asked, verified field by field and allowed to change which papers are judged. The `reflist` disclosure now names it in every branch, and fires on its own when it was the only call. The note, the run-level resolution and the per-claim disclosure name the extractions the model was shown instead of asserting "both texts". Also: `numbering_corroborated` is `bool | None`, so "never computed" is a value and not a comment; `stamp_seen_in` runs after the escalation, so option 3 no longer publishes an empty provenance for a run that computed one; option 3's note says the reading was adopted whole; the resolution reply's entry count is taken before de-duplication; a mid-branch `reflist_attempted` keeps its disclosure without claiming an outcome nobody recorded; and both calls refuse an oversized bibliography before spending anything rather than clipping it into a reading quietly short.
Found by enumerating the states rather than the branches: the not-obtained, failed and unrecognised heads end "the reference numbering rests on the readings above it and nothing else", and the resolution sentence appended after them describes a call that substitutes an entry. The claim is dropped where a resolution happened and kept where none did, since it is load-bearing there. `_also_resolved` is one producer for the sentence both dispute states owe, so a claim citing an unretrieved disputed label and a resolved one no longer loses the resolved caveat to the state added later.
`resolution_readings == []` is never recorded, and on this branch that identifies a manifest written before these fields existed — whose call was also not shown what each reading said at the disputed label. Asserting either half of this build's fuller account over that run is the report describing behaviour the run did not have. `_resolution_sentence` is the one producer for both surfaces.
…dispute `_title_tokens` strips DOIs the way it already strips URLs, for the reason its own docstring gives: `10.1148/radiol.…` yielded `radiol` and `10.1001/jamanetworkopen.…` yielded `jamanetworkopen`, and those identifier substrings were compared against real title words. Only a DOI with no five-letter run reached the abstention added in round 1, so the common publisher shape still disputed every label a bare-DOI deposit voted on. Two consumers move, both toward honesty and both measured in the report: `_same_work` stops reporting a false divergence at a DOI-only deposit, and `_title_check_text` stops carrying a token no first page prints in its denominator — which on a short reference whose slug DID match can now reach `mismatch`, correctly, where a spurious match used to hold it at unverifiable. `labels_uncomparable` is a published subset of `labels_disputed`: the labels withheld because nothing could be compared rather than because the readings named different papers. `labels_disputed` stays the only list that drives withholding, the subset is read off the same `_label_state` call so the two cannot contradict each other, and it is re-intersected once after the escalation so a resolved label cannot survive in it. `None` is never computed and `[]` is a measurement. Every surface — the run-level disclosure, both claim-level dispute states, the withholding note and the console — names the cause through one producer, `disclosures.dispute_causes`, and falls back to the disjunction only where nothing measured the split. `_label_key` moves down to `refs`, the layer below both its users.
…odel call
The `--llm-refs` reading is a voter and reaches neither `reconcile`'s arguments
nor `resolve_all`. Three files generalised that to "no model output can cause a
source to be resolved, downloaded or judged", which the escalation's option 1
falsifies: `_escalate_disputed` substitutes entries `reflist.resolve_disputed`
built from a model's reply, `resolve_all` downloads them and `check.py` judges
them. The substitution is mandatory — returning the list untouched un-disputes
the label while leaving the entry the resolution ruled against as the paper
judged. The exception is now stated where a reader meets it, with its gate: an
interactive run, the disagreement on disk before anything is asked,
`numbering_chosen_by: "user"`, only the resolved labels substituted, and no
answer setting `numbering_verified`.
Constraint 1 cited a module that tests constraint 5; it now cites
`test_a_model_naming_different_papers_changes_no_resolved_entry`. The
escalation's stated precondition ("nothing accounted for the labels AND labels
in dispute") does not exist — `_escalate_disputed` returns early on an empty
`labels_disputed` and nothing else. A disagreement withholds the source
always, the claim only when no other source answered.
Also: `_comparably_same`, not `_same_work`, and why; the third value and the
three dispute causes; `labels_uncomparable` as a published subset;
`numbering_corroborated` as `bool | None`; the DOI strip in `_title_tokens`;
`reflist.py` rather than `refs.py` as the seam's second caller;
`reconcile`'s two candidate readings rather than its arity; the whole reading
discarded by one unverifiable title; the corroboration bullet's other branch;
`CI` and a piped stdout as non-interactive; no integer for the seam's patch
sites, which moved again today. Two wire-format rules earned here go into the
cardinal-rule section.
… anything The seam comment claimed sixty-two monkeypatch sites. That was raw grep hits on `_ask`; the real figure measured 46 at review time and 48 a day later, because it grows with every test that touches the seam. Three values for one claim is the tell — the number is a maintenance cost pretending to be evidence, so the comment now states the constraint without one. ADR 0002 cross-referenced README §"Two independent readings" by title. 0.7.0 took the count from two readings to four and the section was renamed, so the reference did not break loudly, it simply stopped naming anything. It now points at the section that exists and forward to 0003, which is the authorisation 0002 explicitly withheld.
… wrote The reply-shape half of this module built `candidates` and called `reconcile` by hand, so "the model's reading never becomes one of `reconcile`'s arguments" was enforced by the test's own source text. Handing the `llm` reading to `reconcile` in `_refs_pipeline` — the design's central prohibition — left all 29 cells green. So did disabling `reflist._found`, and so did widening the manifest line to `rec.verified or rec.corroborated`. Every cell now drives `cli._refs_pipeline` with `ask._ask` patched per reply shape and reads `refs_manifest.json`. `reflist.propose`, `resolve_disputed`, `reconcile`, `label_agreement` and the manifest construction all run for real; the disputed label is produced by the pipeline rather than assigned by the test, and the raising-resolver cell raises at the seam instead of stubbing the resolver. Each shape carries what it is allowed to have moved, because an invariant asserted alone cannot tell "the reply changed nothing" from "the reply never ran". The grep splits on the first `=` that is not a comparison rather than discarding any line containing one, and `corroborated` joins `_FORBIDDEN`; the self-test plants all three spellings and requires all three caught.
A reply naming only [99] makes every numeral below it absent, which printed 98 consecutive integers — 410 characters — into reflist_numbering_findings, a field the report shows a reader. The first six locate the gap; the rest only prove the reader stopped reading. Nothing is hidden and the finding is unchanged: the total leads the sentence.
… side
`limits` (--max-claims, --max-sources) and this branch's per-label agreement
both decide whether a cited source is judged, so the merge had to keep three
reasons a source yields no verdict apart, not two:
not obtainable retrieval failed fix access, or accept the gap
skipped on request you asked for a slice raise the limit and re-run
withheld the readings disagree check that citation by hand
Verified end to end: paywalled reports `not available (paywalled)`, a capped
source `not available (skipped on request)`, and a disputed label `unchecked`
with the label named. Each says which happened; none is summarised into
another.
They compose by construction: `skipped` is not in ("retrieved", "provided"),
so a skipped entry never reaches the partition that withholding splits, and a
source can never be both.
Resolutions worth knowing. cli.py took both features' flags on `refs` and
`run`, and the stage order `limits` introduced survives —
ingest → extract → refs → scout → check → highlight → report — with extraction
before references so a claims limit retrieves only what the selected claims
cite, and the escalation still inside `_refs_pipeline`. check.py kept this
branch's withheld branch with `limits`' skipped-aware reasons inside it.
CLAUDE.md kept this branch's `ask.py` seam rule and dropped the old "check.py
is the only module that calls a model", which is now false, while taking
`limits`' paragraph on how limits reach `resolve_all`. README took the shorter
rewrite from #6/#7 as the base and swapped in the four-readings bullet, since
theirs still described two. One 0.7.0 section in the changelog carries both.
Two tests needed the merged shape rather than either side's: the wizard's cost
estimate, because the reference-list call survives every limit and `apply_limits`
already accounted for it; and the flag-forwarding test, which had to stub the
new extract stage.
1170 passed, ruff clean.
All ten CI jobs failed on two tests that pass locally. `_check_pipeline` calls `cli._require_claude()`, which does `from .check import claude_available` — and `check.claude_available` is a re-export, a separate module attribute from `ask.claude_available` even though both name one function. The tests patched `ask_mod`, so the binding this path actually reads was never touched: they passed on a machine with the CLI installed and failed on ten that have none. The merge caused it. Before it, `_check_pipeline` imported from `.ask` and the patch was right; `limits` refactored the check into `_require_claude`, which imports from `.check`, and the comment above the patch went on describing the old import. Patching `_require_claude` itself cannot drift with a future refactor of which module it reads. Verified the way CI does, which is the only way this shows: with `claude` off PATH, 1145 passed, 25 skipped.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two plans, worked as one branch: reference numbering (Plan A) and a disputed
label withheld rather than guessed (Plan B).
Neither plan has ever been on
origin/main— Plan A was merged locally andnever pushed, so this branch carries both.
Why
A citation label is the join key between a claim and the source it is judged
against, so a list off by one produces a confident audit of the wrong papers.
On the manuscript that prompted this, 22 printed references parsed as 18 and
every claim citing
[6]or above was judged against a different paper — withtitle_check: verifiedprinted beside it, because the glued reference stringcontained both papers' words.
What it does
Plan A — docling renders a hanging-indent bibliography as a GFM table, so
entries arrive as
| 6. | Author A (2019) … |and the marker regex cannot seethem.
_unwrap_table_rowsrestores the bullet form; printed numerals arechecked against position rather than renumbered, and entries that contradict
their position refuse to resolve instead of risking a wrong paper.
Plan B — up to four independent readings of the bibliography: the Crossref
deposit, the run backend's parse, a flat-text parse of the same PDF, and a
model's reading whose every field must be found verbatim in one of the two
texts it was shown. Agreement is computed per printed label, never by
position. Where the readings name different papers — or where nothing in them
can be compared — the verdict is withheld and reported
unchecked, neverprinted against a paper that may be the wrong one. An interactive escalation
writes the full disagreement to
case/out/reference_disagreement.mdbeforeasking anything.
The model's reading is a voter. It cannot cause a source to be resolved,
downloaded or judged — with one exception, which is the point of the feature: a
person shown the disagreement may ask a model to resolve it, and that is
recorded as
numbering_chosen_by: "user", substitutes only the labelsresolved, and never marks the numbering confirmed.
Review history
Nine tasks, five fix rounds, three final reviews (whole-branch, and one each
for the last two tasks). Two Criticals and fourteen Majors found and fixed.
Almost every one was a two-valued thing asked to carry three meanings —
same/different/cannot-tell, measured-zero/never-measured,
not-attempted/failed/answered — in a codebase whose product is an honest
account of what it does not know.
1135 passed,ruff check src tests scripts evalsclean.limitsfeature, and the overlap is the pointorigin/maingainedlimits(PRs #6, #7) while this was in flight. Merginggives 17 conflict hunks across 6 files, and they are not textual drift —
the two features overlap by design:
limitscheck_claimsavailability filterClaimResult/RefManifestdisclosures.pyBoth partition the same
availlist incheck.py, and both add reader-facingdisclosures about why a source was not judged.
The question a reviewer should hold: "I only audited the first 20 claims"
and "I withheld this verdict because two readings disagree" must never be
confusable. A reader shown one when the other is true is exactly the defect
class this branch spent nine tasks removing — and the test suite cannot catch
it, because both render as honest-looking strings.
I did not resolve the conflicts rather than risk that merge being made blind.
Known, and deliberately not done
--llm-refsdefaults on, so a fresh demo runmakes one extra model call and emits a
reflistdisclosure thatexamples/demo/output/does not show. Recorded as known staleness inCHANGELOG.md; the re-run costs money and rewrites committed showcaseartifacts the README's screenshots describe.