Skip to content

Reference numbering, and a disputed citation label withheld rather than guessed - #8

Merged
defraction0 merged 46 commits into
mainfrom
feat/llm-reference-list
Sep 19, 2026
Merged

defraction0 merged 46 commits into
mainfrom
feat/llm-reference-list

Conversation

@defraction0

Copy link
Copy Markdown
Owner

Two plans, worked as one branch: reference numbering (Plan A) and a disputed
label withheld rather than guessed (Plan B).

Neither plan has ever been on origin/main — Plan A was merged locally and
never pushed, so this branch carries both.

Why

A citation label is the join key between a claim and the source it is judged
against, so a list off by one produces a confident audit of the wrong papers.
On the manuscript that prompted this, 22 printed references parsed as 18 and
every claim citing [6] or above was judged against a different paper — with
title_check: verified printed beside it, because the glued reference string
contained both papers' words.

What it does

Plan A — docling renders a hanging-indent bibliography as a GFM table, so
entries arrive as | 6. | Author A (2019) … | and the marker regex cannot see
them. _unwrap_table_rows restores the bullet form; printed numerals are
checked against position rather than renumbered, and entries that contradict
their position refuse to resolve instead of risking a wrong paper.

Plan B — up to four independent readings of the bibliography: the Crossref
deposit, the run backend's parse, a flat-text parse of the same PDF, and a
model's reading whose every field must be found verbatim in one of the two
texts it was shown. Agreement is computed per printed label, never by
position. Where the readings name different papers — or where nothing in them
can be compared — the verdict is withheld and reported unchecked, never
printed against a paper that may be the wrong one. An interactive escalation
writes the full disagreement to case/out/reference_disagreement.md before
asking anything.

The model's reading is a voter. It cannot cause a source to be resolved,
downloaded or judged — with one exception, which is the point of the feature: a
person shown the disagreement may ask a model to resolve it, and that is
recorded as numbering_chosen_by: "user", substitutes only the labels
resolved, and never marks the numbering confirmed.

Review history

Nine tasks, five fix rounds, three final reviews (whole-branch, and one each
for the last two tasks). Two Criticals and fourteen Majors found and fixed.
Almost every one was a two-valued thing asked to carry three meanings —
same/different/cannot-tell, measured-zero/never-measured,
not-attempted/failed/answered — in a codebase whose product is an honest
account of what it does not know.

1135 passed, ruff check src tests scripts evals clean.

⚠️ This conflicts with the limits feature, and the overlap is the point

origin/main gained limits (PRs #6, #7) while this was in flight. Merging
gives 17 conflict hunks across 6 files, and they are not textual drift —
the two features overlap by design:

this branch limits
check_claims availability filter withholds disputed sources caps sources per claim
ClaimResult / RefManifest ~12 new fields 117 lines of new fields
disclosures.py 4 new keys 183 lines of new disclosures
the four report templates new allow-list branches new allow-list branches

Both partition the same avail list in check.py, and both add reader-facing
disclosures about why a source was not judged.

The question a reviewer should hold: "I only audited the first 20 claims"
and "I withheld this verdict because two readings disagree" must never be
confusable.
A reader shown one when the other is true is exactly the defect
class this branch spent nine tasks removing — and the test suite cannot catch
it, because both render as honest-looking strings.

I did not resolve the conflicts rather than risk that merge being made blind.

Known, and deliberately not done

  • Gate 5 is undischarged. --llm-refs defaults on, so a fresh demo run
    makes one extra model call and emits a reflist disclosure that
    examples/demo/output/ does not show. Recorded as known staleness in
    CHANGELOG.md; the re-run costs money and rewrites committed showcase
    artifacts the README's screenshots describe.
  • No accuracy figure is claimed anywhere, because none has been measured.
  • The pairing eval task is out of scope and left to a sibling plan.

defraction0 and others added 30 commits September 16, 2026 17:10
evals/DESIGN.md described _attribute_label and OCCURRENCE_MIN_RATIO, which
coverage/3 deleted; grep finds neither in src/. Corrected to the ctx-id
lookup it actually uses, keeping the reading-order-zipping rejection.

Spec and Plan A cover the wrong-paper bug found on a real manuscript: 22
printed references parsed as 18, every claim citing [6] or above judged
against a different paper.
… format

The parity guard exempted claim-level keys through a hand-written set that
held a dead key and omitted three real ones, and no check covered the
markdown or editor templates at all. A new claim-level disclosure would
have reached the viewer and silently vanished from the other three.
_title_tokens was [a-z]{5,} over lowercased text, so Späth and Müller
contributed nothing and Cristóbal became crist. That set is what
titles_match, _same_work and _title_check_text all compare. _fold existed
in refs.py and was not used here; NFKD alone leaves ß, ø, æ, đ and ł.
Docling reads a hanging-indent numeral column as a table and emits GFM.
_parse_bulleted treated the pipe rows as wrapped continuations, gluing five
references onto one entry — and _entry then scraped the NEXT reference's DOI
onto it, which _title_check passed because the raw string held both papers.
On the manuscript that surfaced this, 22 references parsed as 18 and every
claim citing [6] or above was judged against a different paper.
…mbered

_usable_printed_numerals required seen[0] == 1, so a list whose converter
stripped the early numerals and kept the later ones was refused wholesale and
then numbered by position anyway. Its docstring licensed that for lists which
'genuinely carry no numerals'; here the list carried ten, and
position-vs-printed agreement was available and unread.
…ontested

The note reported a difference of two totals, which understated a real run
sixfold — '[1]-[19] vs 18 references' where five references were dropped and
one numeral duplicated. _covers already computed the gap set and discarded it.

Reconciliation.contested was declared with a comment saying that burying it
was how a compensating parse error passed unmentioned, and was then never
persisted and never rendered.
Three of parse_references' four return paths number by position, not by a
printed numeral, and reconcile cannot tell the two apart. The failure-branch
note said "carrying N distinct printed numerals" regardless, which is a
positive false statement on exactly the shape Task 4 exists to handle: the
converter stripped every numeral and the labels are 1..N by construction.

Say "distinct labels" instead, and when neither a duplicate nor a gap is
found, say so is not evidence of "none present" — only the printed numerals
that survived the converter can show a duplicate or a gap at all.
_title_tokens folds, and the score is a substring test, so a page that is
only lowercased matches none of the diacritic-bearing tokens. The old broken
tokeniser hid it by accident — it produced "stner", and "stner" is in
"küstner". Measured on a first page carrying the reference's own title
verbatim, 7/7 verified became 2/9 mismatch, and _accept answers a mismatch by
unlinking the downloaded PDF and telling the reader the reference carries a
wrong or mistyped DOI. A correct retrieval destroyed and the manuscript
blamed, on the German, Scandinavian, Polish, Turkish, Spanish and Portuguese
references folding was introduced for.

Every other _fold call site compares folded against folded; _fold now says
that is a requirement, and its lowercase-before-translate ordering — which
_TRANSLITERATE's lowercase-only keys make load-bearing — is commented and
asserted.
…ot a contest

Three findings in reconcile, all of them the report asserting more than the
code knew.

A refused entry keeps its positional label — it needs one to be a slug and a
download path — so _covers' subset test is satisfied by exactly the labels in
doubt and only its length test can bite, which a refusal does not change. A
five-item list refused from entry 4, against a body citing [1]-[5], printed
"numbering confirmed", fired no numbering disclosure and caveated no claim,
in the same run that wrote "printed numbering contradicts position from entry
4 on" onto two of its entries. _refusals_unconfirm closes that on both
refusal routes, the duplicate label included.

contested was "a deposit exists and failed _covers", an extent test requiring
len == max(body) exactly — so a deposit identical for every cited label and
longer by two references nobody cites was reported as another reading naming
different papers. It is now _first_divergence landing at or below the highest
cited label; the disclosure says what that means.

The failure note was four em-dashes in one sentence with CROSSREF_NO_DOI —
which carries an em-dash and a full stop of its own — spliced into the middle
of it. Sentences now, crossref last, and one duplicated label appears rather
than appear.

numbering_ledger's schema description names whose list it describes and
documents its seven keys; nothing is required, and the round-trip test
validates a ledger reconcile actually wrote.
…gate

CHANGELOG gains a 0.7.0 section with the real numbers: 22 printed references
parsed as 18, every claim citing [6] or above judged against a different
paper, and the title check passing on the mis-attributed DOI because the
glued raw string held both papers. Plus this wave's regression, the folding
asymmetry that deleted correct downloads.

CLAUDE.md's refs.py section gains the two rules the spec drafted: a table row
in a bibliography is a reference and not a continuation, and positional
numbering is never restored after a printed numeral contradicts it — nor
confirmed by an extent check that counted exactly the refused labels.
… returns

The case-folder prompt was the only path answer that skipped `clean_path`, so a
dragged path — quoted by Finder whenever it holds a space — became a relative
name starting with a literal quote. The audit landed under the cwd while every
printed line named an absolute folder that did not exist. The other three
prompts were immune only because they check `.exists()`; a folder about to be
created has nothing to contradict it.

A relative answer stays legitimate and is resolved, with the resolution printed
unwrapped: that line is where a mangled path shows up before the money is
spent. Whitespace arrives as `.` and is refused, on `default_case`'s own
grounds that a case folder must not be the directory the user was standing in.
Plan A of two. A docling-rendered bibliography table read 22 references as 18
and judged every claim citing [6] or above against a different paper. Five
changes: the claim-level disclosure parity guard, a diacritic fold used by
`_title_tokens`, the table-row pre-pass, printed numerals checked against
position rather than renumbered, and a numbering ledger that says which labels
went missing.

Two entries the converter splits mid-title still shift [13]+ on this
manuscript. That state is disclosed, not silent, and Plan B is where per-label
agreement withholds those verdicts.
Plan B of two, from the spec approved on 2026-09-16. Nine tasks: extract the
model seam so refs may call it, a flat-text second reading of the same
bibliography, per-label agreement, withholding a disputed label's verdict, the
model's reading verified field by field, the wiring, the show-then-resolve
escalation, the invariant test, and the docs.

Three rulings worth finding here rather than rediscovering: reconcile keeps its
signature and corroboration is a second axis; the model's reading can only
subtract verdicts unless a human chose otherwise and it is recorded; and a
truncated DOI passes both DOI_RE and the verbatim check, so it needs a shape
test the spec does not have.

The pairing eval task is deliberately excluded and left to a sibling plan.
`refs` is about to read the bibliography with a model, so "check.py is the only
module that calls a model" becomes "one file shells out, and a test says which".

The judging model was a single global overwritten by every call. With two
stages calling the seam, a run whose judging made zero calls would print the
reference-list model as the judge of verdicts it never saw. The site travels
out of band in a context manager so `_ask`'s signature stays frozen — sixty-two
test sites monkeypatch it with a two-argument lambda.

The retry now reads ASK_ATTEMPTS instead of hardcoding a second attempt, and
extraction gets the retry it never had. The wizard's advertised ceiling moves
with it. tests/test_coverage.py is included in this commit because its
_parse_json_object import moved with the function, though the task-1 brief's
own file list omitted it.

🤖 Generated with Claude Code

Co-Authored-By: Claude <noreply@anthropic.com>
Three review findings. `_ask_judge` is also extraction's only call site, so the
name and docstring undersold it — a reader of check.py alone could not tell that
extraction now retries, which is what moved the wizard's advertised ceiling.

The seam test matched the literal `subprocess.run(`, so a future file shelling
out via Popen, check_output or os.system would have passed a test whose name
promises it cannot. It now matches every spelling a person would write, and its
docstring says what a grep can and cannot prove.

A zero attempt budget reached a branch whose message called itself unreachable.
It raises the type the rest of the seam raises, and says which setting did it.
…rites nothing

reconcile is about to arbitrate four readings instead of two. The cheapest new
one is a plain pymupdf parse of the PDF the run already has open — free on a
docling run, and a second opinion whose failure modes do not overlap docling's
own (a mis-rendered bibliography table versus a scrambled multi-column reading
order).

references_span_flat is a READING, not an ingest: no out_dir, no
write_outputs, no model, no network. --parse-only already carries the scar of
a read that blurred that line and risked the case's own source_map.json
describing a paper the run never finished auditing. Nothing calls the new
function yet — Task 6 wires it into reconcile.
label_agreement joins readings of the bibliography on the printed numeral, not
position — _first_divergence zips by position and would compare the wrong
entries once a reading has split or merged one. A boundary_ambiguous entry and
a reading carrying its own label twice both speak for nobody, and any
disagreement among the voters that remain is disputed outright: no majority,
matching the "a non-unique match is refused, never ranked" rule this codebase
already applies to provided files.

Nothing calls this yet. Reconciliation gains six fields and RefManifest eight
(plus RefEntry.seen_in), all defaulted and additive, so the wire format exists
before Task 4 withholds a verdict with it and Task 6 wires the two new
candidate readings that make it worth calling.
…dence

`label_agreement` compared voters with `_same_work`, which answers the opposite
question. `_same_work` asks "is there evidence these differ?" and returns True
when there is nothing to compare — correct for `_first_divergence`, which must
not manufacture a divergence out of silence. A per-label vote asks "is there
evidence these agree?", and the same silence must not manufacture agreement.

Measured: the benefit of the doubt applied when EITHER side yielded no title
tokens, so a garbled DOI-less entry paired with a legible reference came back
agreed — two readings naming unrelated papers corroborating each other, and
`agreed` is the one state that lets a verdict through untouched.

`_comparably_same` breaks the tie the other way for this caller only.
`_same_work` is untouched, because `reconcile` derives `contested` through it
and needs the old direction; a test now pins both directions at once. Reverting
the comparator fails exactly the two new laundering tests and leaves the
DOI-less-but-same-title case passing, so the tightening is not too broad.
…ntically

An identical pair exercised the overlap arithmetic trivially and would still
have passed if the 0.34 floor were raised almost to 1.0 — guarding the
structural case and nothing else. The two readings now differ the way two
converters' readings of one entry actually differ, and share two tokens of
five and seven.
A label two or more bibliography readings disagree about is dropped from the
set check_claims will judge, after the retrieval filter so withholding only
ever describes a source that really was fetched. The claim lands unchecked,
never not_retrieved — that would falsely say the source could not be obtained
when it was — and the withheld labels are named in the note.

withheld_refs is a new field, not a reuse of unjudged_refs: CLAUDE.md is
explicit that an entry in unjudged_refs means nobody read the source, and a
withheld source was read. Landing it there would erase the retrieval gap that
does exist and invent one that does not.

claim_pairing is a new claim-level disclosure with one state implemented
(withheld) and three more Tasks 6 and 7 will add as further branches in the
same producer. The reverse-direction disclosure parity test — every key a
producer emits must be declared in CLAIM_KEYS, not just the other way round —
is new in this task and would have caught claim_pairing being added without
it.
…ted set

Three review findings. The targeted-subtraction test asserted the surviving
verdict and the judged slugs, both computed from `avail` — so an edit that only
populated `withheld_refs` in the nothing-survives branch would have passed the
whole suite while silently dropping the claim_pairing disclosure for every
co-cited claim that kept a real verdict. Simulating that edit now fails the
test.

`disputed` is copied into a set rather than passed through. A generator is
truthy with no length, and each membership test consumes it, so the first claim
would withhold and every claim after it would be judged normally — a partial
withholding with nothing to notice it.

The field comment borrowed "numbering gap" for something numbering does not
mean here: `numbering_verified` is about reference positions, this is about two
readings naming different papers under one label.
Every value the model proposes is searched for, verbatim, in one of the two
extractions it was shown. A value not found is discarded; a title not found
discards the whole reading. A DOI is never repaired -- DOI_RE accepts a trailing
hyphen, so a line-broken 10.1038/s41591- passes both the regex and the verbatim
check, and only the last-character test tells a whole DOI from half of one.

Normalisation is two layers and both sides of every comparison get the same one.
A numeral proposed twice is recorded and never renumbered. The reading letter a
model answers with is matched exactly against A/B/AB, never as a substring of
its reply -- a reply like "whichever one was clearer" contains a stray "A" that
a substring test would have mistaken for reading A.

A model agreeing with a parse is a second reading, not confirmation: it read the
same document, so a reference the layout destroyed is one it may also have
missed.

ask.py, check.py, models.py and the refs manifest schema already carried the
parser move and RefEntry.seen_in from earlier work on this branch, so this
commit adds only the new module and its tests.
Three review findings.

A reply of `{"num": "9", "reading": "A"}` became a voter carrying no claim about
any work. The whole-candidate discard only fires when a title was PROPOSED and
failed verification, so a reply that omits the identity field entirely never
reached it. An entry with neither a verified title nor a verified DOI is now
dropped with a recorded reason.

The review diagnosed this as a wrong-verdict risk via `_same_work`'s trivial
agreement. That is not the live path: `label_agreement` compares through
`_comparably_same`, which refuses to call two uncomparable entries agreed, so
the phantom made the label `disputed`, not `agreed`. The hazard is the other
direction — a degenerate reply withholding verdicts it never said anything
against. Measured both ways before fixing.

Numbering findings move to their own list. A discarded field is "this value was
not printed"; a numeral duplicate or gap is "this reading's labels do not add
up", which the spec gives a different consequence — so one list reported a
count of discarded fields including things no field ever lost. No `covers` flag
comes with it, and the comment says why: `label_agreement` already enforces
"not covering", and a flag would be a published field with no consumer.

A `reading` value of the wrong type is now reported as a wrong shape rather than
folded into "the model said nothing".
Task 5's re-review found that Task 6 was written against the old provenance
shape: it read only `fields_discarded`, so every duplicate numeral and gap
`reflist.propose` reports would have been dropped before any reader saw it.

`_llm_reference_reading` now returns the provenance object, which is what the
task's own Interfaces block already claimed and its code contradicted — a
3-tuple of (entries, notes, model) can carry only one of the two kinds of
finding. The findings get their own manifest field, their own schema
description, their own console line and their own clause in the reflist
disclosure, because "3 values discarded" and "it numbered one entry twice" are
different facts.

Step 4 also claimed five new manifest fields that Task 3 already landed. It now
says so, tells the implementer to verify rather than rewrite, and adds only the
one field that is genuinely new.
The flat-text reading and the model's reading are voters. Neither reaches
reconcile's arguments or resolve_all, so no model output can cause a source to be
resolved, downloaded or judged — only a verdict to be withheld. A run whose model
reading names two different papers produces byte-identical entries to --no-llm-refs,
and that equality is a test.

The pymupdf reading is skipped when the run backend already is pymupdf: a reading
agreeing with itself is not corroboration. --parse-only keeps it and loses the
model call, because "no network" is what the flag promises.

numbering_corroborated is a second axis and numbering_verified keeps its exact
meaning. Corroboration is what lets the numbering warning stop being the only
thing a clean reference list gets told about.

claude absent, or the call failing, leaves a stated reason in the manifest and on
the console — never that the model agreed.
… happened

Review round 1 on the four-readings task. Branch on whether the model was
actually asked (ReflistProvenance.attempted / RefManifest.reflist_attempted),
never on whether the reply named itself — a call that verifies cleanly and
votes can still leave reflist_model empty, and reading that emptiness as "no
call was made" is the cardinal rule's own failure mode.

corroborating_readings now names only readings that carried a cited label,
not every reading `others` happened to hold — a reading proposing entries for
labels the body never cites abstains everywhere in label_agreement and must
not be credited as having corroborated anything.

The pymupdf-backend skip is decided where the backend is known, and reports
itself honestly instead of blaming the PDF for "only one of two extractions
had text" when both did and the second was never taken on purpose. The
wizard's own cost estimate and the --llm-refs help text now say the same
conditional thing.

The load-bearing safety test now captures what actually reaches resolve_all,
closing the half of Ruling 2 the entries-equality assertion could not see on
its own.

Plus: an IndexError on an empty converter, a disclosure level that disagreed
with its own console line, a triplicated sort key, and a duplicated PDF
fixture.
Task 7's menu option 1 took a label out of labels_disputed — so check judges it
again — and returned the entry list unchanged. The resolved entries sit in
Resolution.resolved and were computed and thrown away, leaving the verdict to be
printed against the very paper the resolution ruled against, under a report line
saying the user accepted it. The original wrong-paper bug, reached through the
one path a human consented to.

The Interfaces block said the list is changed only by option 3, which is what
made the omission read as intentional. Both now substitute, option 1 per
resolved label, and a test pins it because nothing else in the plan would
notice.
Review round 2. attempted: bool could not tell "never called" apart from
"called and failed" -- a timeout's own note didn't start with "reading
discarded", so it fell into the success branch and printed "1 value it
proposed was not found in the printed text" for a call that proposed
nothing. reflist_outcome is now one of not_attempted/failed/read, published
as REFLIST_OUTCOMES; reflist_failure carries why, kept apart from
fields_discarded so a reason can never be miscounted as a discarded value.

The claude-absent test drove needs_flat=False, which returns before
ask.claude_available is ever consulted -- the exact thing it claimed to
cover went untested, and the "not attempted" substring it matched was the
backend-skip reason. Fixed to reach the guard and assert on the reason's
content; the two reasons no longer share a prefix.

corroborating_readings now intersects the per-label voters label_agreement
itself computes, rather than approximating "carried a cited label" -- which
named a reading that skipped an entry, or whose only carrier was
boundary_ambiguous, as having agreed on all of them.

Walking every outcome by hand caught two more: a plain --no-llm-refs run
started printing a broken disclosure once outcome was always a real string,
and the console line for a reading discarded whole read like an ordinary
success. Both fixed with regression tests.
The escalation. A run that found two readings naming different papers at a
cited label writes every reading of every disputed label to
case/out/reference_disagreement.md, with the verbatim text each was read from,
and prints the path — before any question is asked. Nobody is asked to choose
blind, and the test that pins it asserts the file is on disk at the moment
Prompt.ask is called, because a test that checks afterwards passes just as well
on an implementation that asks first.

Four choices: resolve the disputed labels with a model, withhold them, adopt one
reading whole, or abort. Option 1 substitutes each resolved entry into the list
it returns — un-disputing a label while leaving the entry the resolution ruled
against would judge the label against that paper, under a report line saying a
person accepted it. A non-interactive run does not prompt: it withholds the
disputed labels and keeps every other verdict, recording choice=withheld and
chosen_by=default rather than a choice nobody made.

chosen_by is what makes this honest. It is the only place a model reading can
change what gets resolved and judged, and it is legal because a person is
recorded as having asked for it. numbering_verified is untouched by every branch.

A window onto the printed list now opens ABOVE the disputed label, not at it:
when a reading splits one entry, the damage started in the entry before the
first disputed one, so a window beginning at the label showed the symptom and
hid the cause. The constant's own comment already claimed the entries either
side were visible.

And a run no longer prints "numbering confirmed" beside a line withholding a
label's verdicts. numbering_verified means extent; a disputed label is content.
Widening the flag to cover both — the first attempt — broke two of Plan A's
merged tests, correctly: four report formats read that boolean. The fix belonged
in the sentence, not the field.
Review round 3. reflist_entries_proposed, corroborating_readings and the
console's reply-of-[] line all asked the same empty value to carry two
meanings -- exactly the class table_warnings was already documented against
("a list, [], and null -- nobody watched").

reflist_entries_proposed is now int | None: None means never recorded, 0
means measured and zero. A round-1 manifest that recorded discarded fields
but never this count no longer reads as "proposed no entries at all", which
contradicted its own short text in the repo's own fixture.

corroborating_readings can legitimately name 0 or 1 readings even when
numbering_corroborated is True (it's a per-reading fact, corroborated is
per-label) -- round 2's intersection made that state real. Fixed the
sentence, not the field: the console and the disclosure now gate the "N
readings agree" claim on the list actually having two names, under a
different token when it doesn't, rather than redefine what
numbering_corroborated means (the thing the hard constraint this round
added says not to do).

An unrecognised outcome now fails closed instead of reading as success; the
propose guard names itself instead of blaming a model that was never asked;
the report strips the same prefix the console already does; numbering
findings stop appending to a call that never returned.

Walking the four-axis table (live/loaded, written-by-an-earlier-commit)
past the rows the review named surfaced one more: a round-1-shaped "not
obtained" note, loaded from a case folder, was being read as a discarded
value by this round's own fix for the silence bug -- reopening N1 on the
load path. Split the normalization to tell a real discarded value apart
from round 1's own failure prefixes.
Review round on the escalation (c9fa042): Critical — resolve_disputed built
its RefEntry by hand and never set slug, so resolve_all downloaded every
resolved label to sources_resolved/None.pdf and check.py silently refused to
judge any of them, while the manifest said they were settled. Every test that
reached menu option 1 stubbed resolve_disputed with refs._entry, which sets
the one field the real function forgot, and so hid the bug it looked like it
was testing.

Fixed by building the resolved entry through refs._entry and overlaying the
verified fields, and by re-running _unique_slugs over the whole substituted
list. Wrote the zero-coverage tests/test_reflist.py suite this needed (17
tests against the real resolve_disputed, mocking only ask._ask), plus one
end-to-end test through the real function via _refs_pipeline.

Also from the same review: seen_in was translated through the wrong dict
(names now passed explicitly, matching propose's own label_a/label_b);
option 3 carried a discarded reading's extent claim forward (verified/
source/ledger now cleared, not widened); option 1's console asserted
abstention over a malformed reply or an unprinted title alike (now named per
the real reason resolve_disputed already records); option 3 could adopt a
reading that itself duplicated the disputed label, laundering the coin-flip
label_agreement refuses; the menu's default offer of "3" could raise
IndexError with nothing to offer; a pre-escalation line promised labels
"will be reported unchecked" moments before options 1/3 checked them; and
the escalation had no run-level disclosure, so an 11-of-14-resolved run said
so only on claims that cite one of the 11.

Eleven further minors from the same review, addressed or explicitly declined
with reasons in task-7-report.md.
…laces a convention

ADR 0002 pre-specified this remedy — a third candidate reading behind an opt-in
flag, not a replacement of the parser — and explicitly declined to authorise
it. ADR 0003 is the authorisation, with the trigger now evidenced: 22 printed
references read as 18, six substantive verdicts about papers the manuscript
never cited, every one of them carrying title_check verified. It records the
four rejected alternatives so they are not re-proposed.

CLAUDE.md's "check.py is the only module that calls a model" becomes "one file
in src/ shells out, and a test says which" — the rule was protecting the seam,
not the module. The refs.py section gains per-label agreement, why the join is
the printed numeral and not the position, and why single prints.

README gains a does bullet and a does-not bullet stating the ceiling:
corroborated is not verified, every reading read the same document, and nothing
a model returns or a user answers sets numbering_verified. Line 216 is
untouched and nothing is retracted. No accuracy figure anywhere — none has been
measured.

evals/DESIGN.md gets a pointer to the sibling eval plan, not the plan.
Most of this design's safety lives in one flag. Parametrised over eight model
reply shapes crossed with two extents (a fixture that legitimately verifies
and one that legitimately does not) — including the well-formed reply that
agrees with the parse, which is the tempting case, not the malformed one —
plus all four menu choices crossed with both starting values of the flag, plus
a resolver call that raises.

The brief prescribing this test had two defects of its own: its reply-shape
fixture asserted verified=False against a fixture that legitimately verifies
True, and its monkeypatch target (reflist_mod._ask) is never read by the real
code, which imports the whole ask module and calls ask._ask qualified — as
written it would have let every reply-shape test fall through to a live
subprocess call. Both fixed; see task-8-report.md for detail, including a
mutation from the brief's own verification step that turned out to be a no-op
against the real code.

A companion asserts numbering_corroborated CAN be True in the same run: the two
axes are independent, and a second axis that moves only with the first would
leave the 22-vs-19 false alarm firing.

Plus a grep guard over src/, whose docstring says what a grep cannot prove and
which has a self-test so it has been observed failing.
…nnot corroborate

`_comparably_same` returns `True | False | None`, the same three-valued
vocabulary `models.titles_match` publishes — the two are siblings answering
opposite questions. A two-valued `False` for "nothing to compare" made a
Crossref deposit of bare DOIs dispute every label it voted on while both
parses agreed perfectly, and a run then printed an audit with no verdicts in
it.

`_label_state` is the one place the rule lives now: a cannot-tell voter is
excluded from the comparison rather than counted as dissent, which can leave
one reading standing and print `single`. A label NO voter says anything
comparable about stays `disputed` — it is carried, and nothing establishes
what it names.

`DERIVED_READINGS` excludes the model's reading from crediting an agreement.
`reflist.propose` may only copy from the two extractions that vote here in
their own right, so its agreement is one text read twice. It still votes and
can still dispute.

`stamp_seen_in` matches through the same comparator, so `seen_in` and
`labels_disputed` can no longer contradict each other on one manifest; it
credits a chosen entry to the reading it came from by identity and merges
rather than overwrites, so `[]` keeps one meaning.
… reach disputed

`c.unjudged_refs` is decided above every branch of the claim loop, so the
withheld path can no longer `continue` past it: a claim citing a withheld [3]
and a paywalled [4] reported [4] in none of the three accounts a reader has.
It stays empty in exactly one state — a claim where nothing was judged,
withheld or provided — where the `not_retrieved` verdict is already the whole
report.

`_claim_pairing` splits a disputed label on the manifest entry's own status,
not on the emptiness of `withheld_refs`: a label nobody retrieved gets its own
token and its own sentence instead of "the source fetched under that label",
and a `results.json` written before `withheld_refs` existed still reads as the
fetched case.

`label_agreement` returns `disputed` for two causes and no field tells them
apart, so no surface asserts the contradicting one: the note, both
disclosures, the console, the disagreement file and both schema descriptions
name the disjunction.
…s who made it

`resolve_disputed` is passed `disputing` — what each reading says at each
disputed label, including the deposit and the model's own reading, which have
no printed span and were therefore invisible to a call adjudicating between
them. It is context: `_found` still verifies every returned value against the
printed extraction alone. The prompt renders one block per extraction it
actually has, so a pymupdf-backend run no longer shows an empty `READING B`.

Five `resolution_*` manifest fields carry that call's own provenance. A run
whose backend skips the reading and then escalates recorded
`reflist_outcome: not_attempted` and nothing else, so the report said no model
read the reference list on a run where one was asked, verified field by field
and allowed to change which papers are judged. The `reflist` disclosure now
names it in every branch, and fires on its own when it was the only call.

The note, the run-level resolution and the per-claim disclosure name the
extractions the model was shown instead of asserting "both texts".

Also: `numbering_corroborated` is `bool | None`, so "never computed" is a
value and not a comment; `stamp_seen_in` runs after the escalation, so option
3 no longer publishes an empty provenance for a run that computed one; option
3's note says the reading was adopted whole; the resolution reply's entry
count is taken before de-duplication; a mid-branch `reflist_attempted` keeps
its disclosure without claiming an outcome nobody recorded; and both calls
refuse an oversized bibliography before spending anything rather than
clipping it into a reading quietly short.
Found by enumerating the states rather than the branches: the not-obtained,
failed and unrecognised heads end "the reference numbering rests on the
readings above it and nothing else", and the resolution sentence appended
after them describes a call that substitutes an entry. The claim is dropped
where a resolution happened and kept where none did, since it is
load-bearing there.

`_also_resolved` is one producer for the sentence both dispute states owe, so
a claim citing an unretrieved disputed label and a resolved one no longer
loses the resolved caveat to the state added later.
`resolution_readings == []` is never recorded, and on this branch that
identifies a manifest written before these fields existed — whose call was
also not shown what each reading said at the disputed label. Asserting
either half of this build's fuller account over that run is the report
describing behaviour the run did not have. `_resolution_sentence` is the one
producer for both surfaces.
…dispute

`_title_tokens` strips DOIs the way it already strips URLs, for the reason
its own docstring gives: `10.1148/radiol.…` yielded `radiol` and
`10.1001/jamanetworkopen.…` yielded `jamanetworkopen`, and those identifier
substrings were compared against real title words. Only a DOI with no
five-letter run reached the abstention added in round 1, so the common
publisher shape still disputed every label a bare-DOI deposit voted on.

Two consumers move, both toward honesty and both measured in the report:
`_same_work` stops reporting a false divergence at a DOI-only deposit, and
`_title_check_text` stops carrying a token no first page prints in its
denominator — which on a short reference whose slug DID match can now reach
`mismatch`, correctly, where a spurious match used to hold it at
unverifiable.

`labels_uncomparable` is a published subset of `labels_disputed`: the labels
withheld because nothing could be compared rather than because the readings
named different papers. `labels_disputed` stays the only list that drives
withholding, the subset is read off the same `_label_state` call so the two
cannot contradict each other, and it is re-intersected once after the
escalation so a resolved label cannot survive in it. `None` is never
computed and `[]` is a measurement. Every surface — the run-level
disclosure, both claim-level dispute states, the withholding note and the
console — names the cause through one producer, `disclosures.dispute_causes`,
and falls back to the disjunction only where nothing measured the split.

`_label_key` moves down to `refs`, the layer below both its users.
…odel call

The `--llm-refs` reading is a voter and reaches neither `reconcile`'s arguments
nor `resolve_all`. Three files generalised that to "no model output can cause a
source to be resolved, downloaded or judged", which the escalation's option 1
falsifies: `_escalate_disputed` substitutes entries `reflist.resolve_disputed`
built from a model's reply, `resolve_all` downloads them and `check.py` judges
them. The substitution is mandatory — returning the list untouched un-disputes
the label while leaving the entry the resolution ruled against as the paper
judged. The exception is now stated where a reader meets it, with its gate: an
interactive run, the disagreement on disk before anything is asked,
`numbering_chosen_by: "user"`, only the resolved labels substituted, and no
answer setting `numbering_verified`.

Constraint 1 cited a module that tests constraint 5; it now cites
`test_a_model_naming_different_papers_changes_no_resolved_entry`. The
escalation's stated precondition ("nothing accounted for the labels AND labels
in dispute") does not exist — `_escalate_disputed` returns early on an empty
`labels_disputed` and nothing else. A disagreement withholds the source
always, the claim only when no other source answered.

Also: `_comparably_same`, not `_same_work`, and why; the third value and the
three dispute causes; `labels_uncomparable` as a published subset;
`numbering_corroborated` as `bool | None`; the DOI strip in `_title_tokens`;
`reflist.py` rather than `refs.py` as the seam's second caller;
`reconcile`'s two candidate readings rather than its arity; the whole reading
discarded by one unverifiable title; the corroboration bullet's other branch;
`CI` and a piped stdout as non-interactive; no integer for the seam's patch
sites, which moved again today. Two wire-format rules earned here go into the
cardinal-rule section.
… anything

The seam comment claimed sixty-two monkeypatch sites. That was raw grep hits on
`_ask`; the real figure measured 46 at review time and 48 a day later, because
it grows with every test that touches the seam. Three values for one claim is
the tell — the number is a maintenance cost pretending to be evidence, so the
comment now states the constraint without one.

ADR 0002 cross-referenced README §"Two independent readings" by title. 0.7.0
took the count from two readings to four and the section was renamed, so the
reference did not break loudly, it simply stopped naming anything. It now
points at the section that exists and forward to 0003, which is the
authorisation 0002 explicitly withheld.
… wrote

The reply-shape half of this module built `candidates` and called `reconcile`
by hand, so "the model's reading never becomes one of `reconcile`'s arguments"
was enforced by the test's own source text. Handing the `llm` reading to
`reconcile` in `_refs_pipeline` — the design's central prohibition — left all
29 cells green. So did disabling `reflist._found`, and so did widening the
manifest line to `rec.verified or rec.corroborated`.

Every cell now drives `cli._refs_pipeline` with `ask._ask` patched per reply
shape and reads `refs_manifest.json`. `reflist.propose`, `resolve_disputed`,
`reconcile`, `label_agreement` and the manifest construction all run for real;
the disputed label is produced by the pipeline rather than assigned by the
test, and the raising-resolver cell raises at the seam instead of stubbing the
resolver. Each shape carries what it is allowed to have moved, because an
invariant asserted alone cannot tell "the reply changed nothing" from "the
reply never ran".

The grep splits on the first `=` that is not a comparison rather than
discarding any line containing one, and `corroborated` joins `_FORBIDDEN`; the
self-test plants all three spellings and requires all three caught.
A reply naming only [99] makes every numeral below it absent, which printed 98
consecutive integers — 410 characters — into reflist_numbering_findings, a
field the report shows a reader. The first six locate the gap; the rest only
prove the reader stopped reading. Nothing is hidden and the finding is
unchanged: the total leads the sentence.
… side

`limits` (--max-claims, --max-sources) and this branch's per-label agreement
both decide whether a cited source is judged, so the merge had to keep three
reasons a source yields no verdict apart, not two:

  not obtainable      retrieval failed          fix access, or accept the gap
  skipped on request  you asked for a slice     raise the limit and re-run
  withheld            the readings disagree     check that citation by hand

Verified end to end: paywalled reports `not available (paywalled)`, a capped
source `not available (skipped on request)`, and a disputed label `unchecked`
with the label named. Each says which happened; none is summarised into
another.

They compose by construction: `skipped` is not in ("retrieved", "provided"),
so a skipped entry never reaches the partition that withholding splits, and a
source can never be both.

Resolutions worth knowing. cli.py took both features' flags on `refs` and
`run`, and the stage order `limits` introduced survives —
ingest → extract → refs → scout → check → highlight → report — with extraction
before references so a claims limit retrieves only what the selected claims
cite, and the escalation still inside `_refs_pipeline`. check.py kept this
branch's withheld branch with `limits`' skipped-aware reasons inside it.
CLAUDE.md kept this branch's `ask.py` seam rule and dropped the old "check.py
is the only module that calls a model", which is now false, while taking
`limits`' paragraph on how limits reach `resolve_all`. README took the shorter
rewrite from #6/#7 as the base and swapped in the four-readings bullet, since
theirs still described two. One 0.7.0 section in the changelog carries both.

Two tests needed the merged shape rather than either side's: the wizard's cost
estimate, because the reference-list call survives every limit and `apply_limits`
already accounted for it; and the flag-forwarding test, which had to stub the
new extract stage.

1170 passed, ruff clean.
All ten CI jobs failed on two tests that pass locally. `_check_pipeline` calls
`cli._require_claude()`, which does `from .check import claude_available` — and
`check.claude_available` is a re-export, a separate module attribute from
`ask.claude_available` even though both name one function. The tests patched
`ask_mod`, so the binding this path actually reads was never touched: they
passed on a machine with the CLI installed and failed on ten that have none.

The merge caused it. Before it, `_check_pipeline` imported from `.ask` and the
patch was right; `limits` refactored the check into `_require_claude`, which
imports from `.check`, and the comment above the patch went on describing the
old import. Patching `_require_claude` itself cannot drift with a future
refactor of which module it reads.

Verified the way CI does, which is the only way this shows: with `claude` off
PATH, 1145 passed, 25 skipped.
@defraction0
defraction0 merged commit 6966aee into main Sep 19, 2026
10 checks passed
@defraction0
defraction0 deleted the feat/llm-reference-list branch September 19, 2026 11:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant