Skip to content

Add MCP server: drive PaperTrace from Claude Desktop, Claude Code and other local MCP hosts - #9

Merged
defraction0 merged 13 commits into
mainfrom
claude/gifted-ritchie-ga60nx
Sep 27, 2026
Merged

defraction0 merged 13 commits into
mainfrom
claude/gifted-ritchie-ga60nx

Conversation

@defraction0

@defraction0 defraction0 commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

papertrace mcp serves PaperTrace over stdio to any app that can start a local MCP server: Claude Desktop, Claude Code, Cursor, VS Code (GitHub Copilot in Agent mode), Windsurf, Gemini CLI and Codex CLI. ChatGPT and Claude in a web browser cannot use it: they reach MCP servers over the internet only. This is the roadmap item drive PaperTrace as a tool from any MCP-capable client.

Whichever app drives the tools, every claim is judged with claude -p, so Claude Code must be installed and signed in on the same computer (a paid Claude plan). Reading an audit that already exists needs no Claude Code.

Tools

Tool Kind What it does
start_audit the only tool that writes or spends runs the papertrace run pipeline on a background thread and returns at once
audit_status read-only not_started, running, finished, failed or interrupted, with the log tail; waits up to 50 s per call
audit_summary read-only counts, references, the judging model, coverage and every run-level disclosure; a limited run's scope first and last
list_claims read-only one row per claim (the paper's sentence, the headline verdict, the cited labels, its caveats), optionally one verdict only, with the run-level disclosures
get_claim read-only one claim in full: each source's verdict, rationale, page, block and anchor phrases, with every caveat, run-level ones included
get_evidence read-only the red-box crops as images, captioned with the anchor state
list_references read-only the retrieval manifest and numbering state, without local file paths
list_gaps, get_scout read-only uncited assertions, citation places no claim reached, unchecked claims; the scout's registers with its caveat

How the rules in CLAUDE.md carry over

  • The server computes no verdict. The read tools serialise what the pipeline wrote, with the same run_disclosures / claim_disclosures / judgement_disclosures the four report formats use, through the viewer's report._disclosure_dict. The disclosure parity test now covers this fifth reader.
  • Absent is not zero. A missing or unreadable results.json is refused, never answered with empty counts. The manifest's three-state fields pass through, so null stays null. A coverage audit that never ran is null.
  • Every refusal is a ToolError carrying its reason. In SDK v2, any other exception reaches the host as a bare "Error executing tool".
  • Nobody can be asked. During a job sys.stdin is an empty stream, so a disputed reference label is withheld (numbering_chosen_by: "default"), never settled. MCP elicitation is deliberately not offered in its place: the gate is a person, and a host may be a model.
  • ask.py is still the only file in src/ that starts a process. The audit runs in-process on one thread, one audit per server. Each audit starts with ask.forget_models(), so an audit that judged nothing cannot report the previous audit's model.
  • Nothing but the protocol reaches stdout. The pipeline's console becomes the job's log.
  • No path is served that the tool did not write. Evidence is served only as a .png inside <case>/out/, and no pdf_path leaves the server.
  • An unfinished case is not a result. Each audit writes <case>/mcp_audit.json: running, refreshed every 15 s, then the outcome. The read tools refuse a case whose last MCP audit is running, failed or never finished, because its files may mix two runs. A record not refreshed for 60 s is interrupted. A results.json written after a failed audit supersedes the failure, and a scout.json written after an audit that ran without the scout supersedes that.

Changes

  • src/papertrace/mcp_server.py (new): the nine tools, the job registry and serve().
  • src/papertrace/cli.py: run is split into the keyword-only _run_pipeline and its Typer adapter, with no behaviour change. It gains the papertrace mcp command, which exits 2 with the install line on stderr when the SDK is missing or 1.x. SCOUT_CAVEAT is now one shared constant.
  • src/papertrace/ask.py: forget_models().
  • src/papertrace/disclosures.py: coverage_audited(), extracted unchanged from run_disclosures so the reports and the server share one rule.
  • pyproject.toml: a new [mcp] extra (mcp>=2.2,<3), also carried by [full] and by [dev], so CI runs the MCP tests instead of skipping them.
  • schemas/mcp_tools.schema.json and schemas/mcp_audit.schema.json (new).
  • Tests: tests/test_mcp_server.py (new); additions to tests/test_ask.py, tests/test_packaging.py and tests/test_pipeline_split.py.
  • Docs:
    • README From an MCP host: which apps can and cannot run it, a no-clone install (uv tool install from the GitHub archive), each app's configuration and the Windows notes.
    • CHANGELOG.md (Unreleased) and CLAUDE.md.
    • The spec, plan and ADR: docs/superpowers/specs/2026-09-26-mcp-server-design.md, docs/superpowers/plans/2026-09-26-mcp-server.md, docs/adr/0004-mcp-server.md.

Gates in CLAUDE.md

  1. Tests stay offline: met. Tools are exercised through the SDK's in-memory Client, with the pipeline faked at cli._run_pipeline. One real stdio round trip launches the server itself; it makes no model calls.
  2. New JSON means a schema and a round-trip test: this applied, and it is met. mcp_audit.json and every tool output have schemas. Outputs and every record state are validated against them, and each schema definition's keys are held equal to its TypedDict's.
  3. No PDF or case folder committed: met. Test PDFs are generated in the tests.
  4. refs.py, check.py, scout.py: untouched.
  5. Ingest, highlight and report logic: untouched. report.py is not in the diff, and disclosures.py only gains the extracted function. So the end-to-end demo was not re-run.

Verification

  • CI is green on 3.10, 3.11, 3.12, 3.13 and 3.14 on the head commit.
  • ruff check src tests scripts evals is clean.
  • Locally, the whole suite passes except tests/test_refs.py::test_the_contact_email_reaches_only_the_services_that_need_it. It fails behind the development sandbox's HTTP proxy (httpx.ProxyError: 403) on main as well, and passes in CI.
  • By hand, in a Linux sandbox:
    • Install: uv tool install --python 3.12 "papertrace[mcp] @ <this branch's GitHub archive>" installed it, and papertrace --version printed papertrace 0.7.1.
    • Launched as an app would, by absolute path with a minimal environment, through the SDK's stdio client:
      • nine tools listed, and audit_status answered not_started;
      • start_audit was refused while PATH did not reach claude;
      • with PATH fixed and no email configured, it was refused on the email instead. Nothing was spent and no case folder was created.
    • Claude Code 2.1.283: claude mcp add --env PAPERTRACE_EMAIL=… --transport stdio --scope user papertrace -- <path> mcp, then claude mcp list reported the server Connected. A name placed straight after --env is rejected ("Invalid environment variable format"), which the README warns about.
    • Update: uv tool install --reinstall-package papertrace … re-fetched only PaperTrace, in 11 s.

Tried by the author

After this description was first written, the author installed PaperTrace from GitHub on their own computer, set it up in Claude Code and Claude Desktop, and reported that it worked.

Not verified

  • No live audit was run in the sandbox, because it would spend model calls. The pipeline is the one papertrace run runs; the job around it is tested with the pipeline faked.
  • Cursor, VS Code, Windsurf, Gemini CLI and Codex CLI were not run. Their configurations follow each app's documentation.
  • Beyond the author's own computer, macOS and Windows were not tested; the sandbox is Linux. The install size (6 GB, most of it PyTorch and its GPU libraries) was measured on Linux only.
  • The end-to-end demo was not re-run (gate 5 above).

Deliberate omissions

  • --png is not exposed. It needs playwright and a browser, and the images the server returns are the evidence crops, which need neither.
  • A running audit cannot be cancelled: a thread cannot be stopped safely mid-write.
  • The job's log is not persisted. Its outcome is, in mcp_audit.json.
  • Liveness across processes is a heartbeat window, not a lock.
  • stdio only: no Streamable HTTP, no per-stage tools, no MCP resources.

Review

The branch had two review passes: 20 findings, 19 fixed test-first and 1 rejected with evidence. The plan records them in Tasks 7 and 8.

🤖 Generated with Claude Code

https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf

The roadmap's "drive PaperTrace as a tool from any MCP-capable client",
designed before any code. Nine tools over stdio behind an optional extra:
seven read the case folder with the same disclosures every report format
carries, and one pair runs the existing pipeline as a background job.

The spec walks each CLAUDE.md rule across to the two things that are new
here — a model as the reader, and a process that outlives many audits —
and records what was rejected: a subprocess per audit (a second spawning
file in src/), a blocking run (a TypeScript-SDK host times out at 60 s),
and elicitation to settle a disputed label (the gate is a person).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
`run` was the last stage still reachable only through its Typer command.
The MCP server's audit job is about to call it, and a caller that omits
an option from a Typer command receives an OptionInfo for it — the flaw
cli.py records shipping twice. The body moves unchanged into a
keyword-only function with plain defaults; `run` passes every parameter
by keyword, and a test compares the two signatures so neither can gain
a parameter the other does not.

No behaviour change: the existing run tests fake the inner stages and
still exercise the moved body.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
The per-site model record was correct because one process ran one
command, so it always began empty. A long-lived process — the MCP
server — runs audit after audit, and a second audit whose judging made
no calls would print the first audit's judge as its `Checker:`.
`forget_models()` clears the record; the server calls it at the start of
every audit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
Seven read-only tools over a case folder — audit_summary, list_claims,
get_claim, get_evidence, list_references, list_gaps, get_scout — built on
the official SDK's MCPServer. They compute nothing: they load what the
pipeline wrote and serialise the disclosures disclosures.py decided, with
the viewer's own _disclosure_dict, so a model reading the result gets the
same caveats every report format carries. The parity loops extend
test_disclosure_parity.py's contract to this fifth reader.

What the report formats promise, kept here: the scope sentence first among
the disclosures and again as the last key; a missing or unreadable
results.json refused with its reason rather than returned as empty counts;
three-state manifest fields passed through with their null; no local PDF
path; an evidence crop served only from inside <case>/out/, so an edited
results.json cannot point the server at another file.

Found on the way: below Python 3.12 pydantic cannot build a schema from
typing.TypedDict, and the SDK then drops the tool to unvalidated text
without a word. Every tool here published no output schema on 3.10 and
3.11 until the test asked; typing_extensions.TypedDict fixes it and
structured_output=True makes any repeat an error at registration.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
start_audit runs `papertrace run`'s pipeline (via cli._run_pipeline) on a
background thread and returns at once; audit_status reports running,
finished or failed with the log tail, long-polling at most 50 s. An audit
takes minutes and a TypeScript-SDK host times a request out at 60 s, so a
blocking call would fail while the audit carried on spending unseen.

What the job guarantees, each asserted from inside it by a test that fakes
a terminal on both streams:
- nobody can be asked: stdin is an empty stream for the job, so
  cli._interactive() is False and a prompt reads end-of-file. A disputed
  reference label is withheld with chosen_by "default", as the CLI does
  without a terminal; the gate on settling one is a person, and a host
  may be a model.
- no model from an earlier audit lingers: ask.forget_models() first.
- the pipeline's console becomes the job's log instead of reaching
  stdout, which under stdio is the protocol's wire.

One audit at a time, because the console, the model record and stdin are
process-global; a second start_audit is refused naming the running one,
and read tools refuse a case whose audit is running, since what is on
disk then belongs to an earlier run. Everything that can refuse — no
manuscript, no `claude` on the server's PATH, no contact email, a case
folder holding another paper — refuses before anything is spent. A
failed job keeps the reason the pipeline printed, or the exception's
type and message, which SDK v2 would otherwise reduce to a bare "Error
executing tool". A finished status carries no counts: they travel only
with their disclosures, in audit_summary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
…xtra

The command a host launches. It writes nothing to stdout — under stdio
that is the protocol's wire, and a banner printed before serving begins
corrupts the first message — and without the SDK it says how to install
it on stderr and exits 2. Only the SDK's absence earns that line; any
other missing module keeps its own traceback.

`mcp>=2.2,<3` is an extra, and the same requirement is in [full] (the
quick start's everything-install) and in [dev], which is exactly what CI
installs: the MCP tests skip without the SDK, and a module that skips in
CI is a module nobody runs.

Two tests go through a real process, because the in-memory client skips
exactly the part that can break here: one asserts zero bytes on stdout
when stdin closes, the other launches the command with the SDK's stdio
client, lists the nine tools and reads a case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
…o stdout

The other job tests replace _run_pipeline whole. This one stubs only its
seven stages, so the real body runs inside the job — email, case, guard,
the case folder's own .gitignore, banner, stage order, closing line — and
reads the process's real stdout for anything printed around the captured
console. Integration coverage for the job already committed, not a new
behaviour: it passed on first run, and fails when the console capture is
removed (checked by mutation, as were the stdin swap and the model reset).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
README: a fourth way in, a section for MCP hosts (the command by absolute
path, Claude Code, Claude Desktop, Cursor and VS Code configuration, the
nine tools, what start_audit spends and what it refuses before spending),
the install-matrix row, one line on each side of the does/does-not list —
it serves the audit; it cannot settle a disputed label — and the roadmap
item ticked. "Claude Desktop integration" stays open: a one-click install
may be what it means, and that is not what this is.

CHANGELOG: an Unreleased entry. CLAUDE.md: the commands, why the SDK is in
[dev], forget_models on the ask.py bullet, the server as the disclosures'
fifth reader, and an mcp_server.py bullet recording what a future change
must not undo — in-process, one audit at a time, nobody to ask, no
elicitation in place of the person-gated escalation, and every new piece
of process-global pipeline state reset per audit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
…t its caveats

Fixes from a code review of this branch, each verified against the code
and each with a test that failed first.

- A failed or interrupted audit left one run's manifest beside another's
  verdicts, and the read tools served them. The job now writes
  <case>/mcp_audit.json ("running" before the pipeline starts, the
  outcome when it ends); read tools refuse a case whose last MCP audit
  failed or never finished, and audit_status reports an earlier server's
  audit from the record — "interrupted" when it says "running". A log
  that was not kept is null, not an empty list.
- get_claim and list_claims now carry the run-level disclosures, and the
  evidence caption carries the claim's and the run's: a host can read a
  verdict without ever calling audit_summary, and an unconfirmed source
  identity or a doubtful label belongs beside it.
- With no case given and the paper's folder unwritable, start_audit
  refuses instead of falling back to the server's working directory,
  which the host chose; what the preflight printed opens the job's log.
- get_scout refuses a scout.json the latest audit did not write (it ran
  without the scout), and its caveat carries the scan's own error.
- `papertrace mcp` names the fix for a 1.x SDK too, which fails with a
  plain ImportError, and no longer wraps the install command in two; the
  test module skips on an SDK below 2 instead of failing collection.
- Gate 2: schemas/mcp_tools.schema.json and schemas/mcp_audit.schema.json,
  with every tool output validated against the first, each definition's
  keys held equal to its TypedDict's, and the record validated against
  the second.
- One rule, not two copies: disclosures.coverage_audited() decides
  whether coverage ran for the reports and the server alike, and
  cli.SCOUT_CAVEAT is the scout's sentence for both.

Not taken: the finding that _run_pipeline's OptionInfo guards are dead.
tests/test_pipeline_split.py calls cli.run() directly and omits options,
which hands the adapter's sentinels straight through.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
The spec gains its eleventh rule — a case is not a result until its audit
has ended — and loses two claims the review showed wrong: that stopping
the server stops the audit (a `claude -p` call in flight runs to its
end), and that job state is not persisted (mcp_audit.json now carries
the outcome). The rejected "schema only in tools/list" bullet now says
what gate 2 got instead. Two limits are written down rather than left to
be found: two server processes auditing one case are not detected, and a
CLI run that fails midway leaves no record.

README, CHANGELOG and CLAUDE.md say the same, and the plan records the
review round, including the finding that was not taken and why.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
A second review, of the first round's fixes, found ten more; each held up
against the code and each has a test that failed first.

- A "running" record was reported "interrupted" whatever its age, which
  asserted a cause nobody measured: with two hosts each launching their
  own server, the other server may be alive, and "call start_audit"
  would put a second pipeline in its folder. The job now refreshes its
  record every 15 s; fresh means live in another process — read tools
  and start_audit refuse, audit_status reports "running" and waits on the
  record — and only a record unrefreshed for 60 s is "interrupted".
- The disk record now decides, not this server's memory of its jobs: a
  case completed since by other means reads again. A failed audit is
  superseded by a results.json written after it, and an audit run without
  the scout by a scout.json written after it — the record's mtime is the
  measurement.
- The evidence caption for a named source carried the headline source's
  anchor state beside its own; it now carries only the shown source's.
- list_gaps reported unreached_labels [] for a label audit that never ran
  (no clean.md, occurrences counted); it is null unless labels were read.
- audit_status's manuscript was a full path from this server and a file
  name from a record; it is the absolute path in both, and the record
  stores that.
- A recorded failure's next step sent the host to read a log that was not
  kept; recorded states have their own wording.
- The record is validated — schema version, state, scout as a bool —
  before anything is decided on it; "false" is not False.
- `papertrace mcp` re-raises an import failure when a 2.x SDK is
  installed, instead of an install line contradicting itself over a
  hidden traceback.
- The published schema describes record-derived states; a test validates
  every state a record can produce, and another holds the prose timings
  to the constants.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
The spec's eleventh rule now says how "running" is measured and that the
record on disk, not a server's memory, decides. Its omissions lose "two
server processes auditing one case are not detected", which stopped being
true, and gain the heartbeat's real limit instead: a machine that slept
past the 60 s window reads as stopped — its case refused, the safe side —
while a second audit started inside that window is not.

README, CHANGELOG and CLAUDE.md say the same; the plan records the pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
The README named four hosts and installed from a clone. It now lists the
apps that can start a local MCP server (Claude Desktop, Claude Code,
Cursor, VS Code, Windsurf, Gemini CLI, Codex CLI) with where each takes
its configuration, and says plainly that ChatGPT and Claude in a browser
cannot: they reach MCP servers over the internet only.

Whichever app drives the tools, every claim is judged through `claude -p`,
so Claude Code has to be installed and signed in — said once, beside the
list, rather than left for `start_audit`'s refusal to reveal.

Install is `uv tool install` from the GitHub archive: no Git, and uv
fetches a Python. Verified in a sandbox: the install, the executable
launched by absolute path with a minimal environment (nine tools;
start_audit refused while PATH did not reach `claude`, and refused on the
email next once it did — nothing spent), and Claude Code's `claude mcp
list` reporting it connected. That last is what "tested" now claims; the
other apps' configurations follow their docs and have not each been run.

Also: `claude mcp add` takes `--env` as variadic, so a name straight after
it is read as another KEY=value (reproduced); PowerShell's `where` is
Where-Object, so `where.exe`; a Windows path in JSON doubles each `\`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
@defraction0 defraction0 changed the title Add MCP server: drive PaperTrace from any MCP-capable host Add MCP server: drive PaperTrace from Claude Desktop, Claude Code and other local MCP hosts Sep 27, 2026
@defraction0
defraction0 merged commit 4f1c3ac into main Sep 27, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants