Repository navigation
Add MCP server: drive PaperTrace from Claude Desktop, Claude Code and other local MCP hosts - #9
Merged
Merged
Conversation
The roadmap's "drive PaperTrace as a tool from any MCP-capable client", designed before any code. Nine tools over stdio behind an optional extra: seven read the case folder with the same disclosures every report format carries, and one pair runs the existing pipeline as a background job. The spec walks each CLAUDE.md rule across to the two things that are new here — a model as the reader, and a process that outlives many audits — and records what was rejected: a subprocess per audit (a second spawning file in src/), a blocking run (a TypeScript-SDK host times out at 60 s), and elicitation to settle a disputed label (the gate is a person). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
`run` was the last stage still reachable only through its Typer command. The MCP server's audit job is about to call it, and a caller that omits an option from a Typer command receives an OptionInfo for it — the flaw cli.py records shipping twice. The body moves unchanged into a keyword-only function with plain defaults; `run` passes every parameter by keyword, and a test compares the two signatures so neither can gain a parameter the other does not. No behaviour change: the existing run tests fake the inner stages and still exercise the moved body. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
The per-site model record was correct because one process ran one command, so it always began empty. A long-lived process — the MCP server — runs audit after audit, and a second audit whose judging made no calls would print the first audit's judge as its `Checker:`. `forget_models()` clears the record; the server calls it at the start of every audit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
Seven read-only tools over a case folder — audit_summary, list_claims, get_claim, get_evidence, list_references, list_gaps, get_scout — built on the official SDK's MCPServer. They compute nothing: they load what the pipeline wrote and serialise the disclosures disclosures.py decided, with the viewer's own _disclosure_dict, so a model reading the result gets the same caveats every report format carries. The parity loops extend test_disclosure_parity.py's contract to this fifth reader. What the report formats promise, kept here: the scope sentence first among the disclosures and again as the last key; a missing or unreadable results.json refused with its reason rather than returned as empty counts; three-state manifest fields passed through with their null; no local PDF path; an evidence crop served only from inside <case>/out/, so an edited results.json cannot point the server at another file. Found on the way: below Python 3.12 pydantic cannot build a schema from typing.TypedDict, and the SDK then drops the tool to unvalidated text without a word. Every tool here published no output schema on 3.10 and 3.11 until the test asked; typing_extensions.TypedDict fixes it and structured_output=True makes any repeat an error at registration. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
start_audit runs `papertrace run`'s pipeline (via cli._run_pipeline) on a background thread and returns at once; audit_status reports running, finished or failed with the log tail, long-polling at most 50 s. An audit takes minutes and a TypeScript-SDK host times a request out at 60 s, so a blocking call would fail while the audit carried on spending unseen. What the job guarantees, each asserted from inside it by a test that fakes a terminal on both streams: - nobody can be asked: stdin is an empty stream for the job, so cli._interactive() is False and a prompt reads end-of-file. A disputed reference label is withheld with chosen_by "default", as the CLI does without a terminal; the gate on settling one is a person, and a host may be a model. - no model from an earlier audit lingers: ask.forget_models() first. - the pipeline's console becomes the job's log instead of reaching stdout, which under stdio is the protocol's wire. One audit at a time, because the console, the model record and stdin are process-global; a second start_audit is refused naming the running one, and read tools refuse a case whose audit is running, since what is on disk then belongs to an earlier run. Everything that can refuse — no manuscript, no `claude` on the server's PATH, no contact email, a case folder holding another paper — refuses before anything is spent. A failed job keeps the reason the pipeline printed, or the exception's type and message, which SDK v2 would otherwise reduce to a bare "Error executing tool". A finished status carries no counts: they travel only with their disclosures, in audit_summary. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
…xtra The command a host launches. It writes nothing to stdout — under stdio that is the protocol's wire, and a banner printed before serving begins corrupts the first message — and without the SDK it says how to install it on stderr and exits 2. Only the SDK's absence earns that line; any other missing module keeps its own traceback. `mcp>=2.2,<3` is an extra, and the same requirement is in [full] (the quick start's everything-install) and in [dev], which is exactly what CI installs: the MCP tests skip without the SDK, and a module that skips in CI is a module nobody runs. Two tests go through a real process, because the in-memory client skips exactly the part that can break here: one asserts zero bytes on stdout when stdin closes, the other launches the command with the SDK's stdio client, lists the nine tools and reads a case. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
…o stdout The other job tests replace _run_pipeline whole. This one stubs only its seven stages, so the real body runs inside the job — email, case, guard, the case folder's own .gitignore, banner, stage order, closing line — and reads the process's real stdout for anything printed around the captured console. Integration coverage for the job already committed, not a new behaviour: it passed on first run, and fails when the console capture is removed (checked by mutation, as were the stdin swap and the model reset). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
README: a fourth way in, a section for MCP hosts (the command by absolute path, Claude Code, Claude Desktop, Cursor and VS Code configuration, the nine tools, what start_audit spends and what it refuses before spending), the install-matrix row, one line on each side of the does/does-not list — it serves the audit; it cannot settle a disputed label — and the roadmap item ticked. "Claude Desktop integration" stays open: a one-click install may be what it means, and that is not what this is. CHANGELOG: an Unreleased entry. CLAUDE.md: the commands, why the SDK is in [dev], forget_models on the ask.py bullet, the server as the disclosures' fifth reader, and an mcp_server.py bullet recording what a future change must not undo — in-process, one audit at a time, nobody to ask, no elicitation in place of the person-gated escalation, and every new piece of process-global pipeline state reset per audit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
…t its caveats
Fixes from a code review of this branch, each verified against the code
and each with a test that failed first.
- A failed or interrupted audit left one run's manifest beside another's
verdicts, and the read tools served them. The job now writes
<case>/mcp_audit.json ("running" before the pipeline starts, the
outcome when it ends); read tools refuse a case whose last MCP audit
failed or never finished, and audit_status reports an earlier server's
audit from the record — "interrupted" when it says "running". A log
that was not kept is null, not an empty list.
- get_claim and list_claims now carry the run-level disclosures, and the
evidence caption carries the claim's and the run's: a host can read a
verdict without ever calling audit_summary, and an unconfirmed source
identity or a doubtful label belongs beside it.
- With no case given and the paper's folder unwritable, start_audit
refuses instead of falling back to the server's working directory,
which the host chose; what the preflight printed opens the job's log.
- get_scout refuses a scout.json the latest audit did not write (it ran
without the scout), and its caveat carries the scan's own error.
- `papertrace mcp` names the fix for a 1.x SDK too, which fails with a
plain ImportError, and no longer wraps the install command in two; the
test module skips on an SDK below 2 instead of failing collection.
- Gate 2: schemas/mcp_tools.schema.json and schemas/mcp_audit.schema.json,
with every tool output validated against the first, each definition's
keys held equal to its TypedDict's, and the record validated against
the second.
- One rule, not two copies: disclosures.coverage_audited() decides
whether coverage ran for the reports and the server alike, and
cli.SCOUT_CAVEAT is the scout's sentence for both.
Not taken: the finding that _run_pipeline's OptionInfo guards are dead.
tests/test_pipeline_split.py calls cli.run() directly and omits options,
which hands the adapter's sentinels straight through.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
The spec gains its eleventh rule — a case is not a result until its audit has ended — and loses two claims the review showed wrong: that stopping the server stops the audit (a `claude -p` call in flight runs to its end), and that job state is not persisted (mcp_audit.json now carries the outcome). The rejected "schema only in tools/list" bullet now says what gate 2 got instead. Two limits are written down rather than left to be found: two server processes auditing one case are not detected, and a CLI run that fails midway leaves no record. README, CHANGELOG and CLAUDE.md say the same, and the plan records the review round, including the finding that was not taken and why. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
A second review, of the first round's fixes, found ten more; each held up against the code and each has a test that failed first. - A "running" record was reported "interrupted" whatever its age, which asserted a cause nobody measured: with two hosts each launching their own server, the other server may be alive, and "call start_audit" would put a second pipeline in its folder. The job now refreshes its record every 15 s; fresh means live in another process — read tools and start_audit refuse, audit_status reports "running" and waits on the record — and only a record unrefreshed for 60 s is "interrupted". - The disk record now decides, not this server's memory of its jobs: a case completed since by other means reads again. A failed audit is superseded by a results.json written after it, and an audit run without the scout by a scout.json written after it — the record's mtime is the measurement. - The evidence caption for a named source carried the headline source's anchor state beside its own; it now carries only the shown source's. - list_gaps reported unreached_labels [] for a label audit that never ran (no clean.md, occurrences counted); it is null unless labels were read. - audit_status's manuscript was a full path from this server and a file name from a record; it is the absolute path in both, and the record stores that. - A recorded failure's next step sent the host to read a log that was not kept; recorded states have their own wording. - The record is validated — schema version, state, scout as a bool — before anything is decided on it; "false" is not False. - `papertrace mcp` re-raises an import failure when a 2.x SDK is installed, instead of an install line contradicting itself over a hidden traceback. - The published schema describes record-derived states; a test validates every state a record can produce, and another holds the prose timings to the constants. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
The spec's eleventh rule now says how "running" is measured and that the record on disk, not a server's memory, decides. Its omissions lose "two server processes auditing one case are not detected", which stopped being true, and gain the heartbeat's real limit instead: a machine that slept past the 60 s window reads as stopped — its case refused, the safe side — while a second audit started inside that window is not. README, CHANGELOG and CLAUDE.md say the same; the plan records the pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
The README named four hosts and installed from a clone. It now lists the apps that can start a local MCP server (Claude Desktop, Claude Code, Cursor, VS Code, Windsurf, Gemini CLI, Codex CLI) with where each takes its configuration, and says plainly that ChatGPT and Claude in a browser cannot: they reach MCP servers over the internet only. Whichever app drives the tools, every claim is judged through `claude -p`, so Claude Code has to be installed and signed in — said once, beside the list, rather than left for `start_audit`'s refusal to reveal. Install is `uv tool install` from the GitHub archive: no Git, and uv fetches a Python. Verified in a sandbox: the install, the executable launched by absolute path with a minimal environment (nine tools; start_audit refused while PATH did not reach `claude`, and refused on the email next once it did — nothing spent), and Claude Code's `claude mcp list` reporting it connected. That last is what "tested" now claims; the other apps' configurations follow their docs and have not each been run. Also: `claude mcp add` takes `--env` as variadic, so a name straight after it is read as another KEY=value (reproduced); PowerShell's `where` is Where-Object, so `where.exe`; a Windows path in JSON doubles each `\`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
papertrace mcpserves PaperTrace over stdio to any app that can start a local MCP server: Claude Desktop, Claude Code, Cursor, VS Code (GitHub Copilot in Agent mode), Windsurf, Gemini CLI and Codex CLI. ChatGPT and Claude in a web browser cannot use it: they reach MCP servers over the internet only. This is the roadmap item drive PaperTrace as a tool from any MCP-capable client.Whichever app drives the tools, every claim is judged with
claude -p, so Claude Code must be installed and signed in on the same computer (a paid Claude plan). Reading an audit that already exists needs no Claude Code.Tools
start_auditpapertrace runpipeline on a background thread and returns at onceaudit_statusnot_started,running,finished,failedorinterrupted, with the log tail; waits up to 50 s per callaudit_summarylist_claimsget_claimget_evidencelist_referenceslist_gaps,get_scoutHow the rules in
CLAUDE.mdcarry overrun_disclosures/claim_disclosures/judgement_disclosuresthe four report formats use, through the viewer'sreport._disclosure_dict. The disclosure parity test now covers this fifth reader.results.jsonis refused, never answered with empty counts. The manifest's three-state fields pass through, sonullstaysnull. A coverage audit that never ran isnull.ToolErrorcarrying its reason. In SDK v2, any other exception reaches the host as a bare "Error executing tool".sys.stdinis an empty stream, so a disputed reference label is withheld (numbering_chosen_by: "default"), never settled. MCP elicitation is deliberately not offered in its place: the gate is a person, and a host may be a model.ask.pyis still the only file insrc/that starts a process. The audit runs in-process on one thread, one audit per server. Each audit starts withask.forget_models(), so an audit that judged nothing cannot report the previous audit's model..pnginside<case>/out/, and nopdf_pathleaves the server.<case>/mcp_audit.json:running, refreshed every 15 s, then the outcome. The read tools refuse a case whose last MCP audit is running, failed or never finished, because its files may mix two runs. A record not refreshed for 60 s isinterrupted. Aresults.jsonwritten after a failed audit supersedes the failure, and ascout.jsonwritten after an audit that ran without the scout supersedes that.Changes
src/papertrace/mcp_server.py(new): the nine tools, the job registry andserve().src/papertrace/cli.py:runis split into the keyword-only_run_pipelineand its Typer adapter, with no behaviour change. It gains thepapertrace mcpcommand, which exits 2 with the install line on stderr when the SDK is missing or 1.x.SCOUT_CAVEATis now one shared constant.src/papertrace/ask.py:forget_models().src/papertrace/disclosures.py:coverage_audited(), extracted unchanged fromrun_disclosuresso the reports and the server share one rule.pyproject.toml: a new[mcp]extra (mcp>=2.2,<3), also carried by[full]and by[dev], so CI runs the MCP tests instead of skipping them.schemas/mcp_tools.schema.jsonandschemas/mcp_audit.schema.json(new).tests/test_mcp_server.py(new); additions totests/test_ask.py,tests/test_packaging.pyandtests/test_pipeline_split.py.uv tool installfrom the GitHub archive), each app's configuration and the Windows notes.CHANGELOG.md(Unreleased) andCLAUDE.md.docs/superpowers/specs/2026-09-26-mcp-server-design.md,docs/superpowers/plans/2026-09-26-mcp-server.md,docs/adr/0004-mcp-server.md.Gates in
CLAUDE.mdClient, with the pipeline faked atcli._run_pipeline. One real stdio round trip launches the server itself; it makes no model calls.mcp_audit.jsonand every tool output have schemas. Outputs and every record state are validated against them, and each schema definition's keys are held equal to itsTypedDict's.refs.py,check.py,scout.py: untouched.report.pyis not in the diff, anddisclosures.pyonly gains the extracted function. So the end-to-end demo was not re-run.Verification
ruff check src tests scripts evalsis clean.tests/test_refs.py::test_the_contact_email_reaches_only_the_services_that_need_it. It fails behind the development sandbox's HTTP proxy (httpx.ProxyError: 403) onmainas well, and passes in CI.uv tool install --python 3.12 "papertrace[mcp] @ <this branch's GitHub archive>"installed it, andpapertrace --versionprintedpapertrace 0.7.1.audit_statusanswerednot_started;start_auditwas refused whilePATHdid not reachclaude;PATHfixed and no email configured, it was refused on the email instead. Nothing was spent and no case folder was created.claude mcp add --env PAPERTRACE_EMAIL=… --transport stdio --scope user papertrace -- <path> mcp, thenclaude mcp listreported the server Connected. A name placed straight after--envis rejected ("Invalid environment variable format"), which the README warns about.uv tool install --reinstall-package papertrace …re-fetched only PaperTrace, in 11 s.Tried by the author
After this description was first written, the author installed PaperTrace from GitHub on their own computer, set it up in Claude Code and Claude Desktop, and reported that it worked.
Not verified
papertrace runruns; the job around it is tested with the pipeline faked.Deliberate omissions
--pngis not exposed. It needs playwright and a browser, and the images the server returns are the evidence crops, which need neither.mcp_audit.json.Review
The branch had two review passes: 20 findings, 19 fixed test-first and 1 rejected with evidence. The plan records them in Tasks 7 and 8.
🤖 Generated with Claude Code
https://claude.ai/code/session_01J6JtZg47dHEKHDEfAJ5Kpf