A background bug hunter for the code Claude writes.
After each Claude Code turn that edits files, a cheap reviewer from another model family (DeepSeek V4.1-Flash by default) reads the change in the background. Claude Opus 5.5 then checks the most serious claims one at a time against the code and throws out the ones that do not hold. On your next prompt you see at most two findings, ranked; confirmed defects go back to Claude to fix. You tell polywatch which findings were real, and that sets the ranking.
polywatch: 2 finding(s) worth a look (2 confirmed) · HARD · 2 weaker hidden · $0.0205 · id mufo4bu5-manual
1. [confirmed, high] vote-bug.js handleVoteRequest: the up-to-date check compares only lastIndex and ignores the candidate's last log term
2. [confirmed, high] vote-bug.js onMessage: a higher-term heartbeat never updates node.term, node.role or node.votedFor
When nothing survives: polywatch: nothing worth your attention.
polywatch is not a gate. Replayed on 50 commits of a real project, the reviewer rejected 24 of 25 commits that were later fixed and 25 of 25 that were not; a trust score built on its verdict separated them no better than chance (AUC 0.48). So the verdict is ignored. What the reviewer does produce is specific claims, and some of them are real bugs: in 7 of the 25 later-fixed commits, the bug the fix repaired was already on its list. polywatch's job is to get those few in front of you and keep the rest out of the way. See BaanBaan replay.
| Step | What happens | Cost |
|---|---|---|
| Record | PostToolUse hook notes every Edit, Write and MultiEdit; at Stop, source files changed any other way during the turn (Bash heredocs, sed -i, scripts) are added |
none |
| Start | Stop hook snapshots the turn and starts a background worker; Claude is never blocked (with "deliver": "stop" it reviews before Claude stops instead, and hands confirmed defects straight back) |
none |
| Route | a deterministic router labels the change EASY, MEDIUM or HARD (protocol logic, concurrency, state) | none |
| Review | DeepSeek V4.1-Flash lists specific, checkable issues (one retry on an empty answer) | about $0.002 to $0.02 |
| Check | for HARD changes, testCommand runs if configured |
CPU only |
| Confirm | Opus 5.5 checks the most severe high and medium claims, up to maxClaims (default 2), against the cited code; refuted claims are dropped |
about $0.02 to $0.05 per claim |
| Rank | each finding is scored by how often findings of its kind (confirmed or not, by severity) turned out real | none |
| Deliver | UserPromptSubmit shows the top 2 above the threshold, plus any notes (failed tests, reviewer errors); only confirmed findings and failed tests go to Claude, at most 2 rounds in a row |
none |
| Learn | reviews and your outcomes go to .polywatch/ledger.jsonl |
none |
The adjudicator sees the cited file whole when it fits, otherwise windows around the identifiers the claim names, plus matching excerpts from the other files changed in the turn. Sending the top of a long file instead left 23 of 34 claims "uncertain" in the replay; the targeted excerpts cut that to 27 of 86.
Findings fall into six kinds: confirmed or unconfirmed, times high, medium or low severity. Each kind starts from a prior precision (confirmed high 60%, confirmed medium 50%, unconfirmed high 30%, unconfirmed medium 20%, and so on) worth four outcomes, and every real or false you record for that kind moves it. Findings below minScore (0.25) are hidden, so unconfirmed medium and low claims stay out of sight until your outcomes say they are worth showing. polywatch calibration prints the current table.
The diagrams are generated by archlens from polywatch.analysis.json. Edit the analysis, then render again. Each diagram answers one question:
- What happens after Claude edits code: the timeline from an edit to the next prompt. Claude never waits for a review.
- What happens to a claim: review, confirm, rank, show.
- What leaves your machine: the two outbound calls, and what each one carries.
- How your outcomes change the ranking: from "that was a false alarm" to a lower score next time.
- Measuring on real project history: the git replay.
- Measuring on planted bugs: the sysmobench replay.
The same content as text, with every component and relation: docs/architecture/README.md.
To render again:
node <archlens>/skills/archlens/bin/archlens.mjs validate polywatch.analysis.json
node <archlens>/skills/archlens/bin/archlens.mjs render polywatch.analysis.json docs/architecture --repo-root .
# the rendered pages are published from the gh-pages branch (architecture/), not kept in the plugin# API keys: set them in the plugin's settings (/plugin → polywatch → Configure). Claude Code keeps them in your
# system credential store. Or, for the CLI and scripts, in your environment (not in the repo). Use your own variable name for the Anthropic key:
# if ANTHROPIC_API_KEY is set, Claude Code itself uses it and bills your API account instead of your plan.
setx POLYWATCH_DEEPSEEK_API_KEY "sk-..." # reviewer (export ... on macOS/Linux)
setx POLYWATCH_ANTHROPIC_API_KEY "sk-ant-..." # confirmation (optional; without it nothing is confirmed)
# ~/.polywatch.json: where the keys are, and (optionally) the folders polywatch may act in
# { "providers": { "deepseek": { "apiKeyEnv": "POLYWATCH_DEEPSEEK_API_KEY" },
# "anthropic": { "apiKeyEnv": "POLYWATCH_ANTHROPIC_API_KEY" } },
# "onlyUnder": ["C:\\Users\\me\\code"] }
# install
claude plugin marketplace add cognitive-fab/polywatch
claude plugin install polywatch@polywatch
# or try a checkout for one session
git clone https://github.com/cognitive-fab/polywatch && claude --plugin-dir ./polywatch/pluginConfigure (.polywatch.json in the project root, or ~/.polywatch.json for every project; all optional)
confirm: which claims Opus checks, most severe first up tomaxClaims:top(high and medium),high,allornone.highroughly halves the confirmation cost.includeRequest: whether Opus sees the request the code was written for when checking a claim:short(requests up to 20,000 characters),alwaysornever. A claim whose fix the request rules out is refuted. On etcd specs this removed most false alarms (see below); on long requests it is the largest cost.excludeadds patterns to the built-in ones (.env,*.pem,*.key,*secret*,*credential*, …); it cannot remove them.- The adjudicator can also run on DeepSeek (
"provider": "deepseek", "model": "deepseek-v4-pro"). Each provider reads its own key (DEEPSEEK_API_KEY,ANTHROPIC_API_KEY). - Settings that decide where your keys go or what runs on your machine (
testCommand, andbaseUrl,apiKeyEnvandpriceunderrevieweroradjudicator) are ignored in a project's.polywatch.json, because a cloned repository could set them. Put them in~/.polywatch.json(same format, applies to every project), or list the project there so its own file may set them:{ "trustedProjects": ["C:\\Users\\me\\code\\myapp"] }. Ignored settings and invalid config files are reported when a session starts. - User file only:
"onlyUnder": ["C:\\Users\\me\\code"]limits polywatch to projects inside those folders (elsewhere the hooks do nothing), and"providers": { "anthropic": { "apiKeyEnv": "POLYWATCH_ANTHROPIC_API_KEY" } }changes which variable a provider's key is read from. Use that for Anthropic: ifANTHROPIC_API_KEYis set globally, Claude Code itself uses it and bills your API account instead of your subscription. "provider": "claude-code"(reviewer or adjudicator) runs the call throughclaude -pon your Claude plan instead of the API, so no Anthropic key is needed:"adjudicator": { "provider": "claude-code", "model": "claude-opus-5-5" }. The call runs with--safe-mode(no hooks, plugins or CLAUDE.md),--tools ""(the model can only answer) and a one-line system prompt, from a temp folder, and withoutANTHROPIC_API_KEYin its environment. It adds a few seconds of CLI start-up per claim; the cost shown is the list price of the tokens, which on a plan counts toward usage limits rather than money.price:[input, output]in USD per million tokens, for a model polywatch has no price for (a local server, a new model). Without it, an unknown model is charged at the highest known price so the per-turn budget still holds, and the review says so..polywatch/gets its own.gitignore, so prompts, code and review results stay out of commits.- Anthropic spend limits: a workspace can have its own monthly limit below the organization's. If confirmation reports "workspace API usage limits", use a key from another workspace or raise that workspace's limit.
- Privacy: changed files are sent to the reviewer's API (DeepSeek by default). Files matching
excludeare never sent. For code that must not leave your control, setreviewer.providertoanthropicor pointbaseUrl(in~/.polywatch.json) at a local OpenAI-compatible server. Full policy: PRIVACY.md.
polywatch report # recent reviews with ranked findings and evidence (the /trust skill runs this)
polywatch review <files> --task "…" # review files by hand
polywatch outcome <id> <n> real|false [--by claude] [--dir <project>] [note] # was finding n real? sets the weights; Claude records its own verdicts with --by claude
polywatch calibration # current precision per kind of finding
polywatch stats [dir] # cost, findings, what reached Claude, outcomes, and state problems for a project
polywatch dashboard [dir] [--no-open] # one-file HTML dashboard: running now, waiting on you (with the default), latest reviews, stuckscripts/replay-git.mjs replays a repository's history. Ground truth comes from SZZ: a commit is buggy if a later fix: commit changed its lines, clean if no fix ever touched them. On BaanBaan (25 buggy, 25 clean commits), with the same Flash reviews for both pipelines:
| v0.1 (trust score, every issue shown) | v0.2 (ranked, top 2) | |
|---|---|---|
| items per commit the user reads | 3.2 issues plus a score | 1.3 findings |
| later-fixed bug that Flash found, shown to the user | 7 of 7, buried among 3.2 per commit | 4 of 7, in the top 2 |
| ... confirmed by Opus and sent to Claude | 0 of 7 | 4 of 7 |
| commits with a confirmed finding (buggy / clean) | n/a | 18 / 17 of 25 |
| cost per commit | $0.05 ($0.016 review, $0.033 Opus) | $0.10 ($0.016 review, $0.08 Opus) |
Two limits on this. Clean commits get confirmed findings as often as buggy ones; "clean" only means no fix commit touched the lines, so some of those findings may be real bugs nobody has fixed, but without labels the precision of confirmed findings is not measured. That is what outcome is for. And the three missed bugs were medium-severity claims ranked below the top two or never checked.
scripts/replay-sysmo.mjs runs polywatch on the sysmobench specifications of the price-of-trust study, where a model checker decides correctness, and on bugs planted in verified specs.
| spinlock and lock service | etcd Raft | |
|---|---|---|
| planted bugs the checker catches: confirmed and ranked first | 90 of 91 | 14 of 15 |
| planted bugs the checker misses: confirmed | 67 of 67 | none planted |
| verified specs with a confirmed finding | 1 of 158 | 15 of 15; 4 of 15 with includeRequest |
| checker-failed specs with a confirmed finding | none tested | 19 of 20; 6 of 20 with includeRequest |
polywatch dashboard writes .polywatch/dashboard.html and opens it. It needs no server: it opens with a double-click and reloads itself every 10 seconds, and polywatch rewrites it when a review starts or ends, when results are delivered and when an outcome is recorded. Panels: what is running now, what waits on you and what happens if you do nothing, the latest reviews, anything stuck, and the last 14 days. Style in ~/.polywatch.json: "dashboard": { "theme": "light", "density": "compact", "accent": "#1D56C9" } (theme also takes dark or auto, density takes airy). The layout follows an idea by @voxyz_ai: a dashboard beside every long task.
bench/polyglot.mjs runs Claude Code headless on exercises from Aider's polyglot benchmark (hidden unit tests), in three arms: plain, selfcheck (the session is resumed once with a fixed "check your work" prompt, the control for a second pass) and polywatch (loaded with "deliver": "stop"). All arms run with --setting-sources project, so none of the user's own hooks or plugins run. bench/report.mjs <out> gives pass rate, time and tokens to success, cost per pass and paired McNemar tests.
npm test (the hook tests run the real CLI and background worker; one uses a local fake of the model API).
- The reviewer sees only the files changed in the turn and judges against the request text. Without a stated contract it raises questions about intent as defects (
examples/lock-fixed.js); Opus usually refutes them, not always. - Confirmation is the main cost. At 30 edit turns a day on code like BaanBaan's, expect about three dollars; set
"confirm": "high"or"maxClaims": 1to reduce it. - Results arrive on the next prompt, not immediately after the turn.
- The priors are starting guesses. The ranking means little until the ledger holds a few dozen outcomes.
- Turn: one round of Claude Code work, from the user's prompt until Claude stops.
- Hook: a command Claude Code runs at a fixed event, such as after a file edit or when a turn ends.
- Tier: EASY, MEDIUM or HARD. How hard a change is to get right, from its size and the kind of code it touches.
- Claim: one specific, checkable statement a reviewer makes about a defect: which file, where, and what is wrong.
- Finding: a claim that was not refuted, with a score.
- Confirmed: Opus checked the claim against the code and said it holds.
- Refuted: Opus checked the claim and said it does not hold. Refuted claims are never shown.
- Precision: the share of findings of one kind that turned out to be real bugs.
- Outcome: your verdict on one finding after the fact: real bug or false alarm.
- Ledger:
.polywatch/ledger.jsonl. Every review and every outcome, in order. - SZZ: a way to find the commit that introduced a bug: blame the lines a later fix commit changed.
- Mutant: a copy of a correct specification with one bug planted on purpose, so the right answer is known.
- sysmobench: a benchmark of executable specifications of real systems, checked against traces of the real implementation.
- AUC: the chance that a score ranks a random buggy item above a random clean one. 0.5 is no better than chance.
{ "testCommand": "npm test", // ~/.polywatch.json or a trusted project only, see below "budgetUsdPerTurn": 0.5, // per turn; a project file can lower it, only yours can raise it "budgetUsdPerDay": 5, // money per day across all projects; reviews stop for the day when reached "confirm": "top", "maxFindings": 2, // what you see per turn "maxToClaude": 5, // confirmed findings sent to Claude per turn "minScore": 0.25, "reviewer": { "provider": "deepseek", "model": "deepseek-flash" }, "adjudicator": { "provider": "anthropic", "model": "claude-opus-5-5", "maxClaims": 2, "includeRequest": "short" }, "exclude": [".env", "*.pem", "secrets/**"] }