Skip to content

Repository files navigation

polywatch

A background bug hunter for the code Claude writes.

After each Claude Code turn that edits files, a cheap reviewer from another model family (DeepSeek V4.1-Flash by default) reads the change in the background. Claude Opus 5.5 then checks the most serious claims one at a time against the code and throws out the ones that do not hold. On your next prompt you see at most two findings, ranked; confirmed defects go back to Claude to fix. You tell polywatch which findings were real, and that sets the ranking.

polywatch: 2 finding(s) worth a look (2 confirmed) · HARD · 2 weaker hidden · $0.0205 · id mufo4bu5-manual
  1. [confirmed, high] vote-bug.js handleVoteRequest: the up-to-date check compares only lastIndex and ignores the candidate's last log term
  2. [confirmed, high] vote-bug.js onMessage: a higher-term heartbeat never updates node.term, node.role or node.votedFor

When nothing survives: polywatch: nothing worth your attention.

What it is and is not

polywatch is not a gate. Replayed on 50 commits of a real project, the reviewer rejected 24 of 25 commits that were later fixed and 25 of 25 that were not; a trust score built on its verdict separated them no better than chance (AUC 0.48). So the verdict is ignored. What the reviewer does produce is specific claims, and some of them are real bugs: in 7 of the 25 later-fixed commits, the bug the fix repaired was already on its list. polywatch's job is to get those few in front of you and keep the rest out of the way. See BaanBaan replay.

How it works

Step What happens Cost
Record PostToolUse hook notes every Edit, Write and MultiEdit; at Stop, source files changed any other way during the turn (Bash heredocs, sed -i, scripts) are added none
Start Stop hook snapshots the turn and starts a background worker; Claude is never blocked (with "deliver": "stop" it reviews before Claude stops instead, and hands confirmed defects straight back) none
Route a deterministic router labels the change EASY, MEDIUM or HARD (protocol logic, concurrency, state) none
Review DeepSeek V4.1-Flash lists specific, checkable issues (one retry on an empty answer) about $0.002 to $0.02
Check for HARD changes, testCommand runs if configured CPU only
Confirm Opus 5.5 checks the most severe high and medium claims, up to maxClaims (default 2), against the cited code; refuted claims are dropped about $0.02 to $0.05 per claim
Rank each finding is scored by how often findings of its kind (confirmed or not, by severity) turned out real none
Deliver UserPromptSubmit shows the top 2 above the threshold, plus any notes (failed tests, reviewer errors); only confirmed findings and failed tests go to Claude, at most 2 rounds in a row none
Learn reviews and your outcomes go to .polywatch/ledger.jsonl none

The adjudicator sees the cited file whole when it fits, otherwise windows around the identifiers the claim names, plus matching excerpts from the other files changed in the turn. Sending the top of a long file instead left 23 of 34 claims "uncertain" in the replay; the targeted excerpts cut that to 27 of 86.

Ranking

Findings fall into six kinds: confirmed or unconfirmed, times high, medium or low severity. Each kind starts from a prior precision (confirmed high 60%, confirmed medium 50%, unconfirmed high 30%, unconfirmed medium 20%, and so on) worth four outcomes, and every real or false you record for that kind moves it. Findings below minScore (0.25) are hidden, so unconfirmed medium and low claims stay out of sight until your outcomes say they are worth showing. polywatch calibration prints the current table.

Architecture

The diagrams are generated by archlens from polywatch.analysis.json. Edit the analysis, then render again. Each diagram answers one question:

The same content as text, with every component and relation: docs/architecture/README.md.

To render again:

node <archlens>/skills/archlens/bin/archlens.mjs validate polywatch.analysis.json
node <archlens>/skills/archlens/bin/archlens.mjs render polywatch.analysis.json docs/architecture --repo-root .
# the rendered pages are published from the gh-pages branch (architecture/), not kept in the plugin

Install

# API keys: set them in the plugin's settings (/plugin → polywatch → Configure). Claude Code keeps them in your
# system credential store. Or, for the CLI and scripts, in your environment (not in the repo). Use your own variable name for the Anthropic key:
# if ANTHROPIC_API_KEY is set, Claude Code itself uses it and bills your API account instead of your plan.
setx POLYWATCH_DEEPSEEK_API_KEY "sk-..."        # reviewer (export ... on macOS/Linux)
setx POLYWATCH_ANTHROPIC_API_KEY "sk-ant-..."   # confirmation (optional; without it nothing is confirmed)

# ~/.polywatch.json: where the keys are, and (optionally) the folders polywatch may act in
# { "providers": { "deepseek":  { "apiKeyEnv": "POLYWATCH_DEEPSEEK_API_KEY" },
#                  "anthropic": { "apiKeyEnv": "POLYWATCH_ANTHROPIC_API_KEY" } },
#   "onlyUnder": ["C:\\Users\\me\\code"] }

# install
claude plugin marketplace add cognitive-fab/polywatch
claude plugin install polywatch@polywatch

# or try a checkout for one session
git clone https://github.com/cognitive-fab/polywatch && claude --plugin-dir ./polywatch/plugin

Configure (.polywatch.json in the project root, or ~/.polywatch.json for every project; all optional)

{
  "testCommand": "npm test",          // ~/.polywatch.json or a trusted project only, see below
  "budgetUsdPerTurn": 0.5,            // per turn; a project file can lower it, only yours can raise it
  "budgetUsdPerDay": 5,               // money per day across all projects; reviews stop for the day when reached
  "confirm": "top",
  "maxFindings": 2,                   // what you see per turn
  "maxToClaude": 5,                   // confirmed findings sent to Claude per turn
  "minScore": 0.25,
  "reviewer": { "provider": "deepseek", "model": "deepseek-flash" },
  "adjudicator": { "provider": "anthropic", "model": "claude-opus-5-5", "maxClaims": 2, "includeRequest": "short" },
  "exclude": [".env", "*.pem", "secrets/**"]
}
  • confirm: which claims Opus checks, most severe first up to maxClaims: top (high and medium), high, all or none. high roughly halves the confirmation cost.
  • includeRequest: whether Opus sees the request the code was written for when checking a claim: short (requests up to 20,000 characters), always or never. A claim whose fix the request rules out is refuted. On etcd specs this removed most false alarms (see below); on long requests it is the largest cost.
  • exclude adds patterns to the built-in ones (.env, *.pem, *.key, *secret*, *credential*, …); it cannot remove them.
  • The adjudicator can also run on DeepSeek ("provider": "deepseek", "model": "deepseek-v4-pro"). Each provider reads its own key (DEEPSEEK_API_KEY, ANTHROPIC_API_KEY).
  • Settings that decide where your keys go or what runs on your machine (testCommand, and baseUrl, apiKeyEnv and price under reviewer or adjudicator) are ignored in a project's .polywatch.json, because a cloned repository could set them. Put them in ~/.polywatch.json (same format, applies to every project), or list the project there so its own file may set them: { "trustedProjects": ["C:\\Users\\me\\code\\myapp"] }. Ignored settings and invalid config files are reported when a session starts.
  • User file only: "onlyUnder": ["C:\\Users\\me\\code"] limits polywatch to projects inside those folders (elsewhere the hooks do nothing), and "providers": { "anthropic": { "apiKeyEnv": "POLYWATCH_ANTHROPIC_API_KEY" } } changes which variable a provider's key is read from. Use that for Anthropic: if ANTHROPIC_API_KEY is set globally, Claude Code itself uses it and bills your API account instead of your subscription.
  • "provider": "claude-code" (reviewer or adjudicator) runs the call through claude -p on your Claude plan instead of the API, so no Anthropic key is needed: "adjudicator": { "provider": "claude-code", "model": "claude-opus-5-5" }. The call runs with --safe-mode (no hooks, plugins or CLAUDE.md), --tools "" (the model can only answer) and a one-line system prompt, from a temp folder, and without ANTHROPIC_API_KEY in its environment. It adds a few seconds of CLI start-up per claim; the cost shown is the list price of the tokens, which on a plan counts toward usage limits rather than money.
  • price: [input, output] in USD per million tokens, for a model polywatch has no price for (a local server, a new model). Without it, an unknown model is charged at the highest known price so the per-turn budget still holds, and the review says so.
  • .polywatch/ gets its own .gitignore, so prompts, code and review results stay out of commits.
  • Anthropic spend limits: a workspace can have its own monthly limit below the organization's. If confirmation reports "workspace API usage limits", use a key from another workspace or raise that workspace's limit.
  • Privacy: changed files are sent to the reviewer's API (DeepSeek by default). Files matching exclude are never sent. For code that must not leave your control, set reviewer.provider to anthropic or point baseUrl (in ~/.polywatch.json) at a local OpenAI-compatible server. Full policy: PRIVACY.md.

Commands

polywatch report                             # recent reviews with ranked findings and evidence (the /trust skill runs this)
polywatch review <files> --task "…"          # review files by hand
polywatch outcome <id> <n> real|false [--by claude] [--dir <project>] [note]   # was finding n real? sets the weights; Claude records its own verdicts with --by claude
polywatch calibration                        # current precision per kind of finding
polywatch stats [dir]                        # cost, findings, what reached Claude, outcomes, and state problems for a project
polywatch dashboard [dir] [--no-open]        # one-file HTML dashboard: running now, waiting on you (with the default), latest reviews, stuck

BaanBaan replay

scripts/replay-git.mjs replays a repository's history. Ground truth comes from SZZ: a commit is buggy if a later fix: commit changed its lines, clean if no fix ever touched them. On BaanBaan (25 buggy, 25 clean commits), with the same Flash reviews for both pipelines:

v0.1 (trust score, every issue shown) v0.2 (ranked, top 2)
items per commit the user reads 3.2 issues plus a score 1.3 findings
later-fixed bug that Flash found, shown to the user 7 of 7, buried among 3.2 per commit 4 of 7, in the top 2
... confirmed by Opus and sent to Claude 0 of 7 4 of 7
commits with a confirmed finding (buggy / clean) n/a 18 / 17 of 25
cost per commit $0.05 ($0.016 review, $0.033 Opus) $0.10 ($0.016 review, $0.08 Opus)

Two limits on this. Clean commits get confirmed findings as often as buggy ones; "clean" only means no fix commit touched the lines, so some of those findings may be real bugs nobody has fixed, but without labels the precision of confirmed findings is not measured. That is what outcome is for. And the three missed bugs were medium-severity claims ranked below the top two or never checked.

sysmobench

scripts/replay-sysmo.mjs runs polywatch on the sysmobench specifications of the price-of-trust study, where a model checker decides correctness, and on bugs planted in verified specs.

spinlock and lock service etcd Raft
planted bugs the checker catches: confirmed and ranked first 90 of 91 14 of 15
planted bugs the checker misses: confirmed 67 of 67 none planted
verified specs with a confirmed finding 1 of 158 15 of 15; 4 of 15 with includeRequest
checker-failed specs with a confirmed finding none tested 19 of 20; 6 of 20 with includeRequest

Dashboard

polywatch dashboard writes .polywatch/dashboard.html and opens it. It needs no server: it opens with a double-click and reloads itself every 10 seconds, and polywatch rewrites it when a review starts or ends, when results are delivered and when an outcome is recorded. Panels: what is running now, what waits on you and what happens if you do nothing, the latest reviews, anything stuck, and the last 14 days. Style in ~/.polywatch.json: "dashboard": { "theme": "light", "density": "compact", "accent": "#1D56C9" } (theme also takes dark or auto, density takes airy). The layout follows an idea by @voxyz_ai: a dashboard beside every long task.

Claude Code with and without polywatch

bench/polyglot.mjs runs Claude Code headless on exercises from Aider's polyglot benchmark (hidden unit tests), in three arms: plain, selfcheck (the session is resumed once with a fixed "check your work" prompt, the control for a second pass) and polywatch (loaded with "deliver": "stop"). All arms run with --setting-sources project, so none of the user's own hooks or plugins run. bench/report.mjs <out> gives pass rate, time and tokens to success, cost per pass and paired McNemar tests.

Tests

npm test (the hook tests run the real CLI and background worker; one uses a local fake of the model API).

Limitations

  • The reviewer sees only the files changed in the turn and judges against the request text. Without a stated contract it raises questions about intent as defects (examples/lock-fixed.js); Opus usually refutes them, not always.
  • Confirmation is the main cost. At 30 edit turns a day on code like BaanBaan's, expect about three dollars; set "confirm": "high" or "maxClaims": 1 to reduce it.
  • Results arrive on the next prompt, not immediately after the turn.
  • The priors are starting guesses. The ranking means little until the ledger holds a few dozen outcomes.

Glossary

  • Turn: one round of Claude Code work, from the user's prompt until Claude stops.
  • Hook: a command Claude Code runs at a fixed event, such as after a file edit or when a turn ends.
  • Tier: EASY, MEDIUM or HARD. How hard a change is to get right, from its size and the kind of code it touches.
  • Claim: one specific, checkable statement a reviewer makes about a defect: which file, where, and what is wrong.
  • Finding: a claim that was not refuted, with a score.
  • Confirmed: Opus checked the claim against the code and said it holds.
  • Refuted: Opus checked the claim and said it does not hold. Refuted claims are never shown.
  • Precision: the share of findings of one kind that turned out to be real bugs.
  • Outcome: your verdict on one finding after the fact: real bug or false alarm.
  • Ledger: .polywatch/ledger.jsonl. Every review and every outcome, in order.
  • SZZ: a way to find the commit that introduced a bug: blame the lines a later fix commit changed.
  • Mutant: a copy of a correct specification with one bug planted on purpose, so the right answer is known.
  • sysmobench: a benchmark of executable specifications of real systems, checked against traces of the real implementation.
  • AUC: the chance that a score ranks a random buggy item above a random clean one. 0.5 is no better than chance.

About

Background bug hunting for code Claude Code writes: a cheap reviewer from another model family, Opus confirmation, only confirmed findings go back to Claude.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages