Skip to content

Repository files navigation

Tri-Agent Testing

Tri-Agent Testing

AI writes the code. AI writes the tests. Who tests the tests?
A separation-of-powers acceptance-testing protocol in which three isolated agents set the exam, take the exam, and judge the exam. Nobody grades their own paper.

MIT License Python 3.10+ Platform selftest 21/21 简体中文


When an AI system is tested by the same AI (or the same session) that built it, the test is already lost: the fixer avoids its own known blind spots, the tester who has seen the expected values "discovers" exactly those values, and a green report means nothing. This skill was distilled from a real production fight where exactly that happened: 16 designed chaos cases were never executed, rounds 2-5 were run by the fixer itself, and only the first round ever met the fresh-eyes standard.

Tri-Agent Testing answers with information isolation. Every role gets exactly one piece of the truth: the examiner writes the exam and computes the answer key independently, the test executor sees neither the implementation nor the answers, the judge reads the answers only when grading, and only the builder may modify code. Nothing is trusted that cannot be re-run by a third party.

The visibility contract

Role Sees implementation? Sees answers? Cross-round memory? Write access
Examiner forbidden computes it independently kept (consistent grading lens) tests/<v>/ only, minus judge/ and tickets/
Builder (main agent) yes never kept sole writer (fixes from bug tickets)
Tester forbidden forbidden fresh spawn every round judge/ only
Judge yes on grading only kept, must cite mechanical evidence config only; files bug tickets

Five iron rules

  1. Separation of powers. The fixer never acts as test executor. Anything the fixer runs is dev self-test, not acceptance evidence.
  2. Sealed answers. The answer key is computed by an independent implementation (pure Decimal, never through the engine under test), then sealed with a SHA-256 manifest. Sealing rejects CRLF, verify catches tampering, re-sealing must declare a new hash.
  3. Mechanical grading. Numbers are judged by a script, never by an LLM (LLM judges drift and favor; this is the CAR-bench finding). The judge-LLM only rules on behavior categories, test quality, and leaks, and must cite re-runnable evidence.
  4. Fresh eyes every round. The test agent is spawned with no memory of previous rounds. A tester with memory avoids known mines and goes soft on the fixer.
  5. Cases must actually run, and the harness must have seen red. Designed-but-never-executed cases are debt. On round 1 the harness must be calibrated against at least 3 injected mutations (deliberate known bugs); any surviving mutation is a blind spot.

The pipeline

flowchart TD
    START(["hard acceptance requested"]) --> SURVEY["startup survey: scope+spec / depth / gate N / object type<br/>cost note inline; answer = consent"]
    SURVEY --> P1["gate checks: spec readable, baseline double-run byte-identical, git ready"]
    P1 --> Q["Examiner: SPEC + deterministic generator + sealed answer + caliber + TESTPLAN (no answers)"]
    Q --> ASM["Builder assembles, never reading the sealed dir"]
    ASM --> T["Tester (fresh spawn): runs every case in copy environments,<br/>collects actual + provenance + behaviors + metamorphic pairs"]
    T --> J["Judge: three-channel scoring + leak scan + smell list +<br/>mutation kill-rate + fingerprint re-runs + bug tickets"]
    J --> GATE{"double-run stable, 100%, zero blockers, kill 100%?"}
    GATE -- flaky --> FIX["Builder fixes from tickets; fuse: 3 failures = escalate"]
    GATE -- any fail --> ATTR{"attribute: engine / exam / process"}
    ATTR --> ZERO["gate counter resets"] --> DEEP["Examiner deepens: new chaos, new bounds,<br/>mutation menu rotates 30%/50%"]
    DEEP --> ASM
    GATE -- pass --> NG{"N-1 heterogeneous rounds done?"}
    NG -- no --> DEEP
    NG -- yes --> DONE(["final round: baseline regression + all past reds re-run"])
Loading

Grading is three-channel, and the harness itself is tested

Three channels, one script, frozen tolerances. scripts/score.py grades sealed answers (amount 0.01 / ratio 1e-6 / count and status exact), behavior assertions (enum equality, including mutation verdicts), and metamorphic relations (slice sums equal totals, row order changes nothing, idempotent replay is stable). It fails fast on malformed input, treats a missing field as failure rather than a pass, detects duplicate keys instead of silently merging them, is byte-identical across re-runs, and echoes the answer and caliber hashes into its own output so every grade is auditable.

Mutation testing for the exam itself (the Stryker/mutant school). The judge injects deliberate known bugs into throwaway copies and requires the harness to catch them. A surviving mutation is a blind spot in your tests, and the kill rate is a gate criterion.

Metamorphic relations as a third evidence chain. MRs are algebraic identities derived from the spec that need no expected values at all: no leak (the builder can run them), no precision mirroring problem, and a third independent opinion. When the sealed answer and the engine agree but an MR fails, your caliber itself is wrong.

three-channel grading

One power structure, four kinds of target

The protocol does not change with the target; only the artifacts do (oracle shape, tolerance classes, chaos and mutation menus, baseline restore). The preset table covers:

Data warehouse / reports SaaS / API Document pipeline GUI
Sealed oracle answer rows one row per request case golden files + hashes golden DOM/state tree
Tolerance classes amount / ratio / count status / text / json / set text / json / set text / set / status
Chaos menu BOM, headers, CRLF, dates authz, token expiry, replay, pagination bad formats, macros, huge files invalid submits, expired sessions
Baseline restore file copies DB snapshot + restart file copies session reset + snapshot
Relative cost 1x 1.5-2x 1x 5-10x

For GUI targets the rule is: trigger through the GUI, assert through machine-readable channels (DB, API, files, logs) wherever possible; screenshots are semi-mechanical and never a gate criterion.

Quick start

Copy the skill into your user skills directory:

git clone https://github.com/August06exe/tri-agent-testing.git
mkdir -p ~/.zcode/skills && cp -r tri-agent-testing ~/.zcode/skills/

Then, in a session, say what you want in your own words. These all trigger it:

  • "Run a hard acceptance test on this system, do not fool around"
  • "Three-branch testing with separate exam-setter and grader"
  • "Three consecutive all-green rounds before we call this done"

A startup survey follows (scope and spec path, depth, gate N, target type). Answering it is consent; the cost note is inline in the questions.

Run the self-test any time:

bash ~/.zcode/skills/tri-agent-testing/selftest/run_selftest.sh
# 21 checks: seal/verify/tamper, three-channel pass+fail, idempotency,
# fail-fast, leak hits, CRLF rejection, duplicate keys, invalid rows, column order

Worked example

examples/toy-grades/ is a complete, real run against a tiny grade-summary CLI with two deliberately planted bugs (banker's rounding where HALF_UP was promised; group keys not stripped). Round 1 ended 19/22 with a 75% mutation kill rate, seven bug tickets, and a judge finding that the scoring script itself cannot see tolerance-sized rounding drift (the audit channel caught it). The examiner deepened the exam, the builder fixed three engine bugs, and the gate was re-run with fresh testers every round. Read tests/v1/SCOREBOARD.md for the full trail.

Lessons baked in

Every rule here was paid for in production: precision mirroring failures that produced 44 false failures across 1,135 rows; a config that promised validation the pipeline skipped (caught via a side-effect fingerprint); a missed rebuild step that let stale artifacts fail grading; an overwrite-happy scorer destroying round evidence; a judge going soft with memory; and an "actual" file that was itself an answer mirror. The full trap list lives in SKILL.md, and each trap seeds a regression case or mutation so every lesson is paid once.

Prior art this stands on

  • Inspect AI (UK AISI): task/solver/scorer separation
  • CAR-bench (UC Berkeley RDI): LLM judges are not judges
  • Stryker / mutant / SWE-smith: mutation testing and synthetic bugs
  • superpowers: watch-it-fail discipline and the 3-strike escalation fuse
  • Hypothesis and metamorphic testing: oracle-free third evidence chains

License

MIT

About

Three-agent separation-of-powers acceptance testing for AI-built systems: sealed answers, mechanical grading, mutation-tested harness, fresh-eyes testers, hard N-round gate

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages