Establish end-to-end time-to-answer baselines and regression gates (#335) - #346
Merged
Merged
Conversation
What: - apps/web/src/lib/performance/timeToAnswerEvidence.ts — the closed schema and math for the controlled time-to-answer benchmark: the 12 measured stages with their bucket (backend-collection / transport-notification / browser-model / rendering / search), an explicit `measures` and `doesNotProve` statement per stage, the three reference fixtures (25/100/250 containers) plus the four scenario fixtures, the fixture x stage matrix derived from the stage lists, the pinned-environment allowlist (runner/CPU/OS/kernel, Node/Rust/Docker revisions, Chromium revision and flags, fonts, production build, fixture and source revision), raw-sample validation, nearest-rank p95, median-of-three controlled runs, and the reviewed max(baseline x 1.25, baseline + 2 ms) promotion gate. - apps/web/src/lib/performance/timeToAnswerEvidence.test.ts — shape/math tests plus hostile cases (fabricated summary field, wrong baseline id, truncated matrix, 14 or 2 samples, negative and non-numeric samples, unknown stage, duplicate record, undeclared fixture, extra or unsafe metadata fields, non-chromium engine, incompatible pinned environment, over-limit candidate). Why: #335 owns measurement, fixtures, evidence format and promotion gates only. There is no performance authority outside Atlas today, so no later optimization claim in #336/#337/#338 can be compared against anything. This contract is that authority's closed shape: it stores raw samples only and derives every summary during review, so a supplied summary can never influence a result. How checked: - vitest apps/web/src/lib/performance/timeToAnswerEvidence.test.ts -> 6 passed. - The stage matrix, bucket coverage, fixture declaration and doesNotProve copy are asserted, so a stage cannot be added or re-bucketed silently. Nothing measures anything yet and no production behaviour changes: this is the schema, the fixtures' declaration and the promotion rule.
…capture (#335) What: - tests/perf/dockerFixtureTopology.mjs — one deterministic, secret-free Docker inventory generator: 25/100/250 containers plus the four scenario fixtures (provider-only revision change, Docker topology change, slow-but-bounded Compose projection, unavailable optional provider). Byte-identical for the same inputs, bounded by the published caps, and it never contacts Docker, the network, or the host filesystem. - tests/perf/fake-docker-api.mjs — a fake Docker Engine API over a unix socket serving only the three read-only inventory endpoints the collector uses (/containers/json, /networks, /volumes, plus /_ping, /version, /info) from that generator, with a benchmark-only control route to advance the published topology generation. - tests/perf/dockerFixtureTopology.test.mjs — determinism, exact counts, unique identities, contract bounds, secret-free labels, rejection of unsupported sizes/scenarios, topology-generation semantics, bounded Compose project, and the pinned fixture-name set. - docs/testing/TIME_TO_ANSWER_EVIDENCE.md — the measurement authority: stage buckets, what each number does and does not prove, the controlled-run record rules, the fixture source, the promotion rule, and an explicit statement of what is done versus still outstanding in #335. - package.json — npm run test:perf (node --test tests/perf/*.test.mjs), wired into check:js so the fixture source cannot rot silently. Why: 250/100-container fixtures cannot be invented as real containers reproducibly, and ordinary CI wall-clock time is not a gate. Pointing the daemon at a deterministic fixture socket via the existing DOCKERMAP_DOCKER_GATEWAY_SOCKET exercises the real collector, projection and publication path while keeping the input byte-identical run to run. How checked: - Verified end to end against the real daemon build: with the fixture socket serving 25 containers the daemon reports mode=docker, dockerReachable=true, and publishes 25 containers / 1 network / 5 volumes at a model revision. - node --test tests/perf/*.test.mjs -> 8 passed. - npm run check (audit, typecheck, build, contracts, version/deployment/perf tests, JS tests, Rust fmt/clippy/tests) green. No production behaviour changes: the fixture source and fixture daemon are test-only.
What:
- crates/dockermap-daemon/src/bench_timing.rs — a test-only stage timer. It is
inert unless DOCKERMAP_BENCH_STAGE_TIMING_PATH names an ABSOLUTE path; then it
appends newline-delimited JSON records ({"stage":...,"ms":...}) to that file.
A relative, empty or malformed value disables the hook. Write failures are
dropped, never propagated into a publication.
- Three attribution points, all measurements of the CURRENT implementation:
* dockerObservationMs — the Docker inventory read in
collect_docker_snapshot_candidate;
* composeEnrichmentMs — the Compose filesystem projection for the same
publication;
* findingsDerivationMs — the runtime-map findings projection in
assign_revision.
- No route, no response field, no runtime telemetry, and no behaviour change:
the Compose projection still executes inside the same Docker publication
budget as before. #336 owns moving it off that path; nothing is decoupled
here, and these numbers are what will let #336 prove the improvement.
Why:
#335 must measure the current critical path as it actually exists and show that
Compose projection currently sits inside it. Timing the two phases separately
requires a hook inside the daemon, which the issue permits only when strictly
needed to make a stage measurable.
How checked:
- 4 new unit tests: the sink requires an absolute non-empty path; stage lines are
closed newline-delimited JSON; a DISABLED hook writes nothing and creates no
file; an enabled hook appends exactly one parseable line per stage.
- cargo fmt + clippy --all-targets -D warnings clean; 207 daemon tests pass.
…n the evidence doc
What:
- tests/perf/capture.ts — `npm run perf:time-to-answer`. The single documented
command owns the whole test lifecycle: the deterministic fixture Docker
daemon, the real daemon (bench attribution on), the real API, the production
web build, the benchmark-only probe build, and real Chromium. It records raw
samples only and writes the closed artifact; summaries are recomputed during
review.
- tests/perf/browserProbe.js — test-only instrumentation loaded with
`addInitScript({ path })`: it wraps the REAL EventSource (listening for the
product's "snapshot" event) and observes DOM commits, so stages 6 and 7 have
explicit timestamps. It also exposes the measurement helpers the harness calls
by name through raw string expressions.
- tests/perf/probe/ + tests/perf/benchVite.config.mjs — a benchmark-only Vite
build that imports the REAL `buildModel` and `layoutServices` production
modules and measures them in real Chromium (stages 8 and 10). The ordinary
production build never reads this config or includes this entry.
- tests/perf/staticServer.mjs — deterministic no-store static serving with SPA
fallback and OS-reserved ports.
- tests/perf/emit-metadata.mjs — reads every pinned environment field from the
runner itself (kernel, arch, cpu count, node, rustc, docker, chromium
revision, font environment) instead of hand-authoring it.
- package.json — perf:time-to-answer and perf:metadata scripts.
Why:
#335 must measure the CURRENT implementation end to end. Stage 5 uses the real
Node/API path with its existing polling behaviour, so the baseline exposes
today's publication-observation floor rather than a synthetic watcher; the
poll interval is pinned in the environment. Lifecycle is boring on purpose:
private OS-reserved ports, process-group teardown (never pkill), explicit
readiness waits, fail-closed on any incomplete cell, raw samples preserved on
failure, and the live DockerMap deployment is never touched.
Notes:
- Stage guards come from the closed fixture x stage matrix, so a stage can only
be recorded where the contract declares it.
- provider-only and unavailable-optional-provider fixtures need no trigger:
their revision advance comes from provider state alone.
- Two esbuild/esbuild-keepNames traps are documented in the harness: a
TS-authored init script and TS-authored page.evaluate functions both emit a
`__name` helper that does not exist in the page realm, so all page code is
either a plain .js file or a raw string expression.
What: - Completes the capture command's measurement loops so every declared cell gets the contract's 3 controlled runs x 15 warmed samples: stages 1 and 2 restart the daemon per sample, and the revision-driven browser stages (5, 6, 7) loop over real published revision changes rather than being measured once. Cmd-K is measured from a fresh page per sample so each is a real open, not an already open palette. - tests/perf/summarize.ts + perf:summarize — recomputes every summary from the stored raw samples, so interpretation numbers cannot drift from the evidence. - tests/perf/productionIsolation.test.mjs — proves the ordinary production artifact carries no benchmark entry or probe identifier, no production source or build script reaches the benchmark, the daemon hook is inert by default and unreachable from any API surface, and the shipped bundle has no analytics. - docs/testing/TIME_TO_ANSWER_BASELINE.md — the measured baseline and the evidence-backed interpretation. - docs/testing/TIME_TO_ANSWER_EVIDENCE.md — the documented procedure and the slice's completed state. - Fixes two harness defects found by running it for real: a run-shape bug that nested samples one level too deep, and raw samples not being preserved when artifact assembly failed. `record()` now fails at the measurement site, naming the cell. Measured baseline 1 (median of three run p95, ms; source b6904d5, artifact sha256 67f9b78b4e358c77d3980fbc8752dbcd33141e1898bb9112a9b9312224eaf9d2): 25 / 100 / 250 containers daemon start -> listener 41.99 / 109.84 / 254.00 listener -> first Docker model 12.71 / 59.90 / 165.76 Docker observation 7.26 / 9.10 / 13.07 Compose projection 1.15 / 1.18 / 1.25 publication -> Node observation 1307.47 / 636.12 / 1623.47 (0.42 .. 1976) notification -> coherent model 12.70 / 20.60 / 34.60 coherent -> useful render 12.70 / 20.60 / 34.60 buildModel() 0.50 / 0.90 / 1.50 findings derivation 0.07 / 0.20 / 0.92 legacy topology layout 2.10 / 25.80 / 174.40 Cmd-K open + query 55.10 / 45.40 / 34.90 production bundle load 49.30 / 59.30 / 50.20 Interpretation (no recommendation, no optimization): the fixed-interval publication->Node observation floor is the largest single contributor (87/64/68% of summed stage medians) and is a phase-of-poll distribution, not network latency — #337 owns it. Cold start is dominated by process bring-up, not Docker work, and Compose projection is a measured but small cost on these fixtures even though it currently executes inside the Docker publication budget — #336 owns decoupling and should be judged against these numbers rather than an assumed large win. Legacy Home layout is 174 ms at 250 containers and scales super-linearly — #338 owns that decision. Why: #335 is the measurement authority for the Speed epic. This lands the baseline the later issues must prove themselves against, with the environment pinned and the promotion gate RED-checked before any comparison is trusted.
…ne (#335) Every accepted finding from both adversarial reviewers, except the artifact itself, which is recaptured after this commit. P1 — Cmd-K measured palette-open, not query-to-results: - tests/perf/browserProbe.js commandQuery now snapshots the UNFILTERED command list first (the palette renders every command on open, so "an item exists" passed even with filtering entirely broken), types a query with a known expected result, and only stops the timer when the list has actually changed AND still contains that token. The token is the fixture-derived `fixture-service-0` when present, otherwise derived from the rendered list. P1 — stage 5 phase-locking: - the trigger is now jittered by a uniform sub-interval delay per sample, so the measurement describes the real poll-wait distribution instead of one fixed phase offset between the daemon's 2 s refresh loop and the API's 2 s poller. The contract and the doc state that de-correlation is the harness's doing and that the underlying mechanism still runs on two fixed cycles. P2s: - probe daemons write to their own sink, so the stages documented as warmed no longer mix cold-start first observations into the resident daemon's samples; - the capture refuses a dirty worktree and refuses to run when the metadata's sourceRevision/harnessRevision do not match the checked-out commits, and the environment now records BOTH the product revision and the benchmark-harness revision, so a baseline is reproducible from a commit; - findingsDerivationMs is bucketed as backend-collection (it runs in the daemon during publication, not in the browser) and its doc text says the fixture derives no findings, so it measures the empty-derivation path; - stages 6 and 7 are no longer duplicates: stage 6 ends on the first commit that CHANGES RENDERED TEXT anywhere, stage 7 on a text-changing repaint of the Home content region only, and stage 7 is declared only for fixtures whose published change demonstrably repaints Home (reference-* and docker-topology-change); - the environment records harnessRevision and dockerRevision is now informational for compatibility (no measured stage exercises the host Docker daemon), so an engine upgrade cannot fail an unrelated comparison. P3s: - ssePollIntervalMs is derived from the API's own source default and passed to the API explicitly, so the pin cannot drift from the interval that ran; - scenario premises are asserted: the provider-only fixture fails if its fixture inventory changed during the run, and the unavailable-provider fixture fails if no optional provider was non-empty/fresh; - --fixtures is gated behind DOCKERMAP_BENCH_DEBUG (a partial run can never satisfy the closed matrix, so it was a guaranteed-waste footgun); - promotion RED-checks now pin the median-of-three aggregation (one slow run must pass, two slow runs must fail) and assert that a differing informational dockerRevision does NOT fail a comparison; - the contract's stage/bucket documentation is corrected for stages 4-9. Gates: vitest src/lib/performance 31/31.
…335) Baseline 2, captured from committed revision 0714c87a (product and harness revisions both recorded; artifact sha256 38e0c650121f81ec69c2dd1ddc18191c98c091836652bd8ecd7570fc0d006197, stored outside the repo at /srv/jonas/evidence/dockermap/time-to-answer-baseline-2.json). Every number is the median of three run p95 values recomputed from the raw samples. Baseline 1 is documented as rejected and superseded, with the four reasons it was invalidated. Corrected findings versus the rejected baseline (25 / 100 / 250 containers, median of three run p95, ms): daemon start -> listener 43.54 / 119.47 / 234.29 listener -> first Docker model 13.51 / 54.77 / 156.83 Docker observation 7.91 / 9.36 / 15.34 Compose projection 1.08 / 1.06 / 1.08 publication -> Node observation 1971.87 / 1828.38 / 1918.79 (spread 0.27 .. 1991) notification -> coherent model 15.20 / 19.30 / 52.10 coherent -> useful render 15.20 / 19.30 / 52.10 buildModel() 0.60 / 0.70 / 1.80 findings derivation 0.07 / 0.30 / 0.91 (backend-collection) legacy topology layout 2.20 / 24.20 / 164.40 Cmd-K open + query 51.40 / 37.60 / 26.60 production bundle load 53.10 / 55.60 / 61.50 Stage 5 now spans the interval instead of a 66 ms band: baseline 1's 1283-1349 ms for reference-25 was a phase offset locked by the harness's own startup sequence, not a sample of the poll wait. The de-correlation is the harness's doing and the two fixed 2 s cycles are unchanged, which the contract and the doc both say. Also corrected: 46 cells (not 21); stage 7 declared only for the four fixtures whose published change demonstrably repaints Home; findings derivation described as the daemon-side empty-derivation path; variance statements replaced with the measured ranges; dockerRevision described as informational. Gates: test:perf 13/13, vitest src/lib/performance 31/31, npm run check green.
…3, partial) Round-3 review findings 1, 2 and 4, plus the contract's cold/warm classification. Findings 3 (stage 6 vs 7 seam), 5 (documentation corrections) and the baseline-3 recapture are NOT in this commit — see the issue comment. 1. Cold-start contamination of warmed stages: - the contract now classifies every stage explicitly (TIME_TO_ANSWER_STAGE_KIND: cold-start | warmed-repeated) and identifies scenario cells (isScenarioCell); - new pure `splitWarmedObservations()` discards EXACTLY ONE warm-up observation from a warmed window and returns it for audit; it refuses a window shorter than count + 1 so no arbitrary slow sample can be dropped instead; - the capture now waits for samples + 1 bench observations per warmed daemon stage and records the remainder, keeping the discarded warm-up in `warmUpObservations` for the raw audit trail. Cold-start stages (process start, first Docker publication) are untouched, because there the first observation IS the measurement; - RED-checks: a 99.9 ms cold first observation beside 2 ms steady samples is proven to leave the recorded summary under 10 ms, and a short window throws. 2. Daemon binary provenance: - new `assertDaemonBinaryProvenance()` requires two lowercase sha256 digests and fails on mismatch, so the executable that produced every daemon-side number is bound to the recorded revision. The benchmark does not claim bit-for-bit reproducible Rust builds across machines; it proves which binary THIS capture executed; - emit-metadata now BUILDS the release daemon (`cargo build --release --locked -p dockermap-daemon`), then records `daemonBinarySha256`, `daemonBinaryBuild` and `cargoRevision`; the capture verifies the digest before the run and will verify it again after, failing closed on substitution or drift. 4. Cmd-K hardening: success now also requires the filtered item count to be STRICTLY lower than the unfiltered count, so a reorder-only or no-op filter cannot satisfy the assertion (the palette prepends an "Ask Copilot" item on any query, which would otherwise make a reordered list pass). Environment gains daemonBinarySha256, daemonBinaryBuild and cargoRevision; all three test-suite environment fixtures updated. Gates: vitest src/lib/performance 34/34.
…contract (#335) Round-3 finding 5 (all accepted round-2 documentation/evidence issues except the stage-6/7 definitions, which follow once the benchmark seam lands). - matrix count corrected to 44 cells in both documents; - the bucket table no longer lists findingsDerivationMs under browser-model (the contract buckets it backend-collection; it runs in the daemon); - the promotion gate sentence now states that compatibility excludes sourceRevision (by design) AND dockerRevision (recorded but informational, since no measured stage exercises the host Docker daemon), and that the other 15 fields must match; - new sections defining cold-start vs warmed-repeated vs scenario cells, and the exactly-one-discarded-warm-up rule with the warm-up retained separately in the raw audit trail and excluded from p95; - new section on capture discipline: sourceRevision vs harnessRevision meaning, and the requirement to check out the revision the artifact names; - new section on daemon binary provenance: the canonical locked release build, daemonBinarySha256/daemonBinaryBuild/cargoRevision, digest verified before and after the capture, no bit-for-bit reproducibility claim; - new section stating that the 2000 ms SSE polling mechanism is production behaviour that is deliberately unchanged, that the jitter de-correlates only the measurement, and that the stage must never be described as network latency; - baseline 1 and baseline 2 are labelled REJECTED historical attempts with the reasons, and the baseline document no longer presents either as current authority. No code changes. Gates unchanged (vitest src/lib/performance 34/34 at c622da2).
What: - Stage 6 (notificationToCoherentModelMs) now ends at the REAL application seam: useSystemModel records one opaque acceptance event (timestamp + model revision token) at the instant a fetched snapshot/runtime pair becomes the coherent model the UI renders. It is application state, never derived from a DOM mutation. - Stage 7 (coherentModelToUsefulRenderMs) starts at that stage-6 timestamp and ends when the accepted revision's EXPECTED Home content is present - in a commit the application stamped with that accepted revision - plus exactly one bounded requestAnimationFrame. No sleeps. - The seam is real product source (apps/web/src/lib/performance/modelAcceptance.tsx) gated by the compile-time flag __DOCKERMAP_BENCH_ACCEPTANCE__: the production Vite config defines it "false" (dead-code eliminated), and a benchmark-mode application build (tests/perf/benchAppVite.config.mjs) defines it "true". The application source is built twice, never copied. - The capture serves stages 6/7 from the benchmark-mode build and stages 11/12 (Cmd-K, production bundle) from the ordinary production build, and refuses to run unless the benchmark-mode artifact carries the seam and the production artifact does not (assertBuildIsolation). - New stage-6/7 independence control: 3 control samples per stage-6/7 fixture with a 250 ms artificial presentation delay injected AFTER acceptance. Stage 6 must not move beyond max(30 ms, 25%); stage 7 must absorb at least 70% of the delay. Enforced before any artifact is assembled; raw pairs and a per-sample audit trail are written beside the artifact as <output>.stage-seam.json. - The fixture generation delta is now product-visible (generation g stops the first g containers), so the stage-7 expected-content check discriminates a fresh render from a stale one; expectedExitedCount derives the expectation from the same generator the fixture daemon serves. - Chain of custody is checked per sample: the revision the API observed through the live SSE path must be the revision the browser accepted and rendered. - Docs: stage 6/7 definitions, the two-clock distinction, benchmark-build isolation, the independence control, the generation delta, the 44-cell matrix, and the per-build stage mapping. Why: Baseline 2 was rejected because its stage 7 was element-for-element identical to stage 6 in all 180 samples: both stages came from one DOM-derived clock. This change makes stage 6 the application's acceptance instant and stage 7 the presentation of that accepted revision, and adds the control that proves the two clocks are independent. How checked: - npm run test:perf -> 17/17 (production isolation incl. seam confinement and the flag-is-false check, fixture generation delta, independence RED cases) - npm run test:web -> 554/554 (incl. new timeToAnswerIndependence.test.ts) - production bundle inspected after `npm run build --workspace @dockermap/web`: 0 occurrences of every seam identifier; benchmark-mode build contains the sink - npm run typecheck --workspace @dockermap/web passes Remaining risk / follow-up: No baseline may be captured until the round-3 smoke demonstrates the independence control, the warm-up discard and the artifact assembly on the committed revision.
What: emit-metadata.mjs builds the release daemon as `cargo build --release --locked -p dockermap-daemon --manifest-path crates/Cargo.toml`, and the recorded daemonBinaryBuild slug and the evidence doc name the same command. Why: the cargo workspace lives under crates/, so the previous command failed with "could not find Cargo.toml" and no capture could emit metadata at all - the pinned binary provenance step was unrunnable. How checked: npm run perf:metadata now completes, builds the release daemon and records daemonBinarySha256.
What: - The probe arms on the monotonic acceptance SEQUENCE and waits for the next accepted revision, instead of "any revision other than the last one seen". The previous predicate selected a stale acceptance event whenever a publication landed between arming and the trigger. - Stage 6 starts at the notification of the ACCEPTED revision, recovered from a per-revision notification log, instead of assuming the latest notification belongs to it. - The API observation keeps listening for 200 ms after the first change and keeps every revision it saw for the sample; the browser's accepted revision must be one of them (a single fixture change can publish more than once). - preserveRaw creates its destination directory, so raw samples survive even when the raw directory does not exist yet. Why: the round-3 smoke failed with "the accepted revision ...-15 is not the revision the browser was notified of (...-16)" - the harness's own arming logic mis-attributed a revision, not the seam. How checked: round-3 smoke on reference-25 (1 run x 3 samples) - see the run log; the independence control, warm-up discard and Cmd-K checks are exercised there.
What: - Stage 5's API observation is now ARMED before the trigger and stopped after the sample's browser measurement resolves, instead of running as a one-shot call after the trigger with a fixed post-collection window. Every revision belonging to the sample is captured, however many publications it produces. - Stages 6/7 select the (accepted model, rendered content) PAIR that carries the sample's expected Home content and that revision's stamp, instead of assuming the first acceptance after arming owns the change. A provider-state-only publication between the trigger and the inventory publication no longer mis-attributes the sample. - The audit trail records how many acceptances were skipped and which revisions the API observed for the sample. Why: the second round-3 smoke failed with "the Home content for the accepted revision ...-21 never rendered (expected Offline=4)" - the daemon published an intermediate revision (-21) with unchanged Home metrics before the inventory revision (-22), and the harness assumed one acceptance per sample. How checked: round-3 smoke on reference-25 (1 run x 3 samples) plus the stage-6/7 independence control; see the run log.
…ision (#335) What: PublicationObserver.stop() now waits (bounded by one poll interval plus a 1.5 s margin) until the harness's own SSE connection has emitted the revision the browser accepted, and the sample loop passes that revision in. Why: each SSE connection runs its OWN poll timer (apps/api/src/index.ts holds a per-connection interval and a per-connection lastEmittedRevision), so the harness's stream legitimately lags the browser's by up to one interval. The chain check failed with "the browser accepted revision ...-21, which the API never observed for this sample (observed: ...-20)" even though both streams were fed by the same real mechanism. The recorded stage-5 duration is fixed at the first emission and is unaffected by the extra wait. How checked: round-3 smoke on reference-25 (1 run x 3 samples).
What: every capture now writes <output>.harness-evidence.json beside the artifact, carrying (a) the stage-6/7 independence control (verdict, per-run sample sets and the per-sample acceptance audit) and (b) warm-up retention: for every daemon-side warmed cell the COMPLETE samples+1 observation window in the order the daemon produced it, with the discarded warm-up at index 0, plus proof that the recorded samples equal that window minus the warm-up. Why: splitWarmedObservations discarded the cold first observation and kept the value only in memory, so no reviewer could check that a slow warm-up value was not silently dropped. The capture now refuses to emit an artifact when a window is missing, shorter than samples+1, discards anything other than the first observation, or does not match the stored run. How checked: round-3 smoke on reference-25; the retention checks run on every capture and the evidence file is written before the artifact is validated.
…#335) What: the probe now arms in one of two modes, chosen by the closed matrix: - content mode (every fixture that also declares stage 7): the sample ends on the (accepted model, rendered content) pair carrying the expected Home content for the triggered change; - acceptance-only mode (provider-only-revision-change, unavailable-optional-provider): stage 6 ends at the acceptance instant, because those fixtures publish a provider-state revision with NO inventory change, so no Home repaint exists to wait for and stage 7 is not declared for them. Why: the full baseline-3 attempt aborted on provider-only-revision-change with "the Home content for the triggered change never rendered (expected Offline=1; story=... storyValue 0 ...)" - the content-mode end condition can never fire on a fixture whose inventory does not change. The three reference fixtures had already completed at full scale, so only those two cells were affected. How checked: round-3 smoke over reference-25 (content mode) and provider-only-revision-change (acceptance-only mode).
… model (#335) What: the harness now records which revision each paired /api/snapshot and /api/runtime/map fetch actually delivered, and stage 6 starts at the browser notification that preceded that fetch cycle. The requirement that the accepted revision equal a revision the harness's OWN SSE stream announced is replaced by the browser-side provenance chain, and the stream overlap is recorded as acceptedRevisionInApiStream evidence instead of gating. Why: the second full baseline-3 attempt aborted on unavailable-optional-provider with "the accepted revision ...-32 was never notified to this browser". The daemon is read per request: /daemon/health (what the SSE stream carries) and /daemon/snapshot (what the app accepts) are separate reads and legitimately hold different revisions while provider state churns, so the previous token-equality proxy was wrong, not the seam. Stage 6 is now anchored to the fetch the model came from, which is stronger evidence than the proxy it replaces. How checked: round-3 smoke over reference-25 (content mode), docker-topology-change, provider-only-revision-change and unavailable-optional-provider (acceptance-only).
What: the three static servers now bind an OS-assigned port (port 0) and report the port actually bound, and every spawned child that binds a reserved port is started through a bounded retry (fresh port for the daemon and the per-sample startup probes; the SAME port for the API, whose port is baked into the production build). Why: the third full baseline-3 attempt aborted with 'listen EADDRINUSE: address already in use 127.0.0.1:37501'. reservePort() binds, closes and hands the number back to the pool, so the port can be taken before the child binds it - and the API port in particular stayed reserved-but-unbound for minutes between fixtures. An unrelated process or a not-yet-reaped child could therefore abort an expensive capture. How checked: round-3 smoke over the four fixtures that exercise both stage-6 modes.
…335) What: docs/testing/TIME_TO_ANSWER_BASELINE.md is rewritten as the baseline-3 record: the pinned environment, the verbatim recomputed 44-cell table, the stage-5 distribution, stage 6/7 medians, the enforced stage-6/7 independence control with its per-fixture verdicts, the acceptance audit, bucket shares, top contributors, the Compose contribution and the auditable warm-up retention. Baselines 1 and 2 are kept as rejected history WITHOUT their numbers, so they cannot be read as current authority. TIME_TO_ANSWER_EVIDENCE.md now states that baseline 3 exists and that the round-3 independent review is the gate before it becomes the authority. Why: baseline 3 was captured from committed revision cf77e8b (25.6 min, 44 cells x 3 runs x 15 samples = 1980 raw samples, exit 0) and the numbers had to be recorded from the recompute path rather than from memory. How checked: every table is the verbatim output of ; the artifact, the harness evidence and the recomputed summary are stored in /srv/jonas/evidence/dockermap/time-to-answer/ with recorded sha256 digests.
…-5 phase sweep (#335) What (methodology revision 2, dockermap-v1/time-to-answer-methodology-2): - Stage 5 now DRIVES the publication phase instead of hoping for one. The poll interval is divided into 15 declared phases (one per recorded sample per run, each at the centre of its division), every run sweeps them ascending, and the harness controls the phase by choosing when it connects its observation stream relative to a predicted daemon publication. Each sample records its declared and observed phase, and asserts its own control: observed publication matches the prediction, observed latency lands on the declared phase within 60 ms, and the observation arrived through the real API poller path. - New validity guards (assertPollPhaseSweep): every declared phase represented at least twice, no uncontrolled phase, no non-poller observation, span >= half the interval, and the earliest declared phase at least half an interval slower than the latest. A narrow-band sweep is REJECTED - it is a RED test, not an aspiration. - Stage 5 reports a phase curve plus a PHASE-NORMALIZED p95 computed from the predeclared uniform grid, explicitly documented as a statement about the mechanism under uniformly sampled phase offsets - not an observed user-traffic distribution and not network latency. - Fixed warm-up protocol: five predeclared warm-up observations before the 15 measured samples, chosen from the round-3 windows (the second observation reached 2.15x the window median in 5/30 windows; index 5 onward stayed within 1.29x). Warm-ups are never dropped adaptively; a declared stationarity check (0.5x-1.5x band on the final two warm-ups against the measured median) INVALIDATES a window instead of trimming it. Every warm-up observation is retained. - Provenance is no longer compatibility: daemonBinarySha256, sourceRevision and harnessRevision are recorded but not required to match, so a candidate that changes crates/ (#336) is comparable; methodologyVersion, the toolchain, browser, fonts, fixture revision, build mode and polling configuration still must match. - The documented before-and-after daemon binary verification now exists: the digest is re-hashed after the run and the capture fails closed on a mismatch, with both digests recorded in the harness evidence. - The slow-Compose scenario premise is asserted (the project really declares its 400 services) instead of merely named, and the capture asserts the application page never contacted the daemon directly. - Docs corrected: stage 7 is a bounded render/presentation confirmation rather than "exactly one frame", the promotion section states the provenance/compatibility split with the exact field counts, the rejected-capture paths are right, and the stage-5 sections describe the sweep and the normalized figure. Why: round-3 review rejected baseline 3. Its stage-5 samples were a drift-locked sawtooth (reference-25 covered 6.0% of the 2000 ms interval with consecutive deltas under 19 ms) even though the harness slept a uniform random delay, because both the daemon refresh loop and the API poller are fixed 2 s loops; its single discarded warm-up still left a 2.15x-of-median observation inside the measured window; it gated promotion on a rebuilt binary digest; and it documented a post-run verification the code did not perform. How checked: npm run test:perf 20/20 (including the new methodology drift guard), vitest performance suites 56/56 (including the narrow-band, uncontrolled-phase, missing-phase, non-poller-path and stationarity RED cases), web typecheck clean.
…335) The declared minimum of three samples per phase needs the three controlled runs the full protocol requires, so a debug probe (which can never emit an artifact) would fail the sweep guard for a reason that has nothing to do with the measurement. The minimum is now an explicit guard parameter, and the capture passes 1 only when the declared run count is below the protocol's; every other guard still applies.
…om daemon start (#335) The methodology smoke failed its first stage-5 sample: declared phase 66.7 ms produced 2051.5 ms against an intended 1933.3 ms (118.2 ms error, tolerance 60 ms). Cause: the predicted publication instant came from a single observed gap, which carries one cycle's refresh work plus the tracker's own start transient. Stage 5's deepest declared phases put the publication just after a poll tick, so a period error maps one-for-one into phase error, and the error was large enough to land the sample in a neighbouring phase bucket. Fix: the tracker now runs from daemon start (so the grid is known before the browser stages begin), requires four publications before the first sample is placed, and estimates the period by least-squares over the last eight publication instants instead of one gap. The control tolerance stays at 60 ms - below half the 133 ms grid step, so a sample cannot bucket into a neighbouring phase.
…nge (#335) The methodology smoke failed again at the deepest declared phase (66.7 ms declared, 1933.3 ms intended, 2456.3 ms observed). Cause: the daemon can publish more than one revision per refresh cycle (the inventory publication followed by a provider-state publication), and the tracker fitted its grid over every revision change, so the extra intra-cycle publications skewed the period estimate - which the deepest phases amplify one-for-one into phase error. The tracker now collapses revisions into refresh-cycle leaders (changes within a quarter-interval belong to one cycle), fits the period over the last eight cycle leaders, and each sample attributes itself to the detected publication nearest the predicted cycle. Failure messages now carry the connect, prediction, publication and tick instants so a control failure is diagnosable without a re-run.
…ctually carried (#335) The API's poll tick emits whatever revision is current at tick time, so the publication a sample measured is the LAST revision change at or before the tick. The previous attribution (nearest to a predicted cycle instant) could pick a publication the tick never carried: the smoke showed a sample attributed to a publication 510 ms before the prediction and 444 ms before the connection. Also: cycle boundaries are now identified as publications at least 60% of the poll interval apart, because one refresh cycle can publish a provider-state revision a few hundred milliseconds after the Docker snapshot, and folding those extras into the grid skewed the period estimate. Each sample records whether it was attributed to a cycle boundary.
) The methodology smoke kept missing the deepest declared phase (66.7 ms declared, 1933.3 ms intended, 453-2456 ms observed): the grid was fitted over revision changes, but one refresh cycle publishes the Docker snapshot and can publish a provider-state revision up to ~1.5 s later, so the fitted period carried those extras and the prediction landed hundreds of milliseconds off - which the deepest phases amplify one-for-one into phase error. The tracker now uses the daemon's own snapshot timestamp: /daemon/health carries lastUpdated, which the daemon stamps once per refresh cycle in epoch milliseconds, so the grid is one exact instant per cycle with no detection lag (the stamp is aligned once to the harness's monotonic clock). Attribution stays causal - the tick carries the last cycle boundary at or before it - and each sample records whether the tick also carried a newer intra-cycle revision.
Two fixtures cannot be phase-driven at all: provider-only-revision-change and unavailable-optional-provider advance their revisions from the daemon's own host provider collection, and the harness has no input that makes the daemon publish at a chosen instant. The smoke showed the consequence - waiting for a spontaneous revision landed the observation on the fourth poll tick, 6.1 s after the predicted cycle. Rather than claim a phase those cells never drove, the design now declares which fixtures are phase-controlled (POLL_PHASE_CONTROLLED_FIXTURES: the three reference fixtures and docker-topology-change) and which are free-running: - phase-controlled cells keep the declared grid, the per-sample control check and the full sweep guards (coverage, span, direction, poller path); - free-running cells record the phase they ACHIEVED (declared phase = the bucket the observation landed in), pass a separate guard that requires real poller observations and a minimum sample count, and derive NO phase-normalized figure; - the two guards refuse each other's samples, so a cell cannot claim a sweep it did not drive.
…ndow (#335) Baseline 4's capture aborted on its third fixture: declared phase 200.0 ms produced 1877.9 ms, with the observation arriving 3.5 ms after the connection and attributed to a boundary 2.1 s earlier. Cause: the API emits the current revision immediately on connect, and previousRevision came from the tracker's last poll, which can be up to 10 ms stale. A revision published in that window made the very first frame look like a new publication, so the sample was attributed to the PREVIOUS cycle's boundary. Fix, in two parts: the pre-sample revision is now read fresh from the daemon health endpoint right before the stream opens, and a sample only accepts a frame that could be its own observation - for a controlled sample, at or after the predicted cycle (safe because its poll tick cannot land earlier while the connect frame precedes it by the declared phase); for a free-running sample, a margin after the connect. Frames ignored this way are counted in the sample's evidence.
Provenance: Terra issue-runner; controlled Stage-6/7 protocol implementation and static integration coverage.
Provenance: Terra issue-runner; fixes reviewer-identified output and fixture-capacity contract.
Provenance: Terra follow-up worker; withhold Stage-5 publication until the exact accepted model pair is armed.
Provenance: follow-up worker checkpoint after 7722a79. Validation: node --test tests/perf/productionIsolation.test.mjs
…evidence (#335) What: publish the Baseline-4 authority and reconcile the methodology-8 benchmark runbook. Why: document the composite evidence workflow, controlled Stage-5 capture, and supporting seam-isolation control. How checked: reviewed the seven-step runbook, searched stale methodology terms, verified performance-code provenance, and ran git diff --check.
What: correct the Stage 6/7 seam-isolation protocol documentation.\n\nWhy: the control is a dedicated protocol with its own evidence rather than general-capture harness evidence.\n\nHow checked: inspected the scoped section, remaining harness-evidence references, working-tree scope, and performance-code baseline commit.
Joncallim
marked this pull request as ready for review
September 27, 2026 03:54
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Status: complete — Baseline 4 established as the comparison authority for #336/#337
#335 is a pure measurement authority. It measures; it optimizes nothing, changes no rule,
no severity, no provider cadence and no response contract.
Part of epic #333. Issue: #335.
Baseline 4 identity
dockermap-v1/time-to-answer-methodology-8bdce6ae354d757d3318c514e10d620edb497918eend-to-end+ 6controlled-poll-phase),3 controlled runs × 15 measured samples per record, 1 980 raw samples, each record
naming fixture, stage, measurement protocol, source evidence file and checkpoint
sample is observation 61; burn-in retained for audit and excluded completely from every
timing summary; no stationarity claim
after the run
/srv/jonas/evidence/dockermap/time-to-answer/bdce6ae/—
time-to-answer-baseline-4.json(sha256916610bdd4767fb01a4f57e29f318aa504836875cbb4d87ba615c70cb6f7300),time-to-answer-stage5.json,time-to-answer-baseline-4-composite.json,time-to-answer-baseline-4-composite.json.supporting-evidence.json,time-to-answer-baseline-4-summary.md,baseline-4-verification.json,the gate logs and the review verdicts
Documented record:
docs/testing/TIME_TO_ANSWER_BASELINE.md.Interpretation contract:
docs/testing/TIME_TO_ANSWER_EVIDENCE.md.Protocols, kept separate
end-to-end— every ordinary baseline row, including the normal Stage-6 and Stage-7timings, produced by
npm run perf:time-to-answer.controlled-poll-phase— the dedicated Stage-5 protocol(
npm run perf:stage-five), the sole owner of arm → mark → trigger →identity-acknowledgement. Ten declared phases across the 2000 ms poll interval, 45
distinct trigger ids per fixture, maximum absolute observed-vs-declared phase error
≤ 1.0 ms. Its published figure is phase-normalized and is not user-traffic or network
latency.
controlled-stage6-stage7-seam-isolation— supporting validity evidence only(
npm run perf:independence). Status at this checkpoint: FAIL —[independence] FAIL: Error: control delay did not begin after the acceptance timestamp,after the control had already been re-scoped once. Disclosed limitation:
publication-level causal identity unavailable,validatesDaemonToBrowserAttribution: false. Its samples are never Baseline-4 timingrows and it is not daemon→browser attribution evidence. This result is disclosed, not
hidden or reinterpreted.
The two exclusion gates that must pass before the capture did: sustained preconditioning
PASS across 12 sequential fresh-browser runs, and the Stage-5 phase-control gate PASS
across all six declared fixtures (grid span ≈ 1 795–1 799 ms, tolerance 90 ms).
Reviews at the final head
Baselines 1, 2 and 3 were rejected by earlier independent review and are not authority for
anything; their numbers do not appear as current figures.
Two fresh review lanes were run against the final head:
gpt-5.6-luna): PASS — all five questions PASS, no findings,no gaps; the published tables reproduce against the recomputed summary and the
independent verification records 44/44 raw-provenance matches with zero boundary,
retention or duplicate mismatches.
claude-sonnet-5): APPROVE, no findings, nothing leftunverified. It independently re-ran
npx tsx tests/perf/summarize.ts --artifact <composite>and confirmed the output isbyte-identical to the published summary, then checked the structural counts, the
Stage-5 ownership wording against
captureStageFive/assembleCompositeEvidence, andthe command list against
package.json1:1. Its five questions — validity as comparisonauthority despite the seam-isolation limitation, claim scope of the supporting control,
independence of the normal Stage-6/7 timings, determinism of 60+15, and agreement
between the documents and the executable contract — were all answered YES.
Both reviewers accepted the underlying measurement math and methodology questions in the
first round and blocked only on the stale documentation; the final tranche is
documentation-only because the captured numbers are pinned to harness revision
bdce6ae,and any change to
tests/perf/**orapps/web/src/lib/performance/**would move thebranch head off the harness that produced them.
What this baseline does not claim
Not real-Docker latency (deterministic fixture daemon), not a real network or user-traffic
distribution, not network latency for Stage 5, no claim that host publications occur
uniformly across poll phase, no proof the model is complete (stages 1/2 and stage 6 end at
coherence), no daemon-publication attribution from the failed seam-isolation control, and
no permission to optimize — #336/#337/#338 must be compared under
max(baseline × 1.25, baseline + 2 ms)in a compatible pinned environment.