Skip to content

Establish end-to-end time-to-answer baselines and regression gates (#335) - #346

Merged
Joncallim merged 81 commits into
mainfrom
dm/335-time-to-answer-baseline
Sep 27, 2026
Merged

Joncallim merged 81 commits into
mainfrom
dm/335-time-to-answer-baseline

Conversation

@Joncallim

@Joncallim Joncallim commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Status: complete — Baseline 4 established as the comparison authority for #336/#337

#335 is a pure measurement authority. It measures; it optimizes nothing, changes no rule,
no severity, no provider cadence and no response contract.

Part of epic #333. Issue: #335.

Baseline 4 identity

  • methodology: dockermap-v1/time-to-answer-methodology-8
  • product and capture-harness revision: bdce6ae354d757d3318c514e10d620edb497918e
  • composite authority: 44 records (38 end-to-end + 6 controlled-poll-phase),
    3 controlled runs × 15 measured samples per record, 1 980 raw samples, each record
    naming fixture, stage, measurement protocol, source evidence file and checkpoint
  • conditioning: fixed 60 burn-in observations then exactly 15 measured, first measured
    sample is observation 61; burn-in retained for audit and excluded completely from every
    timing summary; no stationarity claim
  • capture duration 105.1 min; daemon release binary digest verified identical before and
    after the run
  • evidence root (outside the repository, per the evidence-artifact policy):
    /srv/jonas/evidence/dockermap/time-to-answer/bdce6ae/
    — time-to-answer-baseline-4.json (sha256
    916610bdd4767fb01a4f57e29f318aa504836875cbb4d87ba615c70cb6f7300),
    time-to-answer-stage5.json, time-to-answer-baseline-4-composite.json,
    time-to-answer-baseline-4-composite.json.supporting-evidence.json,
    time-to-answer-baseline-4-summary.md, baseline-4-verification.json,
    the gate logs and the review verdicts

Documented record: docs/testing/TIME_TO_ANSWER_BASELINE.md.
Interpretation contract: docs/testing/TIME_TO_ANSWER_EVIDENCE.md.

Protocols, kept separate

  • end-to-end — every ordinary baseline row, including the normal Stage-6 and Stage-7
    timings, produced by npm run perf:time-to-answer.
  • controlled-poll-phase — the dedicated Stage-5 protocol
    (npm run perf:stage-five), the sole owner of arm → mark → trigger →
    identity-acknowledgement. Ten declared phases across the 2000 ms poll interval, 45
    distinct trigger ids per fixture, maximum absolute observed-vs-declared phase error
    ≤ 1.0 ms. Its published figure is phase-normalized and is not user-traffic or network
    latency.
  • controlled-stage6-stage7-seam-isolation — supporting validity evidence only
    (npm run perf:independence). Status at this checkpoint: FAIL —
    [independence] FAIL: Error: control delay did not begin after the acceptance timestamp,
    after the control had already been re-scoped once. Disclosed limitation:
    publication-level causal identity unavailable,
    validatesDaemonToBrowserAttribution: false. Its samples are never Baseline-4 timing
    rows and it is not daemon→browser attribution evidence. This result is disclosed, not
    hidden or reinterpreted.

The two exclusion gates that must pass before the capture did: sustained preconditioning
PASS across 12 sequential fresh-browser runs, and the Stage-5 phase-control gate PASS
across all six declared fixtures (grid span ≈ 1 795–1 799 ms, tolerance 90 ms).

Reviews at the final head

Baselines 1, 2 and 3 were rejected by earlier independent review and are not authority for
anything; their numbers do not appear as current figures.

Two fresh review lanes were run against the final head:

  • acceptance verifier (gpt-5.6-luna): PASS — all five questions PASS, no findings,
    no gaps; the published tables reproduce against the recomputed summary and the
    independent verification records 44/44 raw-provenance matches with zero boundary,
    retention or duplicate mismatches.
  • adversarial issue review (claude-sonnet-5): APPROVE, no findings, nothing left
    unverified. It independently re-ran
    npx tsx tests/perf/summarize.ts --artifact <composite> and confirmed the output is
    byte-identical to the published summary, then checked the structural counts, the
    Stage-5 ownership wording against captureStageFive/assembleCompositeEvidence, and
    the command list against package.json 1:1. Its five questions — validity as comparison
    authority despite the seam-isolation limitation, claim scope of the supporting control,
    independence of the normal Stage-6/7 timings, determinism of 60+15, and agreement
    between the documents and the executable contract — were all answered YES.

Both reviewers accepted the underlying measurement math and methodology questions in the
first round and blocked only on the stale documentation; the final tranche is
documentation-only because the captured numbers are pinned to harness revision bdce6ae,
and any change to tests/perf/** or apps/web/src/lib/performance/** would move the
branch head off the harness that produced them.

What this baseline does not claim

Not real-Docker latency (deterministic fixture daemon), not a real network or user-traffic
distribution, not network latency for Stage 5, no claim that host publications occur
uniformly across poll phase, no proof the model is complete (stages 1/2 and stage 6 end at
coherence), no daemon-publication attribution from the failed seam-isolation control, and
no permission to optimize — #336/#337/#338 must be compared under
max(baseline × 1.25, baseline + 2 ms) in a compatible pinned environment.

What:
- apps/web/src/lib/performance/timeToAnswerEvidence.ts — the closed schema and
  math for the controlled time-to-answer benchmark: the 12 measured stages with
  their bucket (backend-collection / transport-notification / browser-model /
  rendering / search), an explicit `measures` and `doesNotProve` statement per
  stage, the three reference fixtures (25/100/250 containers) plus the four
  scenario fixtures, the fixture x stage matrix derived from the stage lists,
  the pinned-environment allowlist (runner/CPU/OS/kernel, Node/Rust/Docker
  revisions, Chromium revision and flags, fonts, production build, fixture and
  source revision), raw-sample validation, nearest-rank p95, median-of-three
  controlled runs, and the reviewed max(baseline x 1.25, baseline + 2 ms)
  promotion gate.
- apps/web/src/lib/performance/timeToAnswerEvidence.test.ts — shape/math tests
  plus hostile cases (fabricated summary field, wrong baseline id, truncated
  matrix, 14 or 2 samples, negative and non-numeric samples, unknown stage,
  duplicate record, undeclared fixture, extra or unsafe metadata fields,
  non-chromium engine, incompatible pinned environment, over-limit candidate).

Why:
#335 owns measurement, fixtures, evidence format and promotion gates only.
There is no performance authority outside Atlas today, so no later optimization
claim in #336/#337/#338 can be compared against anything. This contract is that
authority's closed shape: it stores raw samples only and derives every summary
during review, so a supplied summary can never influence a result.

How checked:
- vitest apps/web/src/lib/performance/timeToAnswerEvidence.test.ts -> 6 passed.
- The stage matrix, bucket coverage, fixture declaration and doesNotProve copy
  are asserted, so a stage cannot be added or re-bucketed silently.

Nothing measures anything yet and no production behaviour changes: this is the
schema, the fixtures' declaration and the promotion rule.
…capture (#335)

What:
- tests/perf/dockerFixtureTopology.mjs — one deterministic, secret-free Docker
  inventory generator: 25/100/250 containers plus the four scenario fixtures
  (provider-only revision change, Docker topology change, slow-but-bounded
  Compose projection, unavailable optional provider). Byte-identical for the
  same inputs, bounded by the published caps, and it never contacts Docker, the
  network, or the host filesystem.
- tests/perf/fake-docker-api.mjs — a fake Docker Engine API over a unix socket
  serving only the three read-only inventory endpoints the collector uses
  (/containers/json, /networks, /volumes, plus /_ping, /version, /info) from
  that generator, with a benchmark-only control route to advance the published
  topology generation.
- tests/perf/dockerFixtureTopology.test.mjs — determinism, exact counts, unique
  identities, contract bounds, secret-free labels, rejection of unsupported
  sizes/scenarios, topology-generation semantics, bounded Compose project, and
  the pinned fixture-name set.
- docs/testing/TIME_TO_ANSWER_EVIDENCE.md — the measurement authority: stage
  buckets, what each number does and does not prove, the controlled-run record
  rules, the fixture source, the promotion rule, and an explicit statement of
  what is done versus still outstanding in #335.
- package.json — npm run test:perf (node --test tests/perf/*.test.mjs), wired
  into check:js so the fixture source cannot rot silently.

Why:
250/100-container fixtures cannot be invented as real containers
reproducibly, and ordinary CI wall-clock time is not a gate. Pointing the daemon
at a deterministic fixture socket via the existing
DOCKERMAP_DOCKER_GATEWAY_SOCKET exercises the real collector, projection and
publication path while keeping the input byte-identical run to run.

How checked:
- Verified end to end against the real daemon build: with the fixture socket
  serving 25 containers the daemon reports mode=docker,
  dockerReachable=true, and publishes 25 containers / 1 network / 5 volumes at
  a model revision.
- node --test tests/perf/*.test.mjs -> 8 passed.
- npm run check (audit, typecheck, build, contracts, version/deployment/perf
  tests, JS tests, Rust fmt/clippy/tests) green.

No production behaviour changes: the fixture source and fixture daemon are
test-only.
What:
- crates/dockermap-daemon/src/bench_timing.rs — a test-only stage timer. It is
  inert unless DOCKERMAP_BENCH_STAGE_TIMING_PATH names an ABSOLUTE path; then it
  appends newline-delimited JSON records ({"stage":...,"ms":...}) to that file.
  A relative, empty or malformed value disables the hook. Write failures are
  dropped, never propagated into a publication.
- Three attribution points, all measurements of the CURRENT implementation:
  * dockerObservationMs — the Docker inventory read in
    collect_docker_snapshot_candidate;
  * composeEnrichmentMs — the Compose filesystem projection for the same
    publication;
  * findingsDerivationMs — the runtime-map findings projection in
    assign_revision.
- No route, no response field, no runtime telemetry, and no behaviour change:
  the Compose projection still executes inside the same Docker publication
  budget as before. #336 owns moving it off that path; nothing is decoupled
  here, and these numbers are what will let #336 prove the improvement.

Why:
#335 must measure the current critical path as it actually exists and show that
Compose projection currently sits inside it. Timing the two phases separately
requires a hook inside the daemon, which the issue permits only when strictly
needed to make a stage measurable.

How checked:
- 4 new unit tests: the sink requires an absolute non-empty path; stage lines are
  closed newline-delimited JSON; a DISABLED hook writes nothing and creates no
  file; an enabled hook appends exactly one parseable line per stage.
- cargo fmt + clippy --all-targets -D warnings clean; 207 daemon tests pass.
What:
- tests/perf/capture.ts — `npm run perf:time-to-answer`. The single documented
  command owns the whole test lifecycle: the deterministic fixture Docker
  daemon, the real daemon (bench attribution on), the real API, the production
  web build, the benchmark-only probe build, and real Chromium. It records raw
  samples only and writes the closed artifact; summaries are recomputed during
  review.
- tests/perf/browserProbe.js — test-only instrumentation loaded with
  `addInitScript({ path })`: it wraps the REAL EventSource (listening for the
  product's "snapshot" event) and observes DOM commits, so stages 6 and 7 have
  explicit timestamps. It also exposes the measurement helpers the harness calls
  by name through raw string expressions.
- tests/perf/probe/ + tests/perf/benchVite.config.mjs — a benchmark-only Vite
  build that imports the REAL `buildModel` and `layoutServices` production
  modules and measures them in real Chromium (stages 8 and 10). The ordinary
  production build never reads this config or includes this entry.
- tests/perf/staticServer.mjs — deterministic no-store static serving with SPA
  fallback and OS-reserved ports.
- tests/perf/emit-metadata.mjs — reads every pinned environment field from the
  runner itself (kernel, arch, cpu count, node, rustc, docker, chromium
  revision, font environment) instead of hand-authoring it.
- package.json — perf:time-to-answer and perf:metadata scripts.

Why:
#335 must measure the CURRENT implementation end to end. Stage 5 uses the real
Node/API path with its existing polling behaviour, so the baseline exposes
today's publication-observation floor rather than a synthetic watcher; the
poll interval is pinned in the environment. Lifecycle is boring on purpose:
private OS-reserved ports, process-group teardown (never pkill), explicit
readiness waits, fail-closed on any incomplete cell, raw samples preserved on
failure, and the live DockerMap deployment is never touched.

Notes:
- Stage guards come from the closed fixture x stage matrix, so a stage can only
  be recorded where the contract declares it.
- provider-only and unavailable-optional-provider fixtures need no trigger:
  their revision advance comes from provider state alone.
- Two esbuild/esbuild-keepNames traps are documented in the harness: a
  TS-authored init script and TS-authored page.evaluate functions both emit a
  `__name` helper that does not exist in the page realm, so all page code is
  either a plain .js file or a raw string expression.
What:
- Completes the capture command's measurement loops so every declared cell gets
  the contract's 3 controlled runs x 15 warmed samples: stages 1 and 2 restart
  the daemon per sample, and the revision-driven browser stages (5, 6, 7) loop
  over real published revision changes rather than being measured once. Cmd-K is
  measured from a fresh page per sample so each is a real open, not an already
  open palette.
- tests/perf/summarize.ts + perf:summarize — recomputes every summary from the
  stored raw samples, so interpretation numbers cannot drift from the evidence.
- tests/perf/productionIsolation.test.mjs — proves the ordinary production
  artifact carries no benchmark entry or probe identifier, no production source
  or build script reaches the benchmark, the daemon hook is inert by default and
  unreachable from any API surface, and the shipped bundle has no analytics.
- docs/testing/TIME_TO_ANSWER_BASELINE.md — the measured baseline and the
  evidence-backed interpretation.
- docs/testing/TIME_TO_ANSWER_EVIDENCE.md — the documented procedure and the
  slice's completed state.
- Fixes two harness defects found by running it for real: a run-shape bug that
  nested samples one level too deep, and raw samples not being preserved when
  artifact assembly failed. `record()` now fails at the measurement site, naming
  the cell.

Measured baseline 1 (median of three run p95, ms; source b6904d5, artifact
sha256 67f9b78b4e358c77d3980fbc8752dbcd33141e1898bb9112a9b9312224eaf9d2):

  25 / 100 / 250 containers
  daemon start -> listener            41.99 / 109.84 / 254.00
  listener -> first Docker model      12.71 /  59.90 / 165.76
  Docker observation                   7.26 /   9.10 /  13.07
  Compose projection                   1.15 /   1.18 /   1.25
  publication -> Node observation   1307.47 / 636.12 / 1623.47   (0.42 .. 1976)
  notification -> coherent model      12.70 /  20.60 /  34.60
  coherent -> useful render           12.70 /  20.60 /  34.60
  buildModel()                         0.50 /   0.90 /   1.50
  findings derivation                  0.07 /   0.20 /   0.92
  legacy topology layout               2.10 /  25.80 / 174.40
  Cmd-K open + query                  55.10 /  45.40 /  34.90
  production bundle load              49.30 /  59.30 /  50.20

Interpretation (no recommendation, no optimization): the fixed-interval
publication->Node observation floor is the largest single contributor (87/64/68%
of summed stage medians) and is a phase-of-poll distribution, not network
latency — #337 owns it. Cold start is dominated by process bring-up, not Docker
work, and Compose projection is a measured but small cost on these fixtures even
though it currently executes inside the Docker publication budget — #336 owns
decoupling and should be judged against these numbers rather than an assumed
large win. Legacy Home layout is 174 ms at 250 containers and scales
super-linearly — #338 owns that decision.

Why:
#335 is the measurement authority for the Speed epic. This lands the baseline
the later issues must prove themselves against, with the environment pinned and
the promotion gate RED-checked before any comparison is trusted.
…ne (#335)

Every accepted finding from both adversarial reviewers, except the artifact
itself, which is recaptured after this commit.

P1 — Cmd-K measured palette-open, not query-to-results:
- tests/perf/browserProbe.js commandQuery now snapshots the UNFILTERED command
  list first (the palette renders every command on open, so "an item exists"
  passed even with filtering entirely broken), types a query with a known
  expected result, and only stops the timer when the list has actually changed
  AND still contains that token. The token is the fixture-derived
  `fixture-service-0` when present, otherwise derived from the rendered list.

P1 — stage 5 phase-locking:
- the trigger is now jittered by a uniform sub-interval delay per sample, so the
  measurement describes the real poll-wait distribution instead of one fixed
  phase offset between the daemon's 2 s refresh loop and the API's 2 s poller.
  The contract and the doc state that de-correlation is the harness's doing and
  that the underlying mechanism still runs on two fixed cycles.

P2s:
- probe daemons write to their own sink, so the stages documented as warmed no
  longer mix cold-start first observations into the resident daemon's samples;
- the capture refuses a dirty worktree and refuses to run when the metadata's
  sourceRevision/harnessRevision do not match the checked-out commits, and the
  environment now records BOTH the product revision and the benchmark-harness
  revision, so a baseline is reproducible from a commit;
- findingsDerivationMs is bucketed as backend-collection (it runs in the daemon
  during publication, not in the browser) and its doc text says the fixture
  derives no findings, so it measures the empty-derivation path;
- stages 6 and 7 are no longer duplicates: stage 6 ends on the first commit that
  CHANGES RENDERED TEXT anywhere, stage 7 on a text-changing repaint of the Home
  content region only, and stage 7 is declared only for fixtures whose published
  change demonstrably repaints Home (reference-* and docker-topology-change);
- the environment records harnessRevision and dockerRevision is now
  informational for compatibility (no measured stage exercises the host Docker
  daemon), so an engine upgrade cannot fail an unrelated comparison.

P3s:
- ssePollIntervalMs is derived from the API's own source default and passed to
  the API explicitly, so the pin cannot drift from the interval that ran;
- scenario premises are asserted: the provider-only fixture fails if its fixture
  inventory changed during the run, and the unavailable-provider fixture fails if
  no optional provider was non-empty/fresh;
- --fixtures is gated behind DOCKERMAP_BENCH_DEBUG (a partial run can never
  satisfy the closed matrix, so it was a guaranteed-waste footgun);
- promotion RED-checks now pin the median-of-three aggregation (one slow run must
  pass, two slow runs must fail) and assert that a differing informational
  dockerRevision does NOT fail a comparison;
- the contract's stage/bucket documentation is corrected for stages 4-9.

Gates: vitest src/lib/performance 31/31.
…335)

Baseline 2, captured from committed revision 0714c87a (product and harness
revisions both recorded; artifact sha256
38e0c650121f81ec69c2dd1ddc18191c98c091836652bd8ecd7570fc0d006197, stored outside
the repo at /srv/jonas/evidence/dockermap/time-to-answer-baseline-2.json).

Every number is the median of three run p95 values recomputed from the raw
samples. Baseline 1 is documented as rejected and superseded, with the four
reasons it was invalidated.

Corrected findings versus the rejected baseline (25 / 100 / 250 containers,
median of three run p95, ms):

  daemon start -> listener            43.54 / 119.47 / 234.29
  listener -> first Docker model      13.51 /  54.77 / 156.83
  Docker observation                   7.91 /   9.36 /  15.34
  Compose projection                   1.08 /   1.06 /   1.08
  publication -> Node observation   1971.87 / 1828.38 / 1918.79   (spread 0.27 .. 1991)
  notification -> coherent model      15.20 /  19.30 /  52.10
  coherent -> useful render           15.20 /  19.30 /  52.10
  buildModel()                         0.60 /   0.70 /   1.80
  findings derivation                  0.07 /   0.30 /   0.91   (backend-collection)
  legacy topology layout               2.20 /  24.20 / 164.40
  Cmd-K open + query                  51.40 /  37.60 /  26.60
  production bundle load              53.10 /  55.60 /  61.50

Stage 5 now spans the interval instead of a 66 ms band: baseline 1's 1283-1349 ms
for reference-25 was a phase offset locked by the harness's own startup sequence,
not a sample of the poll wait. The de-correlation is the harness's doing and the
two fixed 2 s cycles are unchanged, which the contract and the doc both say.

Also corrected: 46 cells (not 21); stage 7 declared only for the four fixtures
whose published change demonstrably repaints Home; findings derivation described
as the daemon-side empty-derivation path; variance statements replaced with the
measured ranges; dockerRevision described as informational.

Gates: test:perf 13/13, vitest src/lib/performance 31/31, npm run check green.
…3, partial)

Round-3 review findings 1, 2 and 4, plus the contract's cold/warm classification.
Findings 3 (stage 6 vs 7 seam), 5 (documentation corrections) and the baseline-3
recapture are NOT in this commit — see the issue comment.

1. Cold-start contamination of warmed stages:
- the contract now classifies every stage explicitly
  (TIME_TO_ANSWER_STAGE_KIND: cold-start | warmed-repeated) and identifies
  scenario cells (isScenarioCell);
- new pure `splitWarmedObservations()` discards EXACTLY ONE warm-up observation
  from a warmed window and returns it for audit; it refuses a window shorter than
  count + 1 so no arbitrary slow sample can be dropped instead;
- the capture now waits for samples + 1 bench observations per warmed daemon
  stage and records the remainder, keeping the discarded warm-up in
  `warmUpObservations` for the raw audit trail. Cold-start stages (process start,
  first Docker publication) are untouched, because there the first observation IS
  the measurement;
- RED-checks: a 99.9 ms cold first observation beside 2 ms steady samples is
  proven to leave the recorded summary under 10 ms, and a short window throws.

2. Daemon binary provenance:
- new `assertDaemonBinaryProvenance()` requires two lowercase sha256 digests and
  fails on mismatch, so the executable that produced every daemon-side number is
  bound to the recorded revision. The benchmark does not claim bit-for-bit
  reproducible Rust builds across machines; it proves which binary THIS capture
  executed;
- emit-metadata now BUILDS the release daemon (`cargo build --release --locked -p
  dockermap-daemon`), then records `daemonBinarySha256`, `daemonBinaryBuild` and
  `cargoRevision`; the capture verifies the digest before the run and will verify
  it again after, failing closed on substitution or drift.

4. Cmd-K hardening: success now also requires the filtered item count to be
STRICTLY lower than the unfiltered count, so a reorder-only or no-op filter
cannot satisfy the assertion (the palette prepends an "Ask Copilot" item on any
query, which would otherwise make a reordered list pass).

Environment gains daemonBinarySha256, daemonBinaryBuild and cargoRevision; all
three test-suite environment fixtures updated.

Gates: vitest src/lib/performance 34/34.
…contract (#335)

Round-3 finding 5 (all accepted round-2 documentation/evidence issues except the
stage-6/7 definitions, which follow once the benchmark seam lands).

- matrix count corrected to 44 cells in both documents;
- the bucket table no longer lists findingsDerivationMs under browser-model (the
  contract buckets it backend-collection; it runs in the daemon);
- the promotion gate sentence now states that compatibility excludes
  sourceRevision (by design) AND dockerRevision (recorded but informational, since
  no measured stage exercises the host Docker daemon), and that the other 15
  fields must match;
- new sections defining cold-start vs warmed-repeated vs scenario cells, and the
  exactly-one-discarded-warm-up rule with the warm-up retained separately in the
  raw audit trail and excluded from p95;
- new section on capture discipline: sourceRevision vs harnessRevision meaning,
  and the requirement to check out the revision the artifact names;
- new section on daemon binary provenance: the canonical locked release build,
  daemonBinarySha256/daemonBinaryBuild/cargoRevision, digest verified before and
  after the capture, no bit-for-bit reproducibility claim;
- new section stating that the 2000 ms SSE polling mechanism is production
  behaviour that is deliberately unchanged, that the jitter de-correlates only the
  measurement, and that the stage must never be described as network latency;
- baseline 1 and baseline 2 are labelled REJECTED historical attempts with the
  reasons, and the baseline document no longer presents either as current
  authority.

No code changes. Gates unchanged (vitest src/lib/performance 34/34 at c622da2).
What:
- Stage 6 (notificationToCoherentModelMs) now ends at the REAL application seam:
  useSystemModel records one opaque acceptance event (timestamp + model revision
  token) at the instant a fetched snapshot/runtime pair becomes the coherent model
  the UI renders. It is application state, never derived from a DOM mutation.
- Stage 7 (coherentModelToUsefulRenderMs) starts at that stage-6 timestamp and
  ends when the accepted revision's EXPECTED Home content is present - in a commit
  the application stamped with that accepted revision - plus exactly one bounded
  requestAnimationFrame. No sleeps.
- The seam is real product source (apps/web/src/lib/performance/modelAcceptance.tsx)
  gated by the compile-time flag __DOCKERMAP_BENCH_ACCEPTANCE__: the production
  Vite config defines it "false" (dead-code eliminated), and a benchmark-mode
  application build (tests/perf/benchAppVite.config.mjs) defines it "true". The
  application source is built twice, never copied.
- The capture serves stages 6/7 from the benchmark-mode build and stages 11/12
  (Cmd-K, production bundle) from the ordinary production build, and refuses to run
  unless the benchmark-mode artifact carries the seam and the production artifact
  does not (assertBuildIsolation).
- New stage-6/7 independence control: 3 control samples per stage-6/7 fixture with
  a 250 ms artificial presentation delay injected AFTER acceptance. Stage 6 must
  not move beyond max(30 ms, 25%); stage 7 must absorb at least 70% of the delay.
  Enforced before any artifact is assembled; raw pairs and a per-sample audit
  trail are written beside the artifact as <output>.stage-seam.json.
- The fixture generation delta is now product-visible (generation g stops the
  first g containers), so the stage-7 expected-content check discriminates a fresh
  render from a stale one; expectedExitedCount derives the expectation from the
  same generator the fixture daemon serves.
- Chain of custody is checked per sample: the revision the API observed through the
  live SSE path must be the revision the browser accepted and rendered.
- Docs: stage 6/7 definitions, the two-clock distinction, benchmark-build
  isolation, the independence control, the generation delta, the 44-cell matrix,
  and the per-build stage mapping.

Why:
Baseline 2 was rejected because its stage 7 was element-for-element identical to
stage 6 in all 180 samples: both stages came from one DOM-derived clock. This
change makes stage 6 the application's acceptance instant and stage 7 the
presentation of that accepted revision, and adds the control that proves the two
clocks are independent.

How checked:
- npm run test:perf -> 17/17 (production isolation incl. seam confinement and the
  flag-is-false check, fixture generation delta, independence RED cases)
- npm run test:web -> 554/554 (incl. new timeToAnswerIndependence.test.ts)
- production bundle inspected after `npm run build --workspace @dockermap/web`:
  0 occurrences of every seam identifier; benchmark-mode build contains the sink
- npm run typecheck --workspace @dockermap/web passes

Remaining risk / follow-up:
No baseline may be captured until the round-3 smoke demonstrates the independence
control, the warm-up discard and the artifact assembly on the committed revision.
What: emit-metadata.mjs builds the release daemon as
`cargo build --release --locked -p dockermap-daemon --manifest-path crates/Cargo.toml`,
and the recorded daemonBinaryBuild slug and the evidence doc name the same command.

Why: the cargo workspace lives under crates/, so the previous command failed with
"could not find Cargo.toml" and no capture could emit metadata at all - the pinned
binary provenance step was unrunnable.

How checked: npm run perf:metadata now completes, builds the release daemon and
records daemonBinarySha256.
What:
- The probe arms on the monotonic acceptance SEQUENCE and waits for the next
  accepted revision, instead of "any revision other than the last one seen". The
  previous predicate selected a stale acceptance event whenever a publication
  landed between arming and the trigger.
- Stage 6 starts at the notification of the ACCEPTED revision, recovered from a
  per-revision notification log, instead of assuming the latest notification
  belongs to it.
- The API observation keeps listening for 200 ms after the first change and keeps
  every revision it saw for the sample; the browser's accepted revision must be one
  of them (a single fixture change can publish more than once).
- preserveRaw creates its destination directory, so raw samples survive even when
  the raw directory does not exist yet.

Why: the round-3 smoke failed with "the accepted revision ...-15 is not the
revision the browser was notified of (...-16)" - the harness's own arming logic
mis-attributed a revision, not the seam.

How checked: round-3 smoke on reference-25 (1 run x 3 samples) - see the run log;
the independence control, warm-up discard and Cmd-K checks are exercised there.
What:
- Stage 5's API observation is now ARMED before the trigger and stopped after the
  sample's browser measurement resolves, instead of running as a one-shot call
  after the trigger with a fixed post-collection window. Every revision belonging
  to the sample is captured, however many publications it produces.
- Stages 6/7 select the (accepted model, rendered content) PAIR that carries the
  sample's expected Home content and that revision's stamp, instead of assuming
  the first acceptance after arming owns the change. A provider-state-only
  publication between the trigger and the inventory publication no longer
  mis-attributes the sample.
- The audit trail records how many acceptances were skipped and which revisions
  the API observed for the sample.

Why: the second round-3 smoke failed with "the Home content for the accepted
revision ...-21 never rendered (expected Offline=4)" - the daemon published an
intermediate revision (-21) with unchanged Home metrics before the inventory
revision (-22), and the harness assumed one acceptance per sample.

How checked: round-3 smoke on reference-25 (1 run x 3 samples) plus the stage-6/7
independence control; see the run log.
…ision (#335)

What: PublicationObserver.stop() now waits (bounded by one poll interval plus a
1.5 s margin) until the harness's own SSE connection has emitted the revision the
browser accepted, and the sample loop passes that revision in.

Why: each SSE connection runs its OWN poll timer (apps/api/src/index.ts holds a
per-connection interval and a per-connection lastEmittedRevision), so the
harness's stream legitimately lags the browser's by up to one interval. The chain
check failed with "the browser accepted revision ...-21, which the API never
observed for this sample (observed: ...-20)" even though both streams were fed by
the same real mechanism. The recorded stage-5 duration is fixed at the first
emission and is unaffected by the extra wait.

How checked: round-3 smoke on reference-25 (1 run x 3 samples).
What: every capture now writes <output>.harness-evidence.json beside the artifact,
carrying (a) the stage-6/7 independence control (verdict, per-run sample sets and
the per-sample acceptance audit) and (b) warm-up retention: for every daemon-side
warmed cell the COMPLETE samples+1 observation window in the order the daemon
produced it, with the discarded warm-up at index 0, plus proof that the recorded
samples equal that window minus the warm-up.

Why: splitWarmedObservations discarded the cold first observation and kept the
value only in memory, so no reviewer could check that a slow warm-up value was not
silently dropped. The capture now refuses to emit an artifact when a window is
missing, shorter than samples+1, discards anything other than the first
observation, or does not match the stored run.

How checked: round-3 smoke on reference-25; the retention checks run on every
capture and the evidence file is written before the artifact is validated.
…#335)

What: the probe now arms in one of two modes, chosen by the closed matrix:
- content mode (every fixture that also declares stage 7): the sample ends on the
  (accepted model, rendered content) pair carrying the expected Home content for
  the triggered change;
- acceptance-only mode (provider-only-revision-change, unavailable-optional-provider):
  stage 6 ends at the acceptance instant, because those fixtures publish a
  provider-state revision with NO inventory change, so no Home repaint exists to
  wait for and stage 7 is not declared for them.

Why: the full baseline-3 attempt aborted on provider-only-revision-change with
"the Home content for the triggered change never rendered (expected Offline=1;
story=... storyValue 0 ...)" - the content-mode end condition can never fire on a
fixture whose inventory does not change. The three reference fixtures had already
completed at full scale, so only those two cells were affected.

How checked: round-3 smoke over reference-25 (content mode) and
provider-only-revision-change (acceptance-only mode).
… model (#335)

What: the harness now records which revision each paired /api/snapshot and
/api/runtime/map fetch actually delivered, and stage 6 starts at the browser
notification that preceded that fetch cycle. The requirement that the accepted
revision equal a revision the harness's OWN SSE stream announced is replaced by
the browser-side provenance chain, and the stream overlap is recorded as
acceptedRevisionInApiStream evidence instead of gating.

Why: the second full baseline-3 attempt aborted on unavailable-optional-provider
with "the accepted revision ...-32 was never notified to this browser". The daemon
is read per request: /daemon/health (what the SSE stream carries) and
/daemon/snapshot (what the app accepts) are separate reads and legitimately hold
different revisions while provider state churns, so the previous token-equality
proxy was wrong, not the seam. Stage 6 is now anchored to the fetch the model came
from, which is stronger evidence than the proxy it replaces.

How checked: round-3 smoke over reference-25 (content mode), docker-topology-change,
provider-only-revision-change and unavailable-optional-provider (acceptance-only).
What: the three static servers now bind an OS-assigned port (port 0) and report the
port actually bound, and every spawned child that binds a reserved port is started
through a bounded retry (fresh port for the daemon and the per-sample startup
probes; the SAME port for the API, whose port is baked into the production build).

Why: the third full baseline-3 attempt aborted with
'listen EADDRINUSE: address already in use 127.0.0.1:37501'. reservePort() binds,
closes and hands the number back to the pool, so the port can be taken before the
child binds it - and the API port in particular stayed reserved-but-unbound for
minutes between fixtures. An unrelated process or a not-yet-reaped child could
therefore abort an expensive capture.

How checked: round-3 smoke over the four fixtures that exercise both stage-6 modes.
…335)

What: docs/testing/TIME_TO_ANSWER_BASELINE.md is rewritten as the baseline-3
record: the pinned environment, the verbatim recomputed 44-cell table, the stage-5
distribution, stage 6/7 medians, the enforced stage-6/7 independence control with
its per-fixture verdicts, the acceptance audit, bucket shares, top contributors,
the Compose contribution and the auditable warm-up retention. Baselines 1 and 2 are
kept as rejected history WITHOUT their numbers, so they cannot be read as current
authority. TIME_TO_ANSWER_EVIDENCE.md now states that baseline 3 exists and that the
round-3 independent review is the gate before it becomes the authority.

Why: baseline 3 was captured from committed revision cf77e8b (25.6 min, 44 cells x
3 runs x 15 samples = 1980 raw samples, exit 0) and the numbers had to be recorded
from the recompute path rather than from memory.

How checked: every table is the verbatim output of
; the artifact, the harness
evidence and the recomputed summary are stored in
/srv/jonas/evidence/dockermap/time-to-answer/ with recorded sha256 digests.
…-5 phase sweep (#335)

What (methodology revision 2, dockermap-v1/time-to-answer-methodology-2):

- Stage 5 now DRIVES the publication phase instead of hoping for one. The poll
  interval is divided into 15 declared phases (one per recorded sample per run, each
  at the centre of its division), every run sweeps them ascending, and the harness
  controls the phase by choosing when it connects its observation stream relative to
  a predicted daemon publication. Each sample records its declared and observed
  phase, and asserts its own control: observed publication matches the prediction,
  observed latency lands on the declared phase within 60 ms, and the observation
  arrived through the real API poller path.
- New validity guards (assertPollPhaseSweep): every declared phase represented at
  least twice, no uncontrolled phase, no non-poller observation, span >= half the
  interval, and the earliest declared phase at least half an interval slower than
  the latest. A narrow-band sweep is REJECTED - it is a RED test, not an aspiration.
- Stage 5 reports a phase curve plus a PHASE-NORMALIZED p95 computed from the
  predeclared uniform grid, explicitly documented as a statement about the mechanism
  under uniformly sampled phase offsets - not an observed user-traffic distribution
  and not network latency.
- Fixed warm-up protocol: five predeclared warm-up observations before the 15
  measured samples, chosen from the round-3 windows (the second observation reached
  2.15x the window median in 5/30 windows; index 5 onward stayed within 1.29x).
  Warm-ups are never dropped adaptively; a declared stationarity check (0.5x-1.5x
  band on the final two warm-ups against the measured median) INVALIDATES a window
  instead of trimming it. Every warm-up observation is retained.
- Provenance is no longer compatibility: daemonBinarySha256, sourceRevision and
  harnessRevision are recorded but not required to match, so a candidate that
  changes crates/ (#336) is comparable; methodologyVersion, the toolchain, browser,
  fonts, fixture revision, build mode and polling configuration still must match.
- The documented before-and-after daemon binary verification now exists: the digest
  is re-hashed after the run and the capture fails closed on a mismatch, with both
  digests recorded in the harness evidence.
- The slow-Compose scenario premise is asserted (the project really declares its 400
  services) instead of merely named, and the capture asserts the application page
  never contacted the daemon directly.
- Docs corrected: stage 7 is a bounded render/presentation confirmation rather than
  "exactly one frame", the promotion section states the provenance/compatibility
  split with the exact field counts, the rejected-capture paths are right, and the
  stage-5 sections describe the sweep and the normalized figure.

Why: round-3 review rejected baseline 3. Its stage-5 samples were a drift-locked
sawtooth (reference-25 covered 6.0% of the 2000 ms interval with consecutive
deltas under 19 ms) even though the harness slept a uniform random delay, because
both the daemon refresh loop and the API poller are fixed 2 s loops; its single
discarded warm-up still left a 2.15x-of-median observation inside the measured
window; it gated promotion on a rebuilt binary digest; and it documented a post-run
verification the code did not perform.

How checked: npm run test:perf 20/20 (including the new methodology drift guard),
vitest performance suites 56/56 (including the narrow-band, uncontrolled-phase,
missing-phase, non-poller-path and stationarity RED cases), web typecheck clean.
…335)

The declared minimum of three samples per phase needs the three controlled runs the
full protocol requires, so a debug probe (which can never emit an artifact) would
fail the sweep guard for a reason that has nothing to do with the measurement. The
minimum is now an explicit guard parameter, and the capture passes 1 only when the
declared run count is below the protocol's; every other guard still applies.
…om daemon start (#335)

The methodology smoke failed its first stage-5 sample: declared phase 66.7 ms
produced 2051.5 ms against an intended 1933.3 ms (118.2 ms error, tolerance 60 ms).

Cause: the predicted publication instant came from a single observed gap, which
carries one cycle's refresh work plus the tracker's own start transient. Stage 5's
deepest declared phases put the publication just after a poll tick, so a period
error maps one-for-one into phase error, and the error was large enough to land the
sample in a neighbouring phase bucket.

Fix: the tracker now runs from daemon start (so the grid is known before the browser
stages begin), requires four publications before the first sample is placed, and
estimates the period by least-squares over the last eight publication instants
instead of one gap. The control tolerance stays at 60 ms - below half the 133 ms
grid step, so a sample cannot bucket into a neighbouring phase.
…nge (#335)

The methodology smoke failed again at the deepest declared phase (66.7 ms declared,
1933.3 ms intended, 2456.3 ms observed). Cause: the daemon can publish more than one
revision per refresh cycle (the inventory publication followed by a provider-state
publication), and the tracker fitted its grid over every revision change, so the
extra intra-cycle publications skewed the period estimate - which the deepest phases
amplify one-for-one into phase error.

The tracker now collapses revisions into refresh-cycle leaders (changes within a
quarter-interval belong to one cycle), fits the period over the last eight cycle
leaders, and each sample attributes itself to the detected publication nearest the
predicted cycle. Failure messages now carry the connect, prediction, publication and
tick instants so a control failure is diagnosable without a re-run.
…ctually carried (#335)

The API's poll tick emits whatever revision is current at tick time, so the
publication a sample measured is the LAST revision change at or before the tick.
The previous attribution (nearest to a predicted cycle instant) could pick a
publication the tick never carried: the smoke showed a sample attributed to a
publication 510 ms before the prediction and 444 ms before the connection.

Also: cycle boundaries are now identified as publications at least 60% of the poll
interval apart, because one refresh cycle can publish a provider-state revision a
few hundred milliseconds after the Docker snapshot, and folding those extras into
the grid skewed the period estimate. Each sample records whether it was attributed
to a cycle boundary.
)

The methodology smoke kept missing the deepest declared phase (66.7 ms declared,
1933.3 ms intended, 453-2456 ms observed): the grid was fitted over revision changes,
but one refresh cycle publishes the Docker snapshot and can publish a provider-state
revision up to ~1.5 s later, so the fitted period carried those extras and the
prediction landed hundreds of milliseconds off - which the deepest phases amplify
one-for-one into phase error.

The tracker now uses the daemon's own snapshot timestamp: /daemon/health carries
lastUpdated, which the daemon stamps once per refresh cycle in epoch milliseconds, so
the grid is one exact instant per cycle with no detection lag (the stamp is aligned
once to the harness's monotonic clock). Attribution stays causal - the tick carries
the last cycle boundary at or before it - and each sample records whether the tick
also carried a newer intra-cycle revision.
Two fixtures cannot be phase-driven at all: provider-only-revision-change and
unavailable-optional-provider advance their revisions from the daemon's own host
provider collection, and the harness has no input that makes the daemon publish at a
chosen instant. The smoke showed the consequence - waiting for a spontaneous revision
landed the observation on the fourth poll tick, 6.1 s after the predicted cycle.

Rather than claim a phase those cells never drove, the design now declares which
fixtures are phase-controlled (POLL_PHASE_CONTROLLED_FIXTURES: the three reference
fixtures and docker-topology-change) and which are free-running:

- phase-controlled cells keep the declared grid, the per-sample control check and the
  full sweep guards (coverage, span, direction, poller path);
- free-running cells record the phase they ACHIEVED (declared phase = the bucket the
  observation landed in), pass a separate guard that requires real poller observations
  and a minimum sample count, and derive NO phase-normalized figure;
- the two guards refuse each other's samples, so a cell cannot claim a sweep it did
  not drive.
…ndow (#335)

Baseline 4's capture aborted on its third fixture: declared phase 200.0 ms produced
1877.9 ms, with the observation arriving 3.5 ms after the connection and attributed
to a boundary 2.1 s earlier.

Cause: the API emits the current revision immediately on connect, and
previousRevision came from the tracker's last poll, which can be up to 10 ms stale.
A revision published in that window made the very first frame look like a new
publication, so the sample was attributed to the PREVIOUS cycle's boundary.

Fix, in two parts: the pre-sample revision is now read fresh from the daemon health
endpoint right before the stream opens, and a sample only accepts a frame that could
be its own observation - for a controlled sample, at or after the predicted cycle
(safe because its poll tick cannot land earlier while the connect frame precedes it by
the declared phase); for a free-running sample, a margin after the connect. Frames
ignored this way are counted in the sample's evidence.
Provenance: Terra issue-runner; controlled Stage-6/7 protocol implementation and static integration coverage.
Provenance: Terra issue-runner; fixes reviewer-identified output and fixture-capacity contract.
Provenance: Terra follow-up worker; withhold Stage-5 publication until the exact accepted model pair is armed.
Provenance: follow-up worker checkpoint after 7722a79.

Validation: node --test tests/perf/productionIsolation.test.mjs
…evidence (#335)

What: publish the Baseline-4 authority and reconcile the methodology-8 benchmark runbook.

Why: document the composite evidence workflow, controlled Stage-5 capture, and supporting seam-isolation control.

How checked: reviewed the seven-step runbook, searched stale methodology terms, verified performance-code provenance, and ran git diff --check.
What: correct the Stage 6/7 seam-isolation protocol documentation.\n\nWhy: the control is a dedicated protocol with its own evidence rather than general-capture harness evidence.\n\nHow checked: inspected the scoped section, remaining harness-evidence references, working-tree scope, and performance-code baseline commit.
@Joncallim Joncallim changed the title Establish end-to-end time-to-answer baselines and regression gates (partial) (#335) Establish end-to-end time-to-answer baselines and regression gates (#335) Sep 27, 2026
@Joncallim
Joncallim marked this pull request as ready for review September 27, 2026 03:54
@Joncallim
Joncallim merged commit aa023a7 into main Sep 27, 2026
11 checks passed
@Joncallim
Joncallim deleted the dm/335-time-to-answer-baseline branch September 27, 2026 03:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant