Skip to content

#336: first Docker answer independent of Compose and cold-start work (parked — Baseline-4 promotion FAILED) - #347

Draft
Joncallim wants to merge 10 commits into
mainfrom
dm/336-docker-first-answer
Draft

Joncallim wants to merge 10 commits into
mainfrom
dm/336-docker-first-answer

Conversation

@Joncallim

Copy link
Copy Markdown
Owner

#336 — make the first Docker answer independent of Compose and cold-start work

Child of epic #333. This PR is parked as a draft: the change is implemented and deterministically
green, but the pinned Baseline-4 promotion rule FAILS it. Details below — do not merge without a
decision on the stage-boundary tension described at the end.

What changed

slice commit change
1 ae7a2e0 the daemon binds its HTTP listener before the first collection; PublicationReadiness gates every cache-backed route with a redacted 503 while Initializing
2 9f999b1 container/network/volume inventory is issued concurrently under one observation deadline (barrier-proven against the fake gateway)
3 56e91c9 the Docker base candidate (snapshot + runtime map + Docker-only findings) is published before any Compose filesystem work starts
4 550c908 Compose enrichment runs detached and single-flight, keyed to the opaque DockerObservationRevision plus source_generation; a result that no longer matches the current Docker observation is discarded, never attached
5 6a0467e harness/docs: capture listener readiness separated from Docker-model readiness; the promotion comparison is now runnable; docs describe the asynchronous enrichment path
6 92d7cbf listener readiness proven by a real HTTP probe (tests/perf/listenerReadiness.mjs + four behavioural tests)
7 562d4f4 candidate-assembly example corrected in the evidence document

The Compose mount-drift rule (compose.declared_mount_missing_at_bound_container) is now derived
only after a coherent Compose binding exists; the Docker-only #69 rules keep exactly their previous
truth semantics; gateway authority, route allowlist and provider cadence are untouched.

A defect found while reconciling this issue (fixed here)

The merged Baseline-4 authority could not actually be applied. assertTimeToAnswerPromotion
documented itself as "the benchmark job calls this after reading two closed JSON artifacts", but no
such job existed — perf:time-to-answer --baseline deliberately refuses (it only produces the
end-to-end section), while docs/testing/TIME_TO_ANSWER_EVIDENCE.md step 7 documented a comparison
command that could not run. Every performance child of #333 needs this comparison, so this PR adds
the missing read-only entrypoint:

npm run perf:promote -- --baseline <baseline composite> --candidate <candidate composite>

It reuses the existing contract, limit (max(baseline × 1.25, baseline + 2 ms)) and summary math,
prints the per-row comparison plus the regression list, and fails closed on a violated limit, an
incompatible environment or a malformed artifact (tests/perf/promote.ts,
tests/perf/promote.test.mjs).

Measurement against Baseline 4 — PROMOTION FAILED

Methodology dockermap-v1/time-to-answer-methodology-8; baseline pinned to
bdce6ae354d757d3318c514e10d620edb497918e; candidate captured at
562d4f46657047c414972e73ecbf03bfd8979cbd in the same pinned environment; fixed 60 burn-in + 15
measured observations, three controlled runs per cell, full declared fixture matrix, 82.4 min.
Preconditioning PASS and Stage-5 phase-control PASS preceded the capture; the daemon binary digest
was verified identical before and after. Raw samples, gates, the promotion table and the chain log
are under /srv/jonas/evidence/dockermap/time-to-answer/562d4f46/.

PROMOTION: FAIL: candidate exceeds the reviewed promotion limit for
reference-100 / listenerToFirstDockerModelMs, notificationToCoherentModelMs, commandQueryMs
fixture stage baseline reviewed limit candidate delta verdict
reference-100 listenerToFirstDockerModelMs 19.11 23.88 43.92 +24.81 FAIL
reference-100 notificationToCoherentModelMs 37.20 46.50 48.30 +11.10 FAIL
reference-100 commandQueryMs 22.80 28.50 37.40 +14.60 FAIL

Improvements the same run records (all within limit):

fixture stage baseline candidate delta
reference-100 daemonStartToListenerMs 118.14 48.10 −70.04
reference-250 daemonStartToListenerMs 311.26 77.46 −233.80
reference-250 listenerToFirstDockerModelMs 123.58 72.82 −50.77
reference-25/100/250 dockerObservationMs 1.74 / 3.24 / 7.26 1.59 / 2.59 / 6.37 −0.15 / −0.65 / −0.89

Stage-boundary tension (the reason this is parked). The change deliberately moves the
listener/publication boundary, so the two adjacent stages are re-partitioned, not both made slower:

fixture start → listener + listener → first model (baseline) same sum (candidate)
reference-25 38.22 + 11.50 = 49.72 41.25 + 11.35 = 52.60
reference-100 118.14 + 19.11 = 137.25 48.10 + 43.92 = 92.02
reference-250 311.26 + 123.58 = 434.84 77.46 + 72.82 = 150.28

Baseline 4 measured daemonStartToListenerMs against a daemon that bound its listener after the
first collection, so it folded the first collection into that stage; the candidate measures the true
listener boundary and the remainder lands in listenerToFirstDockerModelMs. Which of those two rows
promotes therefore depends on a stage definition this issue is explicitly allowed to change — that is
a measurement-authority question, not something this PR may settle by reinterpreting the rule. The
other two failures are on web-side rows at the same fixture and are consistent across all three runs
(the baseline's own commandQueryMs runs were 18.20 / 40.10 / 22.80 — its median sits well below its
own spread). The candidate capture also records 112 out-of-phase accepted-revision notes against 10
in the baseline capture, i.e. the browser now observes more intermediate revisions.

Verification

  • npm run check (audit, typecheck, contracts, builds, deployment/version/perf suites, JS tests,
    Rust fmt check + clippy + tests) — exit 0 at 562d4f4.
  • npm run fmt:rust:check, npm run test:rust:daemon (209 passed), npm run test:perf (48 tests)
    — green.
  • npm run test:live-docker — PASS (1 passed, 33.3 s) against the host Docker daemon. The first
    attempt failed with web=exited(1),signals=permission_denied: the root-run worker sessions had
    left apps/web/node_modules/.vite root-owned, so the web server could not start as the repo
    owner. Ownership restored (no permission widened); the daemon and API were healthy in both
    attempts.
  • Production-image E2E runs in CI on this branch.
  • perf:preconditioning — PASS (12 sequential fresh-browser runs). perf:phase-control — PASS
    across all six declared fixtures (grid span 1794.2–1800.4 ms, tolerance 90 ms).
  • Benchmark: 82.4 min, full 44-row matrix, 3 controlled runs × 15 measured observations per cell
    after the fixed 60-observation burn-in, daemon binary digest verified identical before/after.

Not changed

Provider scheduling cadence, gateway authority/route allowlist, generated contracts, the Baseline-4
methodology and artifacts, and the disclosed Stage-6/7 seam-isolation supporting-control limitation
(neither reopened nor claimed to be changed).

Refs #336 — not closed by this PR.

What: bind the daemon listener before starting refresh and gate cache-backed routes until the first collection attempt completes.\n\nWhy: bootstrap mock data must never be mistaken for an authoritative model, while listener availability must not wait on collection.\n\nHow checked: cargo test -p dockermap-daemon --manifest-path crates/Cargo.toml initializing_cache_is_not_published_even_when_mock_is_allowed
What: issue the fixed container, network, and volume inventory reads concurrently and cover their synchronized gateway arrival.\n\nWhy: independent Docker inventory latency must not extend the authoritative observation path.\n\nHow checked: npm run fmt:rust; cargo test -p dockermap-daemon --manifest-path crates/Cargo.toml snapshot_timeout_invalidates_stalled_client_and_fresh_client_recovers.
…336)

What: publish base Docker candidates without awaiting Compose filesystem projection.\n\nWhy: authoritative Docker topology must be available before optional Compose enrichment.\n\nHow checked: cargo test -p dockermap-daemon --manifest-path crates/Cargo.toml repeated_projection_timeouts_are_single_flight_and_keep_the_docker_client.
What: carry keyed Compose work after Docker publication, attach only to the matching Docker source observation, and include private binding identity in revisioning.\n\nWhy: late or timed-out filesystem work must never relabel newer Docker topology.\n\nHow checked: cargo test -p dockermap-daemon --manifest-path crates/Cargo.toml repeated_projection_timeouts_are_single_flight_and_keep_the_docker_client.
What: add the read-only composite promotion CLI, listener-boundary probe, focused spawn tests, and documentation.\n\nWhy: make the pinned Baseline-4 authority runnable while keeping Docker-model readiness distinct from listener readiness.\n\nHow checked: npm run test:perf; npm run typecheck; npm run fmt:rust:check.
What: extract listener readiness into a reusable HTTP probe and cover successful, refused, and timed-out listener responses with real local servers.\n\nWhy: an HTTP response, including a truthful 503, proves listener availability while Docker-model readiness remains a later boundary.\n\nHow checked: npm run test:perf; npm run typecheck.
What: Removed the duplicate candidate --output flag from the Step 7 command.\n\nWhy: The example now shows the valid general → stageFive → output argument order.\n\nHow checked: Verified the Step 7 block, one-line documentation diff, whitespace check, and working-tree scope.
What: retain a coherent Compose projection for an unchanged Docker observation and compare normalized public mount findings before republishing enrichment.

Why: avoid transient Docker-only revisions and equivalent Compose confirmations triggering duplicate browser model work.

How checked: npm run test:rust:daemon
What: cover retained Compose findings for unchanged Docker confirmation and one coherent revision for changed mount projections.

Why: pin the no-churn behavior while ensuring genuine Compose drift remains observable.

How checked: npm run fmt:rust; npm run test:rust:daemon
What: align the shared promotion scripts and perf TypeScript configuration with main, and make the candidate evidence sequence runnable.\n\nWhy: prevent formatting-only divergence after the promotion-tooling PR merged while retaining #336-specific Compose evidence prose.\n\nHow checked: npm run test:perf; npm run typecheck; origin/main comparisons for package.json and tests/perf/tsconfig.json; git diff --check.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant