Skip to content

perf: make the Baseline-4 promotion comparison runnable - #348

Merged
Joncallim merged 2 commits into
mainfrom
dm/perf-promotion-entrypoint
Sep 27, 2026
Merged

Joncallim merged 2 commits into
mainfrom
dm/perf-promotion-entrypoint

Conversation

@Joncallim

Copy link
Copy Markdown
Owner

Make the Baseline-4 promotion comparison runnable

Standalone tooling extraction. No product behaviour change — the diff is one new read-only CLI,
its test, the package.json script wiring, one tsconfig include entry, and the corresponding
step-7/8 correction in the evidence document.

The defect

The merged Baseline-4 authority (#335, squash aa023a7) could not actually be applied:

  • assertTimeToAnswerPromotion documents itself as “The benchmark job calls this after reading
    two closed JSON artifacts”
    , but no such job exists — its only callers are unit tests;
  • perf:time-to-answer --baseline deliberately refuses with “promotion requires the assembled
    composite evidence, not the incomplete end-to-end section”
    , because the capture can only
    produce the end-to-end section;
  • so step 7 of “Running the benchmark” in docs/testing/TIME_TO_ANSWER_EVIDENCE.md documented a
    comparison command that cannot run.

Every performance child of epic #333 (#336, #337, #338) must compare its candidate against
the pinned Baseline 4 under max(baseline × 1.25, baseline + 2 ms). Until this PR there was no
runnable way to do that.

What this adds

npm run perf:promote -- --baseline <baseline composite.json> --candidate <candidate composite.json>

tests/perf/promote.ts is a read-only comparator that:

  • loads and validates both artifacts through the existing contract
    (validateTimeToAnswerEvidence), calls the existing assertTimeToAnswerPromotion, and reuses
    the existing derivedTimeToAnswerSummaries / timeToAnswerLimit — it implements no measurement
    logic of its own and changes no rule;
  • prints a per-row table (fixture | stage | protocol | baseline reviewed | limit | candidate reviewed | delta | verdict), then the list of rows slower than baseline, then
    PROMOTION: PASS / PROMOTION: FAIL:<reason>;
  • never creates, modifies or moves an artifact, and fails closed on a missing input, an unknown
    flag, a duplicate flag or a malformed artifact.

tests/perf/promote.test.mjs spawns the CLI against tiny synthetic composites and asserts the
passing pair exits 0, the over-limit pair exits non-zero naming the offending key, and an
incompatible-environment pair fails with the environment message.

The documentation now describes the real sequence: capture the end-to-end section, capture the
dedicated Stage-5 section, assemble a candidate composite, then compare composites with
perf:promote.

Provenance

Extracted, without modification, from the #336 work on branch dm/336-docker-first-answer
(draft PR #347) where it was implemented and reviewed; tests/perf/promote.ts and
promote.test.mjs are byte-identical to that branch's versions (verified with git show … | diff -). The commit is registered under the #336 managed run.

Verification

  • npm run test:perf — 43 tests pass, including the new spawn-based comparator tests.
  • npm run typecheck — pass (tests/perf/tsconfig.json now includes promote.ts).
  • npm run check — exit 0 on df0e12e (audit, typecheck, contracts, builds, deployment/version/perf suites, JS tests, Rust fmt check + clippy + tests).
  • CI on this PR covers the rest.

Not included

No #336 product behaviour: the listener-first startup, concurrent Docker reads, Docker-first
publication, Compose enrichment and churn-reduction commits all remain on dm/336-docker-first-answer
and are untouched here. Every sentence in the evidence document about composeEnrichmentMs, the
Compose publication budget and #336 is left exactly as it is on main.

Refs #333, #335, #336.

What: add the perf:promote command and fail-closed composite comparison CLI, with coverage and evidence instructions.\n\nWhy: make reviewed Baseline-4 promotion comparisons directly runnable from composite evidence artifacts.\n\nHow checked: npm run test:perf; npm run typecheck; git diff --check.
What: document checkpoints and complete candidate capture steps.\n\nWhy: make the performance evidence sequence executable from capture through promotion.\n\nHow checked: inspected the updated command block and ran git diff --check.
@Joncallim
Joncallim merged commit cd70425 into main Sep 27, 2026
11 checks passed
@Joncallim
Joncallim deleted the dm/perf-promotion-entrypoint branch September 27, 2026 10:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant