perf: make the Baseline-4 promotion comparison runnable - #348
Merged
Merged
Conversation
What: add the perf:promote command and fail-closed composite comparison CLI, with coverage and evidence instructions.\n\nWhy: make reviewed Baseline-4 promotion comparisons directly runnable from composite evidence artifacts.\n\nHow checked: npm run test:perf; npm run typecheck; git diff --check.
What: document checkpoints and complete candidate capture steps.\n\nWhy: make the performance evidence sequence executable from capture through promotion.\n\nHow checked: inspected the updated command block and ran git diff --check.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Make the Baseline-4 promotion comparison runnable
Standalone tooling extraction. No product behaviour change — the diff is one new read-only CLI,
its test, the
package.jsonscript wiring, onetsconfiginclude entry, and the correspondingstep-7/8 correction in the evidence document.
The defect
The merged Baseline-4 authority (
#335, squashaa023a7) could not actually be applied:assertTimeToAnswerPromotiondocuments itself as “The benchmark job calls this after readingtwo closed JSON artifacts”, but no such job exists — its only callers are unit tests;
perf:time-to-answer --baselinedeliberately refuses with “promotion requires the assembledcomposite evidence, not the incomplete end-to-end section”, because the capture can only
produce the end-to-end section;
docs/testing/TIME_TO_ANSWER_EVIDENCE.mddocumented acomparison command that cannot run.
Every performance child of epic
#333(#336,#337,#338) must compare its candidate againstthe pinned Baseline 4 under
max(baseline × 1.25, baseline + 2 ms). Until this PR there was norunnable way to do that.
What this adds
tests/perf/promote.tsis a read-only comparator that:(
validateTimeToAnswerEvidence), calls the existingassertTimeToAnswerPromotion, and reusesthe existing
derivedTimeToAnswerSummaries/timeToAnswerLimit— it implements no measurementlogic of its own and changes no rule;
fixture | stage | protocol | baseline reviewed | limit | candidate reviewed | delta | verdict), then the list of rows slower than baseline, thenPROMOTION: PASS/PROMOTION: FAIL:<reason>;flag, a duplicate flag or a malformed artifact.
tests/perf/promote.test.mjsspawns the CLI against tiny synthetic composites and asserts thepassing pair exits 0, the over-limit pair exits non-zero naming the offending key, and an
incompatible-environment pair fails with the environment message.
The documentation now describes the real sequence: capture the end-to-end section, capture the
dedicated Stage-5 section, assemble a candidate composite, then compare composites with
perf:promote.Provenance
Extracted, without modification, from the
#336work on branchdm/336-docker-first-answer(draft PR #347) where it was implemented and reviewed;
tests/perf/promote.tsandpromote.test.mjsare byte-identical to that branch's versions (verified withgit show … | diff -). The commit is registered under the#336managed run.Verification
npm run test:perf— 43 tests pass, including the new spawn-based comparator tests.npm run typecheck— pass (tests/perf/tsconfig.jsonnow includespromote.ts).npm run check— exit 0 ondf0e12e(audit, typecheck, contracts, builds, deployment/version/perf suites, JS tests, Rust fmt check + clippy + tests).Not included
No
#336product behaviour: the listener-first startup, concurrent Docker reads, Docker-firstpublication, Compose enrichment and churn-reduction commits all remain on
dm/336-docker-first-answerand are untouched here. Every sentence in the evidence document about
composeEnrichmentMs, theCompose publication budget and
#336is left exactly as it is onmain.Refs #333, #335, #336.