feat(bench): resolve runs against a workload catalog and qualify every comparison against its clock - #728
Merged
Merged
Conversation
added 13 commits
September 11, 2026 15:59
…bsent-arm figures The page led with whichever profile was measured last, which is a profile that may carry one domain: a physics-only run finished three hours after a rendering-only one buried every rendering comparison in the collapsed list below. Coverage now outranks the measurement date, so a profile carrying both domains leads. A comparison cell declared a ten-rem text track beside a four-rem minimum plot, whose intrinsic width exceeded the column a five-pair table gives it: the table pushed past the viewport and the log axis collapsed to about sixty pixels, on which every ratio short of 10x drew a few pixels. Text and bar each take a lane now, and the centre line the lengths are read against is visible. Column widths were declared on the header cells, which a fixed layout ignores when the first row is the group row and carries colspans; they move to a colgroup, and the count column stops holding a few hundred pixels of white space beside a four-digit number. An arm that sat a comparison out is stored as a zero rather than as a missing value, so the medians, p95, frame share and GPU rows printed 0.000 ms - the fastest figure on the page - for the arm that never ran. Archetype descriptions stay with the page rather than travelling in every profile: they are prose for a reader, identical on every machine, and two published profiles could otherwise disagree about what one archetype means. Schema 7 therefore adds only the GPU frame time, and reads version 6 unchanged. The local probe drained its timer queries immediately after ending them, so no sample was ever available and its untimed fallback reported the submission time under a GPU-inclusive label. Claude-Session: https://claude.ai/code/session_01PXSHGYhLZX3zbLbPFKQVmg
…view The page put both measured domains one under the other, so a reader after a renderer scrolled past twenty-one physics rows to reach the tables they came for and the two scoreboards, legends, methodology notes and provenance blocks were paid for twice. The domains share no row and answer different questions, so each gets its own page and the section's own URL serves rendering. The two views are ordinary links drawn as tabs: two static pages need no panel, no selection state and no arrow-key model, and either can be linked to. The machine profile is not swapped when a view is opened - a profile that measured one domain and not the other says so, because two machines' numbers do not belong in one reading. A comparison cell carried its own disclosure, so a five-pair row offered five toggles that each opened the same six terms; they become one disclosure per archetype, which also gives the terms room to be laid out as a list rather than wrapped into a column. Stacked into cards, an arm with no cell no longer takes a labelled block with a rule and the height of a result: the gap is visible in the matrix, where the columns beside it show it, and the detail row names the arms that were not compared either way. A pair inside the noise band now reads as "similar" rather than as 1.00x. The ladder pins the factor of a level rung to exactly one, so the cell printed a two-decimal ratio above two medians that work out to a different one - a precision the verdict never measured. Nothing about the ladder, its thresholds or the stored numbers changes; only what the cell prints. Where a table's columns fall into groups, a filter narrows it to one of them. It is an enhancement, not a gate: with no script every column is shown, which is what the indexed and printed page carries. Claude-Session: https://claude.ai/code/session_01PXSHGYhLZX3zbLbPFKQVmg
A fixed two decimals on a factor and three on a millisecond value printed digits the measurement never resolved: 669.00x for a ratio whose denominator is one clock tick, and 13.380 ms for a value whose pooled runs spanned 12.73 to 14.10. The decimals now follow the magnitude, so 669x, 13.4 and 8.67 replace them and a sub-millisecond value keeps the three decimals its column aligns on. Trailing zeros are kept rather than stripped. 0.80 beside 0.12 is a column a reader scans down, and the zero is a digit the clock did resolve; what the rule removes is the one it did not. Claude-Session: https://claude.ai/code/session_01PXSHGYhLZX3zbLbPFKQVmg
…ally The scoreboard opened every view: coverage, five outcome states, a strip whose unfilled part is not a result, and a paragraph on how not to read any of it - all of it asked before a reader had seen a single number. It moves under the tables, where it summarises rows already met, and the key above them keeps only what a bar's colour means. Stacked into cards, an archetype carried every one of its comparisons one under the other. A picker now names the pair on screen, and the narrow view opens on the arm ExoJS leads on fewest rows rather than the one it leads on most - the choice decides which comparison is visible first and publishes no figure, so nothing is summed, ranked or counted from it. Wide, the picker steps aside for the matrix, which shows the pairs side by side by design. A factor is read for its size rather than its digits, so it drops to whole numbers from ten up and one decimal below: 72x and 2x say what 72.25x and 2.00x said. Milliseconds keep their own rule - a time is compared against a frame budget, not against another factor. A pair inside the noise band prints the word alone, and what the word means is stated once in the key. Claude-Session: https://claude.ai/code/session_01PXSHGYhLZX3zbLbPFKQVmg
A benchmark that reports milliseconds without recording the grid its clock delivers them on cannot tell a fast cell from an unresolved one. The physics harness already measured that grid; the probe moves to the shared layer so the rendering harness can take the same reading in its own page, which is where it has to be taken - a value read in the driver process describes Node's clock. An unobserved resolution is reported as absent rather than as zero. The previous probe returned 0 when its loop saw no positive step, which reads downstream as a perfectly fine clock - the opposite of what the failed observation established. The report states what it is: the smallest step the probe saw, not a calibrated error bound. Engines may coarsen and jitter timestamps, so one observed minimum bounds neither the error of a sample nor the confidence of a comparison. Claude-Session: https://claude.ai/code/session_01PXSHGYhLZX3zbLbPFKQVmg
A duration a step or two above the grid performance.now() delivers on carries no ratio worth printing: measured again the same scene lands on the neighbouring step, and the factor built from it moves by a whole multiple. One guard now decides that, and the cell, the bar, the label and the scoreboard all read it from `outcomeOf` rather than each testing their own condition. Both durations are checked on their own against the step observed in the context that measured them - how large a reading is against the grid it was read on, not how far the two arms are apart - and a pooled figure inherits the coarsest step of the runs behind it, because a repetition measured on a coarse clock is not repaired by one measured on a fine one. The threshold is a guard set clear of the one- and two-step readings a coarse clock produces, and it is the same for every library, browser and profile. It is not a standard and not a precision claim: clearing it establishes that this one check did not trip, never that a comparison is accurate. A profile written before the step was recorded yields neither a pass nor a failure but the absence of the check, and reads exactly as it did before. A cell the guard stops keeps both measured times and loses the factor, the bar and the winner, and the scoreboard counts it on no summary line. That is a refusal to publish a comparison, not a claim that the two libraries are equally fast; a rule for keeping a direction where the gap is wide would need evidence of its own and is deliberately not attempted here. Claude-Session: https://claude.ai/code/session_01PXSHGYhLZX3zbLbPFKQVmg
…k backs The check ran against the pooled medians and one coarsest step, which lets a well-resolved repetition carry a limited one past the threshold: three runs stepping 0.020, 0.005 and 0.005 ms with durations of 0.180, 0.300 and 0.300 pool to 0.300 against 0.020 - fifteen steps, a pass - while the first run stood nine steps above its own clock and did not. The published profile carries no per-run durations, so the site cannot make that judgement at all; the check moves to a per-run primitive and a merge rule, and the cell reads the verdict the harness will publish beside it. Merging is deliberately asymmetric: a limitation any run established stands, and a run whose clock was never recorded cannot lift it, because missing information does not cancel an established finding. Only a comparison whose every run cleared the check is reported as resolved. A profile that recorded no clock no longer keeps its factors. It was the state every published profile is in, so the exemption would have left exactly the figures the check exists to withhold - and the page is not emptied by removing them: scenarios, both measured times, the machine and the detail all stay, and the reason is stated once for the profile rather than in each of its cells. `timer-unknown` stays apart from `timer-limited`, because a check that could not be made and a check that tripped are different statements, and the scoreboard counts neither as a win, a loss or a level row. Claude-Session: https://claude.ai/code/session_01PXSHGYhLZX3zbLbPFKQVmg
A benchmark run had one shape: every archetype's full ladder against every arm. That is the development matrix, and publishing a comparison meant paying for it in full even though most of its rungs answer a scaling question no reader of the comparison asks. Runs now resolve against a named plan. `reference` selects the loads the project publishes - one headline load per scenario, a short ladder only where the scaling is itself the finding - and `full` keeps the whole development matrix, including the ExoJS-internal probes the published catalog deliberately omits. `--extreme` admits the named million-scale loads, and only together with `full`: an extreme load exists to find where a scenario stops being viable, which is not a published claim. Every load carries the unit it is counted in, because the number alone is ambiguous across the catalog - a million world tiles and a million visible particles are not the same measurement. A plan carries a hash over its semantic content only, so two runs can be checked for having measured the same contract without a path or a timestamp deciding it. `--domain=all` runs both domains serially into separate subdirectories, and `--dry-run` reports the resolved workload without starting a browser. It names the catalogued scenarios that have no archetype behind them yet rather than quietly planning around them. Claude-Session: https://claude.ai/code/session_01PXSHGYhLZX3zbLbPFKQVmg
The clock probe landed without a producer: nothing measured a rendering page's grid, nothing attached one to a cell, and the check that reads it sat in the site with no caller. Every rendering comparison therefore published as "timer not recorded", which reads like a finding and is merely a missing wire. The harness now probes the grid in the page, once per session, and every cell carries the grid of the session that produced it. That attribution matters here: a rendering run opens one browser session per arm, so a backend-wide figure read from whichever page happened to be open would qualify cells it never timed. The comparison builder checks each arm's duration against its own session's grid and the pooling stage merges the per-run results, so a limitation one run established cannot be lifted by a better-resolved repetition. Physics is checked against the batch the clock actually bracketed rather than the per-step quotient the report divides out, which would mark a well-resolved cell unresolved purely for being batched. The rendering block also stops collapsing onto a single table-wide node count. A published page offers the reader a load to pick, and one count discarded every other measurement before the profile was written. Each row is now one archetype at one load, stating the load and the unit it is counted in, so a hundred thousand world tiles can never read as a hundred thousand sprites. What the single count protected - that no load is chosen to suit an outcome - is protected by the plan instead, which fixes the loads before the run starts. Schema version 7 carries all of it, and the result verifier requires it of a version 7 document rather than trusting the stage that writes it. Version 6 profiles still read exactly as they did. Claude-Session: https://claude.ai/code/session_01PXSHGYhLZX3zbLbPFKQVmg
…cene Two scenes the matrix never had, both built from the sprite path every arm already implements. `dynamic-all` is `dynamic-heavy` with every leaf moving instead of 7.5 % of them. The existing row is the shape a real scene has - a few actors over a mostly still background - so the delta between the two is what the still 92.5 % costs once it stops being still. It is a new archetype rather than a raised mutation fraction on the old one, because both questions are worth publishing and changing the old one would silently invalidate every number measured under its name. `fill-layers` is a stack of translucent full-screen layers, the workload a parallax background plus a weather pass plus a few tint overlays adds up to. It shares its geometry with `overdraw`, which sweeps thousands of viewport-sized quads to find the fill ceiling - not something anything ships. This one sweeps 8 to 128 and fixes a low per-layer alpha, so every layer has to be composited rather than skipped by an occlusion policy. The geometry the two share used to be an archetype-id test repeated in each of the four arms, which is how an archetype added with the same shape under another name would have been laid out four different ways. It is a trait predicate now, like the text and mask questions beside it. The structural gate leaves `fill-layers` unguarded for the reason it already leaves `overdraw` unguarded: the fill is enormous under a software rasterizer and its draw structure is one call `static-heavy` guards already. Claude-Session: https://claude.ai/code/session_01PXSHGYhLZX3zbLbPFKQVmg
The one scene shape practically every 2D game has and that no sprite archetype describes: a world far larger than the viewport, drawn through a dedicated tile path rather than one node per tile. `scrolling-world` is not this test - it lays independent sprite nodes over a few viewports and its per-node cost is the finding, while here the world is a hundred thousand tiles, the visible window never changes size, and what is compared is each arm's tile path: an instanced chunk renderer, an imperatively painted quad buffer, a shader over a data texture. Two scenes. `tilemap-scroll` scrolls a fully populated map past a fixed window, so its ladder says how large a map an arm can hold rather than how much of it it draws. `tilemap-edit` replaces visible tile ids every frame, which is where the three paths differ most sharply, because each has to get the change to the GPU a different way. The editing scene holds its camera still, and that is a determinism requirement rather than a simplification. The harness may cut an arm's warmup short on wall clock, so two arms can reach the timed window having run a different number of frames; with a moving camera the edited cells move too, and the arms then draw worlds differing by whatever the extra frames edited. Holding the camera makes each frame set the same cells from that frame alone, and keeps the delta against the scrolling scene the edit cost by itself. Excalibur sits both scenes out. Its `TileMap` is a grid of tiles each holding its own graphics list, drawn through the ordinary graphics path - there is no dedicated tile submission path of the kind the other three arms are being compared on, and building one would mean writing the arm's missing feature rather than adapting to it. The driver can now capture each measured cell's final frame, which is what found two real defects in this work: the Pixi arm scrolled at double speed (the tile pipe combines the rendered root's transform with its own, so the scrolled container is nested now), and repainting one `Tilemap` instance twice renders nothing at all (an edited chunk gets a fresh instance). A cell that draws the wrong scene still produces a number, and that number looks like a result.
…lating Two scenes on one draw path, kept apart because they answer different questions and a figure from one would be read as the other. `particles-draw` submits a fixed set of small translucent quads and simulates nothing, so it measures the submission path alone: the ExoJS particle renderer, Pixi's `ParticleContainer` with every property declared static, a Phaser emitter whose simulation is never stepped. A million quads there is a million quads drawn, not a million interactive sprites. `particles-lifecycle` puts the shared effect on top of that same path - a two second life, a linear drift, a linear fade and a respawn at the end - with the pool held at the node count. ExoJS and Phaser each run their own emitter, which is what is under comparison. Pixi has no emitter at all, so its arm pairs `ParticleContainer` with exactly the shared update rule written out in the harness; that code sits inside the measured bracket and is named for what it is rather than passed off as a Pixi feature. Verified by capture rather than by reading the adapters: all three arms light the identical 161233 pixels in the draw scene, and land within half a percent of each other in the lifecycle scene, where their particles are not on the same coordinates by design. The harness was resolving the official extension packages to their BUILT output while resolving the engine beside them to source, so an extension arm measured a different tree from the rest of the matrix - and the two copies of a class failed every `instanceof` across the boundary, which is how a `ParticleSystem` silently kept its default capacity of 4096 instead of the ten thousand its cell asked for. The dev server now maps the extension entries and the engine's subpath exports to source, and resolves `#*` per importer so each package's own map still reaches its own sources.
Rendering and physics were two pages fronted by a scoreboard and a wall of tables. A reader arriving with "how does ExoJS do on the work I am about to do" had to pick a domain, learn a table and read a tally before anything answered them. One page now, and it opens on result cards: the scenario in plain words, the load it was measured at, and one horizontal bar per library with its time on it. Bars are linear from zero within a card, so a bar's length is the time it shows and two bars can be read directly against each other - no log scale and no ratio axis, both of which have to be learned before they can be read and both of which make a small difference look large. A card compares the arms within one load and never two loads or two cards, because those are different scenes. Which scenarios open each section is fixed in the source before any run happens, and a profile that carries only some of them is topped up in its own order rather than by result, so the top of the page cannot become a selection of whatever ExoJS won. Losses keep the same treatment as wins: the clipping card leads with ExoJS's bar being the long one. Where a scenario was measured at several loads the card offers them as buttons, each showing figures that were actually measured - nothing is interpolated between rungs. The full tables keep every row, spread, p95 and omission one disclosure below, and the methodology stays at the foot where a reader reaches for it once a row surprises them. The old physics URL keeps working: it has been linked to, so it redirects into the physics section rather than 404ing, and it is a redirect rather than a second copy because two pages carrying one profile would eventually disagree about it.
Exoridus
enabled auto-merge (squash)
September 11, 2026 14:20
Bundle ReportChanges will increase total bundle size by 19.36MB (59.94%) ⬆️
Affected Assets, Files, and Routes:view changes for bundle: site-server-esmAssets Changed:
App Routes Affected:
|
❌ 1 Tests Failed:
View the top 1 failed test(s) by shortest run time
To view more test analytics, go to the Test Analytics Dashboard |
added 2 commits
September 11, 2026 16:35
…and the unit lane The rendering adapters import @codexo/exojs-particles and @codexo/exojs-tilemap, whose package entries point at dist. Neither the bench type-check nor the jsdom unit lane builds the packages, so both resolved nothing while a tree with a built dist lying around resolved fine. Mapping the entries to source is not enough on its own for particles: it resolves its own #* imports through a package imports map whose default arm is dist/esm/*.d.ts, so the type-check read stale declarations instead of the sources beside them. @codexo/exojs-particles-source in customConditions is what steers that map to src. The unit lane needed the same distinction one level down. Its blanket #* alias sent #distributions/Curve, imported from inside exojs-particles, into the engine tree. A resolver that declines for importers inside an extension package leaves those specifiers to Vite's own imports resolution, which srcConditions already steers to source; everything else it delegates back through resolution rather than assuming an extension, so shader specifiers keep working. Claude-Session: https://claude.ai/code/session_01UWQw3PuiCFjTVJJBY4AHQG
Vitest sizes its fork pool at one worker per core, so the unit lane saturated a 16-core workstation for the three minutes it runs - on every push, since the pre-push hook runs it. Half the cores costs about 6% wall time here (209s to 197s) because the lane is not purely CPU-bound. CI keeps the default: its runners have few cores and nothing else to serve. EXOJS_TEST_MAX_WORKERS overrides both. Claude-Session: https://claude.ai/code/session_01UWQw3PuiCFjTVJJBY4AHQG
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Splits the benchmark work that accumulated on the scoreboard branch after #727 was cut.
Workload catalog
Runs now resolve against a published workload catalog instead of carrying their own ad-hoc scene descriptions. A result that names a workload the catalog does not define is rejected rather than rendered as a comparison, so a stale result file cannot quietly claim a scenario that no longer exists.
Clock qualification
A shared probe measures the page's clock resolution, and every comparison carries the clock that timed it. Where the resolution cannot back the difference between two arms, the page withholds the comparison instead of printing a ratio the timer never resolved. The judgement is per run, not per page.
New rendering arms
Site
The benchmarks split into a rendering view and a physics view, each led by its results rather than by the tally. Figures print to three significant digits, and the lead profile, ratio bars and absent-arm figures render correctly again.
Claude-Session: https://claude.ai/code/session_01UWQw3PuiCFjTVJJBY4AHQG