Skip to content

Correct the scale claim and publish the evidence behind it - #507

Merged
godofecht merged 2 commits into
mainfrom
site-evidence
Sep 26, 2026
Merged

godofecht merged 2 commits into
mainfrom
site-evidence

Conversation

@godofecht

Copy link
Copy Markdown
Owner

The correction

The README claimed CI wins 34 of 34 scaled rows. Main's own run says 32 of 34, and the two rows it lost are not the two the earlier runs lost. Across five consecutive runs the counts are 32, 32, 34, 34, 32.

row run 1 run 2 run 3 run 4 run 5
KernelSVC_RBF 1000x8 0.68 0.66 1.29 1.30 1.16
KernelSVC_RBF 1000x32 0.88 0.83 1.29 1.34 1.36
Lasso 1000x32 3.83 4.07 1.51 1.47 0.92
LinearRegression 1000x32 2.18 1.79 1.80 1.87 0.90

All four are at 1000 samples, where a fit finishes in well under a millisecond. The bottom two swing by 4x on code that differs in nothing touching them, so that is the runner. The top two show a step change at the support-vector commit and have won every run since, so that is the code.

The published claim is now the spread: 30 of the 34 rows win in all five runs, and the ten rows at 10000 samples, the only sizes large enough to measure without fighting the noise, win in all five.

What backs it

summarize_scaled_ci.py folds each run's scaled_comparison.json into scaled_ci_history.json, keyed by run id so a new run extends the history and a repeated one replaces it. publish_headline_v2.py validates it and copies it into the Pages tree. The benchmarks page renders it as a table sorted by median, with runs-won and the observed range beside each row.

Drift the same shape, found on the way

The check that already self-heals the win counts now covers the substrate means too. It immediately caught real drift in both READMEs: 20.99x / 6.84x / 3.96x against an artifact holding 19.04x / 7.39x / 4.29x, and 26.55x against 19.04x.

The site was worse. It still said Flow wins 75% of Python-bound rows, 45% of mixed and none of the external-native-bound ones, and the architecture page said it loses every external-native-bound comparison. Both predate winning all 19. index.html now builds that sentence from the artifact it already fetches, so it cannot go stale again, and its hardcoded 8 / 19 fallback is gone.

Removed

docs/benchmarks.js and docs/benchmark-corrections.js. No page has loaded either for some time, but benchmarks.js still shipped 600 lines of hand-written old timings to a public URL, and two CI steps existed only to sanitize it. Those steps and the two scripts behind them go with it.

Also

  • benchmarks.html had a duplicated unterminated <section> opener, leaving the markup one tag out of balance.
  • Site copy brought in line with the house rules. The commit hook found these, including some I introduced.
  • The generated docs/*.json evidence copies are now gitignored; their sources live in benchmarks/.

Verification

Every page script parses under node --check and was rendered against the real artifacts in a DOM harness; the pages were also loaded in a browser. All four evidence validators pass.

The README said CI wins 34 of 34 scaled rows. Main's own run says 32 of 34,
and the two rows it lost are not the two the earlier runs lost. Across five
consecutive runs the counts are 32, 32, 34, 34, 32.

Four rows account for all of it, and all four are at 1000 samples, where a fit
finishes in well under a millisecond. Lasso at 1000 rows and 32 features was
recorded at 3.83x and at 0.92x on code that differs in nothing touching Lasso.
The claim is now the spread across runs rather than one run's number:
30 of the 34 rows win in all five, and the ten rows at 10000 samples, the only
sizes large enough to measure without fighting the noise, win in all five.

Both KernelSVC_RBF rows at 1000 samples are in the four for a different
reason. They lost at 0.68x and 0.88x until the support-vector change and have
won every run since at 1.16x to 1.36x, which tracks the commit rather than the
runner.

summarize_scaled_ci.py folds each run's scaled_comparison.json into
scaled_ci_history.json, keyed by run id so a new run extends the history.
publish_headline_v2.py validates it and copies it into the Pages tree, and
the benchmarks page renders it as a table sorted by median with the observed
range beside it.

Two published numbers had gone stale the same way, so the check that already
self-heals the win counts now covers the substrate means as well. It caught
real drift on both READMEs: 20.99x, 6.84x and 3.96x against an artifact
holding 19.04x, 7.39x and 4.29x, and 26.55x against 19.04x.

The site said Flow wins 75 percent of Python-bound rows, 45 percent of mixed
and none of the external-native-bound ones, and the architecture page said it
loses every external-native-bound comparison. Both predate winning all 19.
index.html now builds that sentence from the artifact it already fetches, so
it cannot go stale again, and its hardcoded 8/19 fallback is gone.

Removed docs/benchmarks.js and docs/benchmark-corrections.js. No page has
loaded either for some time, but benchmarks.js still shipped a hand-written
table of old timings to a public URL, and two CI steps existed only to
sanitize it. Those steps and the two scripts behind them go with it.

Also fixed a duplicated unterminated section opener in benchmarks.html that
left the markup with one more section tag than it closed, and brought the
site copy in line with the house rules.
The refreshed baseline failed CI on its first run: RandomForest at 1000 rows
and 8 features measured 2.88 ms against a 2.15 ms allowance. Nothing had
changed. That row has been measured at 1.50, 1.58, 1.63, 2.31, 2.88 and
3.32 ms on identical code, and taking the baseline from one run happened to
pick 1.58, near the fastest of them.

A self-regression gate should fire when the code is slower than it has ever
legitimately been. Each baseline row is now the slowest observation across the
runs it was built from, which the same script writes with --baseline. All five
runs on record pass against it, including the 2.88 ms measurement that failed,
and a 1.5x slowdown injected into RandomForest at 10000 rows still fails it.
@godofecht
godofecht merged commit 7e96ee8 into main Sep 26, 2026
10 checks passed
@godofecht
godofecht deleted the site-evidence branch September 26, 2026 11:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant