Correct the scale claim and publish the evidence behind it - #507
Merged
Merged
Conversation
The README said CI wins 34 of 34 scaled rows. Main's own run says 32 of 34, and the two rows it lost are not the two the earlier runs lost. Across five consecutive runs the counts are 32, 32, 34, 34, 32. Four rows account for all of it, and all four are at 1000 samples, where a fit finishes in well under a millisecond. Lasso at 1000 rows and 32 features was recorded at 3.83x and at 0.92x on code that differs in nothing touching Lasso. The claim is now the spread across runs rather than one run's number: 30 of the 34 rows win in all five, and the ten rows at 10000 samples, the only sizes large enough to measure without fighting the noise, win in all five. Both KernelSVC_RBF rows at 1000 samples are in the four for a different reason. They lost at 0.68x and 0.88x until the support-vector change and have won every run since at 1.16x to 1.36x, which tracks the commit rather than the runner. summarize_scaled_ci.py folds each run's scaled_comparison.json into scaled_ci_history.json, keyed by run id so a new run extends the history. publish_headline_v2.py validates it and copies it into the Pages tree, and the benchmarks page renders it as a table sorted by median with the observed range beside it. Two published numbers had gone stale the same way, so the check that already self-heals the win counts now covers the substrate means as well. It caught real drift on both READMEs: 20.99x, 6.84x and 3.96x against an artifact holding 19.04x, 7.39x and 4.29x, and 26.55x against 19.04x. The site said Flow wins 75 percent of Python-bound rows, 45 percent of mixed and none of the external-native-bound ones, and the architecture page said it loses every external-native-bound comparison. Both predate winning all 19. index.html now builds that sentence from the artifact it already fetches, so it cannot go stale again, and its hardcoded 8/19 fallback is gone. Removed docs/benchmarks.js and docs/benchmark-corrections.js. No page has loaded either for some time, but benchmarks.js still shipped a hand-written table of old timings to a public URL, and two CI steps existed only to sanitize it. Those steps and the two scripts behind them go with it. Also fixed a duplicated unterminated section opener in benchmarks.html that left the markup with one more section tag than it closed, and brought the site copy in line with the house rules.
The refreshed baseline failed CI on its first run: RandomForest at 1000 rows and 8 features measured 2.88 ms against a 2.15 ms allowance. Nothing had changed. That row has been measured at 1.50, 1.58, 1.63, 2.31, 2.88 and 3.32 ms on identical code, and taking the baseline from one run happened to pick 1.58, near the fastest of them. A self-regression gate should fire when the code is slower than it has ever legitimately been. Each baseline row is now the slowest observation across the runs it was built from, which the same script writes with --baseline. All five runs on record pass against it, including the 2.88 ms measurement that failed, and a 1.5x slowdown injected into RandomForest at 10000 rows still fails it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The correction
The README claimed CI wins 34 of 34 scaled rows. Main's own run says 32 of 34, and the two rows it lost are not the two the earlier runs lost. Across five consecutive runs the counts are 32, 32, 34, 34, 32.
KernelSVC_RBF1000x8KernelSVC_RBF1000x32Lasso1000x32LinearRegression1000x32All four are at 1000 samples, where a fit finishes in well under a millisecond. The bottom two swing by 4x on code that differs in nothing touching them, so that is the runner. The top two show a step change at the support-vector commit and have won every run since, so that is the code.
The published claim is now the spread: 30 of the 34 rows win in all five runs, and the ten rows at 10000 samples, the only sizes large enough to measure without fighting the noise, win in all five.
What backs it
summarize_scaled_ci.pyfolds each run'sscaled_comparison.jsonintoscaled_ci_history.json, keyed by run id so a new run extends the history and a repeated one replaces it.publish_headline_v2.pyvalidates it and copies it into the Pages tree. The benchmarks page renders it as a table sorted by median, with runs-won and the observed range beside each row.Drift the same shape, found on the way
The check that already self-heals the win counts now covers the substrate means too. It immediately caught real drift in both READMEs:
20.99x / 6.84x / 3.96xagainst an artifact holding19.04x / 7.39x / 4.29x, and26.55xagainst19.04x.The site was worse. It still said Flow wins 75% of Python-bound rows, 45% of mixed and none of the external-native-bound ones, and the architecture page said it loses every external-native-bound comparison. Both predate winning all 19.
index.htmlnow builds that sentence from the artifact it already fetches, so it cannot go stale again, and its hardcoded8 / 19fallback is gone.Removed
docs/benchmarks.jsanddocs/benchmark-corrections.js. No page has loaded either for some time, butbenchmarks.jsstill shipped 600 lines of hand-written old timings to a public URL, and two CI steps existed only to sanitize it. Those steps and the two scripts behind them go with it.Also
benchmarks.htmlhad a duplicated unterminated<section>opener, leaving the markup one tag out of balance.docs/*.jsonevidence copies are now gitignored; their sources live inbenchmarks/.Verification
Every page script parses under
node --checkand was rendered against the real artifacts in a DOM harness; the pages were also loaded in a browser. All four evidence validators pass.