Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 0 additions & 5 deletions .github/workflows/flow.yml
Original file line number Diff line number Diff line change
Expand Up @@ -37,11 +37,6 @@ jobs:
PYTHONPATH: ${{ github.workspace }}/.flow-toolchain/src
run: python .github/scripts/verify_flow_sources.py

- name: Audit published benchmark timing units
run: |
python benchmarks/normalize_published_timings.py
python benchmarks/validate_published_timings.py

- name: Validate benchmark contracts and fixtures
run: |
python -m pip install numpy
Expand Down
4 changes: 0 additions & 4 deletions .github/workflows/pages.yml
Original file line number Diff line number Diff line change
Expand Up @@ -53,10 +53,6 @@ jobs:
run: python benchmarks/check_disparity_regression.py --baseline /tmp/no-disparity-baseline.json --commit "${GITHUB_SHA}"
- name: Publish canonical benchmark, disparity, history and architecture evidence
run: python benchmarks/publish_headline_v2.py
- name: Normalize legacy sklearn timings to milliseconds
run: python benchmarks/normalize_published_timings.py
- name: Validate published benchmark timing units
run: python benchmarks/validate_published_timings.py
- uses: actions/configure-pages@v5
- uses: actions/upload-pages-artifact@v3
with:
Expand Down
8 changes: 8 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -23,3 +23,11 @@ __pycache__/
*~
.idea/
.vscode/

# Evidence copies that benchmarks/publish_headline_v2.py writes into the Pages
# tree at build time. The sources live in benchmarks/.
docs/headline-result-v2.json
docs/architecture-performance-map.json
docs/disparity-report.json
docs/disparity-history.json
docs/scaled-ci-history.json
12 changes: 8 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,19 +55,23 @@ The win count is machine-dependent and the committed artifact says which machine

Do not read that 23x as a property of the library. scikit-learn's own fit of that row takes about 25 ms on the Intel runner and about 181 ms on the AMD one, for the same code and the same data, so the AMD figure is measuring an OpenBLAS path that suits that machine badly rather than anything Flow does well. Flow's own time on the two runners is 8.9 ms and 7.5 ms. The Intel ratio is the honest one to quote, and a row whose margin sits near 1x can still land either way. The parity contract gates on correctness and measurement resolution rather than on the win count.

The current grouped headline evidence is descriptive rather than causal: Flow wins every row in all three substrate groups, at a mean of 20.99x on Python-bound rows, 6.84x on mixed rows and 3.96x on external-native-bound rows in the committed architecture map.
The current grouped headline evidence is descriptive rather than causal: Flow wins every row in all three substrate groups, at a mean of 19.04x on Python-bound rows, 7.39x on mixed rows and 4.29x on external-native-bound rows in the committed architecture map.

Two earlier readings of this table were wrong, and both were artifacts of how Flow was built rather than of the substrate. While the Flow side was compiled unoptimized, external-native-bound rows all lost, which read as sklearn-owned compiled code being out of reach. The grouping is a guide to where the Python boundary costs most. It is not a ceiling.

One row is not a like-for-like comparison, and the disparity report records it. Flow fits a forest's trees concurrently; scikit-learn's default is one worker, and the benchmark leaves it at its default. Both RandomForest rows therefore carry a declared `n_jobs` difference. Single-threaded, RandomForest on digits runs at 1.82x rather than 5.01x, so the row wins either way. Asking scikit-learn for all cores does not close the gap on this workload: at `n_jobs=-1` its own fit measured slower than at `n_jobs=1`, because joblib's pool costs more than ten small trees save.

## Larger data

The 19 canonical rows sit on iris, digits and diabetes, none of which exceeds 1797 samples. A separate matrix runs five estimators at 100, 1000 and 10000 rows against 8 and 32 features, which is where an implementation that only suits small inputs would show it.
The 19 canonical rows sit on iris, digits and diabetes, none of which exceeds 1797 samples. A separate matrix runs five estimators at 100, 1000 and 10000 rows against 8 and 32 features, which is where an implementation that only suits small inputs would show it. CI measures it on every run and reports it without gating the build.

CI measures this matrix on every run, on an Intel Xeon with OpenBLAS, and Flow wins 34 of 34 of those rows. The narrowest are `LinearRegression` at 10000 rows and 8 features (1.26x) and both `KernelSVC_RBF` rows at 1000 samples (1.29x); the widest is `GaussianNB` at 100 rows and 32 features (74.4x). Four rows were losses before the coordinate-descent, Cholesky, kernel-cache and support-vector changes in this repo's history, the worst at 0.24x.
One run of that matrix does not settle a row. `benchmarks/scaled_ci_history.json` holds five consecutive runs. Flow wins 30 of the 34 rows in all five. The other four dipped below 1x in at least one run, and all four are at 1000 samples, where a fit finishes in well under a millisecond and the runner moves the measurement by more than the difference being measured: Lasso at 1000 rows and 32 features was recorded at 3.83x and at 0.92x on code that differs in nothing touching Lasso.

The matrix is reported without gating the build, because a shared runner moves a tenth-of-a-millisecond row by more than the row itself. What does gate is `benchmarks/scaled_flow_baseline.json`, a Flow-against-itself comparison. Refresh it from a CI artifact rather than from a developer machine, or the gate will read a fast laptop as the standard and fail every CI run.
Both `KernelSVC_RBF` rows at 1000 samples are in that group for a different reason. They lost consistently, at 0.68x and 0.88x, until a fitted one-vs-one model stopped keeping the whole training set and started keeping its support vectors. They have won every run since, at 1.16x to 1.36x. That is a step change that tracks the commit rather than the runner.

The ten rows at 10000 samples, the only sizes in the matrix large enough to measure without fighting the noise, win in all five runs, from a median 1.15x on RandomForest at 32 features to 8.06x on GaussianNB at 8.

`benchmarks/scaled_flow_baseline.json` is a separate Flow-against-itself gate that does fail the build. Refresh it from a CI artifact rather than from a developer machine, or the gate will read a fast laptop as the standard and fail every CI run.

Detailed artifacts:

Expand Down
28 changes: 27 additions & 1 deletion benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -195,10 +195,36 @@ The main generated views are:

A speedup is `sklearn_ms / flow_ms`. Values above `1x` mean Flow is faster; values below `1x` mean scikit-learn is faster.

The current architecture map shows a useful but non-causal pattern: Flow wins every headline row in all three substrate classes, and the margin varies with the class, from a mean of 26.55x on Python-bound rows down to 3.79x on external-native-bound ones. That ordering is evidence for prioritization. It is not proof that execution substrate alone determines performance.
The current architecture map shows a useful but non-causal pattern: Flow wins every headline row in all three substrate classes, and the margin varies with the class, from a mean of 19.04x on Python-bound rows down to 4.29x on external-native-bound ones. That ordering is evidence for prioritization. It is not proof that execution substrate alone determines performance.

Mature BLAS/LAPACK, liblinear, libsvm and other native backends are treated as native competitors. The optimization roadmap deliberately prefers retaining those kernels unless benchmark and parity evidence justify replacement.

## Scaled matrix

The canonical rows run on iris, digits and diabetes and stop at 1797 samples. `bench_scaled.flow` and `bench_scaled_sklearn.py` run five estimators at 100, 1000 and 10000 rows against 8 and 32 features. CI runs both sides every push and reports the comparison without gating the build.

One run does not settle a row. At 100 and 1000 samples a fit finishes in well under a millisecond, and the runner moves that by more than the Flow-versus-sklearn difference: Lasso at 1000 rows and 32 features was recorded at 3.83x on one run and 0.92x on another, on code that differs in nothing touching Lasso.

So the published claim is the spread. [`summarize_scaled_ci.py`](summarize_scaled_ci.py) folds each run's `scaled_comparison.json` into `scaled_ci_history.json`, which records every observation per row along with its minimum, median and maximum:

```
gh run download <id> -n scaled-benchmark-<id> -D /tmp/<id>
python benchmarks/summarize_scaled_ci.py <id>=/tmp/<id>/scaled_comparison.json
```

Runs merge by id, so adding a new one extends the history and re-adding an existing one replaces it.

`scaled_flow_baseline.json` is a different artifact for a different job. It compares Flow against its own earlier CI timings and does fail the build, with a 20% relative tolerance and a 0.25 ms absolute floor.

Each of its rows is the slowest observation across the runs it was built from, so the gate fires when the code is slower than it has ever legitimately been and stays quiet when a run is merely unlucky. Taken from one run it does the opposite: RandomForest at 1000 rows and 8 features has been measured at 1.50, 1.58, 1.63, 2.31, 2.88 and 3.32 ms on identical code, and a baseline taken from the 1.58 run failed the build on the 2.88 one. Rebuild it with the same script:

```
python benchmarks/summarize_scaled_ci.py --baseline benchmarks/scaled_flow_baseline.json \
<id>=/tmp/<id>/scaled_comparison.json ...
```

Each run directory needs `scaled_flow.json` beside `scaled_comparison.json`. Build it from CI artifacts; a developer machine's numbers would make a fast laptop the standard CI has to meet.

## Pages publication

[`publish_headline_v2.py`](publish_headline_v2.py) validates the committed canonical benchmark and architecture map, then copies the JSON artifacts into `docs/` for the static site. The public benchmark and architecture pages render those artifacts directly instead of embedding hand-maintained timing claims.
52 changes: 0 additions & 52 deletions benchmarks/normalize_published_timings.py

This file was deleted.

24 changes: 23 additions & 1 deletion benchmarks/publish_headline_v2.py
Original file line number Diff line number Diff line change
Expand Up @@ -14,10 +14,12 @@
ARCH = BENCH / "architecture_performance_map.json"
DISPARITY = BENCH / "disparity_report.json"
HISTORY = BENCH / "disparity_history.json"
SCALED = BENCH / "scaled_ci_history.json"
DOC_RESULT = DOCS / "headline-result-v2.json"
DOC_ARCH = DOCS / "architecture-performance-map.json"
DOC_DISPARITY = DOCS / "disparity-report.json"
DOC_HISTORY = DOCS / "disparity-history.json"
DOC_SCALED = DOCS / "scaled-ci-history.json"


def validate_headline(result: dict) -> None:
Expand Down Expand Up @@ -60,6 +62,20 @@ def validate_history(history: dict, total_rows: int) -> None:
raise SystemExit("disparity history snapshot does not cover every canonical row")


def validate_scaled(scaled: dict) -> None:
counts = scaled["counts"]
if counts["runs"] < 2:
raise SystemExit("the scaled matrix needs at least two runs before it says anything")
if counts["rows"] != len(scaled["rows"]):
raise SystemExit("scaled history row count does not match counts.rows")
won = sum(1 for r in scaled["rows"] if r["runs_won"] == r["runs_observed"])
if won != counts["rows_won_in_every_run"]:
raise SystemExit("scaled history win count disagrees with its own rows")
for row in scaled["rows"]:
if len(row["observations"]) != row["runs_observed"]:
raise SystemExit(f"scaled row {row['algorithm']} has a stale observation list")


def main() -> int:
parser = argparse.ArgumentParser()
parser.add_argument("--check", action="store_true")
Expand All @@ -76,6 +92,11 @@ def main() -> int:
validate_disparity(disparity, result["counts"]["total_rows"])
elif not args.check:
raise SystemExit("disparity_report.json must be generated before Pages publication")
scaled = json.loads(SCALED.read_text()) if SCALED.exists() else None
if scaled is not None:
validate_scaled(scaled)
elif not args.check:
raise SystemExit("scaled_ci_history.json must be present before Pages publication")
if history is not None:
validate_history(history, result["counts"]["total_rows"])
elif not args.check:
Expand All @@ -86,8 +107,9 @@ def main() -> int:
shutil.copyfile(ARCH, DOC_ARCH)
shutil.copyfile(DISPARITY, DOC_DISPARITY)
shutil.copyfile(HISTORY, DOC_HISTORY)
shutil.copyfile(SCALED, DOC_SCALED)
counts = result["counts"]
print(f"published evidence: {counts['flow_wins']}/{counts['eligible_comparisons']} Flow wins; {disparity['counts']['rows_with_tracked_disparity']} rows with tracked disparities; {len(history['snapshots'])} history snapshots")
print(f"published evidence: {counts['flow_wins']}/{counts['eligible_comparisons']} Flow wins; {disparity['counts']['rows_with_tracked_disparity']} rows with tracked disparities; {len(history['snapshots'])} history snapshots; scaled matrix over {scaled['counts']['runs']} runs")
else:
print("canonical benchmark, architecture, disparity, and available history evidence are internally consistent")
return 0
Expand Down
Loading
Loading