diff --git a/.github/workflows/flow.yml b/.github/workflows/flow.yml index 9f567ba..0b27b5f 100644 --- a/.github/workflows/flow.yml +++ b/.github/workflows/flow.yml @@ -37,11 +37,6 @@ jobs: PYTHONPATH: ${{ github.workspace }}/.flow-toolchain/src run: python .github/scripts/verify_flow_sources.py - - name: Audit published benchmark timing units - run: | - python benchmarks/normalize_published_timings.py - python benchmarks/validate_published_timings.py - - name: Validate benchmark contracts and fixtures run: | python -m pip install numpy diff --git a/.github/workflows/pages.yml b/.github/workflows/pages.yml index 3b07301..2d27efe 100644 --- a/.github/workflows/pages.yml +++ b/.github/workflows/pages.yml @@ -53,10 +53,6 @@ jobs: run: python benchmarks/check_disparity_regression.py --baseline /tmp/no-disparity-baseline.json --commit "${GITHUB_SHA}" - name: Publish canonical benchmark, disparity, history and architecture evidence run: python benchmarks/publish_headline_v2.py - - name: Normalize legacy sklearn timings to milliseconds - run: python benchmarks/normalize_published_timings.py - - name: Validate published benchmark timing units - run: python benchmarks/validate_published_timings.py - uses: actions/configure-pages@v5 - uses: actions/upload-pages-artifact@v3 with: diff --git a/.gitignore b/.gitignore index 2951b90..be0e7d1 100644 --- a/.gitignore +++ b/.gitignore @@ -23,3 +23,11 @@ __pycache__/ *~ .idea/ .vscode/ + +# Evidence copies that benchmarks/publish_headline_v2.py writes into the Pages +# tree at build time. The sources live in benchmarks/. +docs/headline-result-v2.json +docs/architecture-performance-map.json +docs/disparity-report.json +docs/disparity-history.json +docs/scaled-ci-history.json diff --git a/README.md b/README.md index b95e6ea..6d83f79 100644 --- a/README.md +++ b/README.md @@ -55,7 +55,7 @@ The win count is machine-dependent and the committed artifact says which machine Do not read that 23x as a property of the library. scikit-learn's own fit of that row takes about 25 ms on the Intel runner and about 181 ms on the AMD one, for the same code and the same data, so the AMD figure is measuring an OpenBLAS path that suits that machine badly rather than anything Flow does well. Flow's own time on the two runners is 8.9 ms and 7.5 ms. The Intel ratio is the honest one to quote, and a row whose margin sits near 1x can still land either way. The parity contract gates on correctness and measurement resolution rather than on the win count. -The current grouped headline evidence is descriptive rather than causal: Flow wins every row in all three substrate groups, at a mean of 20.99x on Python-bound rows, 6.84x on mixed rows and 3.96x on external-native-bound rows in the committed architecture map. +The current grouped headline evidence is descriptive rather than causal: Flow wins every row in all three substrate groups, at a mean of 19.04x on Python-bound rows, 7.39x on mixed rows and 4.29x on external-native-bound rows in the committed architecture map. Two earlier readings of this table were wrong, and both were artifacts of how Flow was built rather than of the substrate. While the Flow side was compiled unoptimized, external-native-bound rows all lost, which read as sklearn-owned compiled code being out of reach. The grouping is a guide to where the Python boundary costs most. It is not a ceiling. @@ -63,11 +63,15 @@ One row is not a like-for-like comparison, and the disparity report records it. ## Larger data -The 19 canonical rows sit on iris, digits and diabetes, none of which exceeds 1797 samples. A separate matrix runs five estimators at 100, 1000 and 10000 rows against 8 and 32 features, which is where an implementation that only suits small inputs would show it. +The 19 canonical rows sit on iris, digits and diabetes, none of which exceeds 1797 samples. A separate matrix runs five estimators at 100, 1000 and 10000 rows against 8 and 32 features, which is where an implementation that only suits small inputs would show it. CI measures it on every run and reports it without gating the build. -CI measures this matrix on every run, on an Intel Xeon with OpenBLAS, and Flow wins 34 of 34 of those rows. The narrowest are `LinearRegression` at 10000 rows and 8 features (1.26x) and both `KernelSVC_RBF` rows at 1000 samples (1.29x); the widest is `GaussianNB` at 100 rows and 32 features (74.4x). Four rows were losses before the coordinate-descent, Cholesky, kernel-cache and support-vector changes in this repo's history, the worst at 0.24x. +One run of that matrix does not settle a row. `benchmarks/scaled_ci_history.json` holds five consecutive runs. Flow wins 30 of the 34 rows in all five. The other four dipped below 1x in at least one run, and all four are at 1000 samples, where a fit finishes in well under a millisecond and the runner moves the measurement by more than the difference being measured: Lasso at 1000 rows and 32 features was recorded at 3.83x and at 0.92x on code that differs in nothing touching Lasso. -The matrix is reported without gating the build, because a shared runner moves a tenth-of-a-millisecond row by more than the row itself. What does gate is `benchmarks/scaled_flow_baseline.json`, a Flow-against-itself comparison. Refresh it from a CI artifact rather than from a developer machine, or the gate will read a fast laptop as the standard and fail every CI run. +Both `KernelSVC_RBF` rows at 1000 samples are in that group for a different reason. They lost consistently, at 0.68x and 0.88x, until a fitted one-vs-one model stopped keeping the whole training set and started keeping its support vectors. They have won every run since, at 1.16x to 1.36x. That is a step change that tracks the commit rather than the runner. + +The ten rows at 10000 samples, the only sizes in the matrix large enough to measure without fighting the noise, win in all five runs, from a median 1.15x on RandomForest at 32 features to 8.06x on GaussianNB at 8. + +`benchmarks/scaled_flow_baseline.json` is a separate Flow-against-itself gate that does fail the build. Refresh it from a CI artifact rather than from a developer machine, or the gate will read a fast laptop as the standard and fail every CI run. Detailed artifacts: diff --git a/benchmarks/README.md b/benchmarks/README.md index 858c81c..19ba316 100644 --- a/benchmarks/README.md +++ b/benchmarks/README.md @@ -195,10 +195,36 @@ The main generated views are: A speedup is `sklearn_ms / flow_ms`. Values above `1x` mean Flow is faster; values below `1x` mean scikit-learn is faster. -The current architecture map shows a useful but non-causal pattern: Flow wins every headline row in all three substrate classes, and the margin varies with the class, from a mean of 26.55x on Python-bound rows down to 3.79x on external-native-bound ones. That ordering is evidence for prioritization. It is not proof that execution substrate alone determines performance. +The current architecture map shows a useful but non-causal pattern: Flow wins every headline row in all three substrate classes, and the margin varies with the class, from a mean of 19.04x on Python-bound rows down to 4.29x on external-native-bound ones. That ordering is evidence for prioritization. It is not proof that execution substrate alone determines performance. Mature BLAS/LAPACK, liblinear, libsvm and other native backends are treated as native competitors. The optimization roadmap deliberately prefers retaining those kernels unless benchmark and parity evidence justify replacement. +## Scaled matrix + +The canonical rows run on iris, digits and diabetes and stop at 1797 samples. `bench_scaled.flow` and `bench_scaled_sklearn.py` run five estimators at 100, 1000 and 10000 rows against 8 and 32 features. CI runs both sides every push and reports the comparison without gating the build. + +One run does not settle a row. At 100 and 1000 samples a fit finishes in well under a millisecond, and the runner moves that by more than the Flow-versus-sklearn difference: Lasso at 1000 rows and 32 features was recorded at 3.83x on one run and 0.92x on another, on code that differs in nothing touching Lasso. + +So the published claim is the spread. [`summarize_scaled_ci.py`](summarize_scaled_ci.py) folds each run's `scaled_comparison.json` into `scaled_ci_history.json`, which records every observation per row along with its minimum, median and maximum: + +``` +gh run download -n scaled-benchmark- -D /tmp/ +python benchmarks/summarize_scaled_ci.py =/tmp//scaled_comparison.json +``` + +Runs merge by id, so adding a new one extends the history and re-adding an existing one replaces it. + +`scaled_flow_baseline.json` is a different artifact for a different job. It compares Flow against its own earlier CI timings and does fail the build, with a 20% relative tolerance and a 0.25 ms absolute floor. + +Each of its rows is the slowest observation across the runs it was built from, so the gate fires when the code is slower than it has ever legitimately been and stays quiet when a run is merely unlucky. Taken from one run it does the opposite: RandomForest at 1000 rows and 8 features has been measured at 1.50, 1.58, 1.63, 2.31, 2.88 and 3.32 ms on identical code, and a baseline taken from the 1.58 run failed the build on the 2.88 one. Rebuild it with the same script: + +``` +python benchmarks/summarize_scaled_ci.py --baseline benchmarks/scaled_flow_baseline.json \ + =/tmp//scaled_comparison.json ... +``` + +Each run directory needs `scaled_flow.json` beside `scaled_comparison.json`. Build it from CI artifacts; a developer machine's numbers would make a fast laptop the standard CI has to meet. + ## Pages publication [`publish_headline_v2.py`](publish_headline_v2.py) validates the committed canonical benchmark and architecture map, then copies the JSON artifacts into `docs/` for the static site. The public benchmark and architecture pages render those artifacts directly instead of embedding hand-maintained timing claims. diff --git a/benchmarks/normalize_published_timings.py b/benchmarks/normalize_published_timings.py deleted file mode 100644 index 28de30d..0000000 --- a/benchmarks/normalize_published_timings.py +++ /dev/null @@ -1,52 +0,0 @@ -#!/usr/bin/env python3 -"""Normalize legacy sklearn timing literals in docs/benchmarks.js to milliseconds. - -The historical BENCH object copied Python ``perf_counter`` durations in seconds -into ``sk_ms``, ``sk_fit`` and ``sk_pred`` fields. Flow fields were already ms. -This transform is intentionally narrow and idempotent; it is used by Pages until -the benchmark site is generated directly from benchmark artifacts. -""" -from __future__ import annotations - -import re -from pathlib import Path - -PATH = Path("docs/benchmarks.js") -MARKER = "// SKLEARN_TIMINGS_NORMALIZED_TO_MS\n" -FIELD_RE = re.compile(r"\b(sk_(?:ms|fit|pred)):\s*(-?\d+(?:\.\d+)?)") - - -def normalize_text(text: str) -> tuple[str, int]: - if MARKER.strip() in text: - return text, 0 - - replacements = 0 - - def repl(match: re.Match[str]) -> str: - nonlocal replacements - field, raw = match.groups() - value = float(raw) * 1000.0 - replacements += 1 - if value == 0: - rendered = "0.000" - elif value < 1: - rendered = f"{value:.3f}" - else: - rendered = f"{value:.3f}".rstrip("0").rstrip(".") - return f"{field}: {rendered}" - - normalized = FIELD_RE.sub(repl, text) - if replacements == 0: - raise RuntimeError("no legacy sklearn timing fields found") - return MARKER + normalized, replacements - - -def main() -> None: - original = PATH.read_text() - normalized, count = normalize_text(original) - PATH.write_text(normalized) - print(f"normalized {count} sklearn timing literals to milliseconds") - - -if __name__ == "__main__": - main() diff --git a/benchmarks/publish_headline_v2.py b/benchmarks/publish_headline_v2.py index 3db26ef..e201e05 100644 --- a/benchmarks/publish_headline_v2.py +++ b/benchmarks/publish_headline_v2.py @@ -14,10 +14,12 @@ ARCH = BENCH / "architecture_performance_map.json" DISPARITY = BENCH / "disparity_report.json" HISTORY = BENCH / "disparity_history.json" +SCALED = BENCH / "scaled_ci_history.json" DOC_RESULT = DOCS / "headline-result-v2.json" DOC_ARCH = DOCS / "architecture-performance-map.json" DOC_DISPARITY = DOCS / "disparity-report.json" DOC_HISTORY = DOCS / "disparity-history.json" +DOC_SCALED = DOCS / "scaled-ci-history.json" def validate_headline(result: dict) -> None: @@ -60,6 +62,20 @@ def validate_history(history: dict, total_rows: int) -> None: raise SystemExit("disparity history snapshot does not cover every canonical row") +def validate_scaled(scaled: dict) -> None: + counts = scaled["counts"] + if counts["runs"] < 2: + raise SystemExit("the scaled matrix needs at least two runs before it says anything") + if counts["rows"] != len(scaled["rows"]): + raise SystemExit("scaled history row count does not match counts.rows") + won = sum(1 for r in scaled["rows"] if r["runs_won"] == r["runs_observed"]) + if won != counts["rows_won_in_every_run"]: + raise SystemExit("scaled history win count disagrees with its own rows") + for row in scaled["rows"]: + if len(row["observations"]) != row["runs_observed"]: + raise SystemExit(f"scaled row {row['algorithm']} has a stale observation list") + + def main() -> int: parser = argparse.ArgumentParser() parser.add_argument("--check", action="store_true") @@ -76,6 +92,11 @@ def main() -> int: validate_disparity(disparity, result["counts"]["total_rows"]) elif not args.check: raise SystemExit("disparity_report.json must be generated before Pages publication") + scaled = json.loads(SCALED.read_text()) if SCALED.exists() else None + if scaled is not None: + validate_scaled(scaled) + elif not args.check: + raise SystemExit("scaled_ci_history.json must be present before Pages publication") if history is not None: validate_history(history, result["counts"]["total_rows"]) elif not args.check: @@ -86,8 +107,9 @@ def main() -> int: shutil.copyfile(ARCH, DOC_ARCH) shutil.copyfile(DISPARITY, DOC_DISPARITY) shutil.copyfile(HISTORY, DOC_HISTORY) + shutil.copyfile(SCALED, DOC_SCALED) counts = result["counts"] - print(f"published evidence: {counts['flow_wins']}/{counts['eligible_comparisons']} Flow wins; {disparity['counts']['rows_with_tracked_disparity']} rows with tracked disparities; {len(history['snapshots'])} history snapshots") + print(f"published evidence: {counts['flow_wins']}/{counts['eligible_comparisons']} Flow wins; {disparity['counts']['rows_with_tracked_disparity']} rows with tracked disparities; {len(history['snapshots'])} history snapshots; scaled matrix over {scaled['counts']['runs']} runs") else: print("canonical benchmark, architecture, disparity, and available history evidence are internally consistent") return 0 diff --git a/benchmarks/scaled_ci_history.json b/benchmarks/scaled_ci_history.json new file mode 100644 index 0000000..92fdb74 --- /dev/null +++ b/benchmarks/scaled_ci_history.json @@ -0,0 +1,617 @@ +{ + "schema_version": 1, + "source": "GitHub Actions scaled-benchmark artifacts, godofecht/flow-scikit", + "runs": [ + { + "run_id": "36156144714", + "rows_compared": 34, + "flow_wins": 32 + }, + { + "run_id": "36164983993", + "rows_compared": 34, + "flow_wins": 32 + }, + { + "run_id": "36166087543", + "rows_compared": 34, + "flow_wins": 34 + }, + { + "run_id": "36169006565", + "rows_compared": 34, + "flow_wins": 34 + }, + { + "run_id": "36173242643", + "rows_compared": 34, + "flow_wins": 32 + } + ], + "counts": { + "runs": 5, + "rows": 34, + "rows_won_in_every_run": 30, + "rows_lost_in_at_least_one_run": 4 + }, + "rows": [ + { + "algorithm": "GaussianNB", + "rows": 10000, + "features": 32, + "observations": [ + 2.6841, + 2.566, + 2.8451, + 2.7577, + 2.8422 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 2.566, + "median_speedup": 2.7577, + "max_speedup": 2.8451 + }, + { + "algorithm": "GaussianNB", + "rows": 10000, + "features": 8, + "observations": [ + 8.0605, + 7.9017, + 8.1038, + 8.2239, + 7.3528 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 7.3528, + "median_speedup": 8.0605, + "max_speedup": 8.2239 + }, + { + "algorithm": "GaussianNB", + "rows": 1000, + "features": 32, + "observations": [ + 11.4624, + 11.3796, + 11.6477, + 11.9571, + 9.6402 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 9.6402, + "median_speedup": 11.4624, + "max_speedup": 11.9571 + }, + { + "algorithm": "GaussianNB", + "rows": 1000, + "features": 8, + "observations": [ + 35.2834, + 34.1568, + 34.5062, + 35.4365, + 28.3833 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 28.3833, + "median_speedup": 34.5062, + "max_speedup": 35.4365 + }, + { + "algorithm": "GaussianNB", + "rows": 100, + "features": 32, + "observations": [ + 72.8536, + 72.9485, + 74.41, + 74.0101, + 50.7279 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 50.7279, + "median_speedup": 72.9485, + "max_speedup": 74.41 + }, + { + "algorithm": "GaussianNB", + "rows": 100, + "features": 8, + "observations": [ + 64.8624, + 66.4084, + 66.7392, + 67.7841, + 36.6261 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 36.6261, + "median_speedup": 66.4084, + "max_speedup": 67.7841 + }, + { + "algorithm": "KMeans", + "rows": 10000, + "features": 32, + "observations": [ + 1.8732, + 1.8738, + 1.648, + 1.7541, + 1.2107 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 1.2107, + "median_speedup": 1.7541, + "max_speedup": 1.8738 + }, + { + "algorithm": "KMeans", + "rows": 10000, + "features": 8, + "observations": [ + 1.9508, + 1.8724, + 1.8299, + 1.8713, + 1.7668 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 1.7668, + "median_speedup": 1.8713, + "max_speedup": 1.9508 + }, + { + "algorithm": "KMeans", + "rows": 1000, + "features": 32, + "observations": [ + 3.2426, + 3.1766, + 3.1882, + 3.2591, + 2.9195 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 2.9195, + "median_speedup": 3.1882, + "max_speedup": 3.2591 + }, + { + "algorithm": "KMeans", + "rows": 1000, + "features": 8, + "observations": [ + 8.359, + 8.1273, + 8.3222, + 8.54, + 7.4468 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 7.4468, + "median_speedup": 8.3222, + "max_speedup": 8.54 + }, + { + "algorithm": "KMeans", + "rows": 100, + "features": 32, + "observations": [ + 31.1124, + 30.4329, + 30.2153, + 31.3033, + 20.3215 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 20.3215, + "median_speedup": 30.4329, + "max_speedup": 31.3033 + }, + { + "algorithm": "KMeans", + "rows": 100, + "features": 8, + "observations": [ + 54.2592, + 56.2389, + 54.48, + 56.1383, + 34.3141 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 34.3141, + "median_speedup": 54.48, + "max_speedup": 56.2389 + }, + { + "algorithm": "KernelSVC_RBF", + "rows": 1000, + "features": 32, + "observations": [ + 0.8801, + 0.8276, + 1.288, + 1.3407, + 1.3633 + ], + "runs_observed": 5, + "runs_won": 3, + "min_speedup": 0.8276, + "median_speedup": 1.288, + "max_speedup": 1.3633 + }, + { + "algorithm": "KernelSVC_RBF", + "rows": 1000, + "features": 8, + "observations": [ + 0.6781, + 0.6648, + 1.2883, + 1.2969, + 1.1573 + ], + "runs_observed": 5, + "runs_won": 3, + "min_speedup": 0.6648, + "median_speedup": 1.1573, + "max_speedup": 1.2969 + }, + { + "algorithm": "KernelSVC_RBF", + "rows": 100, + "features": 32, + "observations": [ + 6.0569, + 5.5135, + 5.8496, + 6.0266, + 5.5063 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 5.5063, + "median_speedup": 5.8496, + "max_speedup": 6.0569 + }, + { + "algorithm": "KernelSVC_RBF", + "rows": 100, + "features": 8, + "observations": [ + 4.3978, + 4.3108, + 4.3423, + 4.5396, + 1.3945 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 1.3945, + "median_speedup": 4.3423, + "max_speedup": 4.5396 + }, + { + "algorithm": "Lasso", + "rows": 10000, + "features": 32, + "observations": [ + 1.4031, + 1.5449, + 1.4837, + 1.481, + 1.7445 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 1.4031, + "median_speedup": 1.4837, + "max_speedup": 1.7445 + }, + { + "algorithm": "Lasso", + "rows": 10000, + "features": 8, + "observations": [ + 3.7841, + 4.2454, + 3.9141, + 4.1963, + 3.316 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 3.316, + "median_speedup": 3.9141, + "max_speedup": 4.2454 + }, + { + "algorithm": "Lasso", + "rows": 1000, + "features": 32, + "observations": [ + 3.8283, + 4.072, + 1.5077, + 1.466, + 0.921 + ], + "runs_observed": 5, + "runs_won": 4, + "min_speedup": 0.921, + "median_speedup": 1.5077, + "max_speedup": 4.072 + }, + { + "algorithm": "Lasso", + "rows": 1000, + "features": 8, + "observations": [ + 9.9444, + 9.8279, + 9.4535, + 9.8622, + 7.5214 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 7.5214, + "median_speedup": 9.8279, + "max_speedup": 9.9444 + }, + { + "algorithm": "Lasso", + "rows": 100, + "features": 32, + "observations": [ + 5.7589, + 4.9942, + 5.6982, + 5.8756, + 3.9334 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 3.9334, + "median_speedup": 5.6982, + "max_speedup": 5.8756 + }, + { + "algorithm": "Lasso", + "rows": 100, + "features": 8, + "observations": [ + 23.9078, + 32.2604, + 31.4913, + 37.3463, + 23.1179 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 23.1179, + "median_speedup": 31.4913, + "max_speedup": 37.3463 + }, + { + "algorithm": "LinearRegression", + "rows": 10000, + "features": 32, + "observations": [ + 1.4966, + 1.6501, + 1.6043, + 1.6174, + 1.32 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 1.32, + "median_speedup": 1.6043, + "max_speedup": 1.6501 + }, + { + "algorithm": "LinearRegression", + "rows": 10000, + "features": 8, + "observations": [ + 1.4231, + 1.5147, + 1.2582, + 1.3045, + 1.1102 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 1.1102, + "median_speedup": 1.3045, + "max_speedup": 1.5147 + }, + { + "algorithm": "LinearRegression", + "rows": 1000, + "features": 32, + "observations": [ + 2.1822, + 1.7896, + 1.7991, + 1.8715, + 0.8973 + ], + "runs_observed": 5, + "runs_won": 4, + "min_speedup": 0.8973, + "median_speedup": 1.7991, + "max_speedup": 2.1822 + }, + { + "algorithm": "LinearRegression", + "rows": 1000, + "features": 8, + "observations": [ + 6.2297, + 6.4194, + 5.4502, + 6.6422, + 5.1362 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 5.1362, + "median_speedup": 6.2297, + "max_speedup": 6.6422 + }, + { + "algorithm": "LinearRegression", + "rows": 100, + "features": 32, + "observations": [ + 11.4533, + 12.5131, + 12.3174, + 12.5428, + 8.243 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 8.243, + "median_speedup": 12.3174, + "max_speedup": 12.5428 + }, + { + "algorithm": "LinearRegression", + "rows": 100, + "features": 8, + "observations": [ + 17.2824, + 21.6447, + 21.9858, + 22.7798, + 11.038 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 11.038, + "median_speedup": 21.6447, + "max_speedup": 22.7798 + }, + { + "algorithm": "RandomForest", + "rows": 10000, + "features": 32, + "observations": [ + 1.1454, + 1.2595, + 1.4667, + 1.1394, + 1.0324 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 1.0324, + "median_speedup": 1.1454, + "max_speedup": 1.4667 + }, + { + "algorithm": "RandomForest", + "rows": 10000, + "features": 8, + "observations": [ + 1.7194, + 1.8356, + 2.0071, + 1.9202, + 1.6407 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 1.6407, + "median_speedup": 1.8356, + "max_speedup": 2.0071 + }, + { + "algorithm": "RandomForest", + "rows": 1000, + "features": 32, + "observations": [ + 2.9509, + 3.862, + 3.7723, + 3.1295, + 2.674 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 2.674, + "median_speedup": 3.1295, + "max_speedup": 3.862 + }, + { + "algorithm": "RandomForest", + "rows": 1000, + "features": 8, + "observations": [ + 4.2202, + 6.0155, + 8.6103, + 8.4982, + 5.2407 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 4.2202, + "median_speedup": 6.0155, + "max_speedup": 8.6103 + }, + { + "algorithm": "RandomForest", + "rows": 100, + "features": 32, + "observations": [ + 11.3405, + 10.897, + 9.9421, + 10.6379, + 8.6659 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 8.6659, + "median_speedup": 10.6379, + "max_speedup": 11.3405 + }, + { + "algorithm": "RandomForest", + "rows": 100, + "features": 8, + "observations": [ + 11.0781, + 10.7355, + 10.8695, + 11.2911, + 9.372 + ], + "runs_observed": 5, + "runs_won": 5, + "min_speedup": 9.372, + "median_speedup": 10.8695, + "max_speedup": 11.2911 + } + ] +} diff --git a/benchmarks/scaled_flow_baseline.json b/benchmarks/scaled_flow_baseline.json index 4b7b4de..14de1e0 100644 --- a/benchmarks/scaled_flow_baseline.json +++ b/benchmarks/scaled_flow_baseline.json @@ -1,40 +1,380 @@ { - "schema_version": 1, - "source": "GitHub Actions run 36166087543, artifact scaled-benchmark-36166087543", + "schema_version": 2, + "source": "slowest observation per row across GitHub Actions runs 36156144714, 36164983993, 36166087543, 36169006565, 36173242643", "rows": [ - {"implementation":"flow","algorithm":"LinearRegression","rows":100,"features":8,"fit_ms":0.027541,"pred_ms":0.004419,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"Lasso","rows":100,"features":8,"fit_ms":0.02145,"pred_ms":3e-05,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"RandomForest","rows":100,"features":8,"fit_ms":1.130270958,"pred_ms":0.005771,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"GaussianNB","rows":100,"features":8,"fit_ms":0.004058,"pred_ms":0.01081,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"KMeans","rows":100,"features":8,"fit_ms":0.028463,"pred_ms":0.000901,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"KernelSVC_RBF","rows":100,"features":8,"fit_ms":0.088125996,"pred_ms":0.147138,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"LinearRegression","rows":100,"features":32,"fit_ms":0.067396998,"pred_ms":0.000721,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"Lasso","rows":100,"features":32,"fit_ms":0.135665998,"pred_ms":3e-05,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"RandomForest","rows":100,"features":32,"fit_ms":1.247370958,"pred_ms":0.006071,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"GaussianNB","rows":100,"features":32,"fit_ms":0.009487,"pred_ms":0.004068,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"KMeans","rows":100,"features":32,"fit_ms":0.052237999,"pred_ms":0.001452,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"KernelSVC_RBF","rows":100,"features":32,"fit_ms":0.169199005,"pred_ms":0.019848,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"LinearRegression","rows":1000,"features":8,"fit_ms":0.141497001,"pred_ms":0.001603,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"Lasso","rows":1000,"features":8,"fit_ms":0.076274,"pred_ms":3e-05,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"RandomForest","rows":1000,"features":8,"fit_ms":1.579267025,"pred_ms":0.024536001,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"GaussianNB","rows":1000,"features":8,"fit_ms":0.025467999,"pred_ms":0.007313,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"KMeans","rows":1000,"features":8,"fit_ms":0.306737989,"pred_ms":0.003877,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"KernelSVC_RBF","rows":1000,"features":8,"fit_ms":2.592648029,"pred_ms":0.262003988,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"LinearRegression","rows":1000,"features":32,"fit_ms":0.776654005,"pred_ms":0.002234,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"Lasso","rows":1000,"features":32,"fit_ms":0.65140897,"pred_ms":5e-05,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"RandomForest","rows":1000,"features":32,"fit_ms":4.035840034,"pred_ms":0.027171001,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"GaussianNB","rows":1000,"features":32,"fit_ms":0.077735998,"pred_ms":0.023604,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"KMeans","rows":1000,"features":32,"fit_ms":1.026185036,"pred_ms":0.011852,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"KernelSVC_RBF","rows":1000,"features":32,"fit_ms":2.995167017,"pred_ms":0.275148004,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"LinearRegression","rows":10000,"features":8,"fit_ms":1.151620984,"pred_ms":0.009177,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"Lasso","rows":10000,"features":8,"fit_ms":0.354048014,"pred_ms":3e-05,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"RandomForest","rows":10000,"features":8,"fit_ms":12.04566288,"pred_ms":0.204225004,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"GaussianNB","rows":10000,"features":8,"fit_ms":0.231436998,"pred_ms":0.063951001,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"KMeans","rows":10000,"features":8,"fit_ms":2.740686893,"pred_ms":0.035445999,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"LinearRegression","rows":10000,"features":32,"fit_ms":2.645308018,"pred_ms":0.0105,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"Lasso","rows":10000,"features":32,"fit_ms":2.626912117,"pred_ms":5.1e-05,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"RandomForest","rows":10000,"features":32,"fit_ms":27.210836411,"pred_ms":0.220786005,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"GaussianNB","rows":10000,"features":32,"fit_ms":0.752909005,"pred_ms":0.224684,"timing_unit":"ms","status":"ok"}, - {"implementation":"flow","algorithm":"KMeans","rows":10000,"features":32,"fit_ms":3.527801037,"pred_ms":0.083888002,"timing_unit":"ms","status":"ok"} + { + "implementation": "flow", + "algorithm": "GaussianNB", + "rows": 10000, + "features": 32, + "fit_ms": 0.816300988, + "pred_ms": 0.238667995, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "GaussianNB", + "rows": 10000, + "features": 8, + "fit_ms": 0.238407001, + "pred_ms": 0.066775002, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "GaussianNB", + "rows": 1000, + "features": 32, + "fit_ms": 0.081202, + "pred_ms": 0.024816999, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "GaussianNB", + "rows": 1000, + "features": 8, + "fit_ms": 0.026039001, + "pred_ms": 0.007674, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "GaussianNB", + "rows": 100, + "features": 32, + "fit_ms": 0.009688, + "pred_ms": 0.004518, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "GaussianNB", + "rows": 100, + "features": 8, + "fit_ms": 0.004479, + "pred_ms": 0.011592, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "KMeans", + "rows": 10000, + "features": 32, + "fit_ms": 3.527801037, + "pred_ms": 0.084668003, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "KMeans", + "rows": 10000, + "features": 8, + "fit_ms": 2.746421099, + "pred_ms": 0.041097, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "KMeans", + "rows": 1000, + "features": 32, + "fit_ms": 1.065410972, + "pred_ms": 0.011852, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "KMeans", + "rows": 1000, + "features": 8, + "fit_ms": 0.325080991, + "pred_ms": 0.004468, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "KMeans", + "rows": 100, + "features": 32, + "fit_ms": 0.05302, + "pred_ms": 0.001463, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "KMeans", + "rows": 100, + "features": 8, + "fit_ms": 0.029966, + "pred_ms": 0.000991, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "KernelSVC_RBF", + "rows": 1000, + "features": 32, + "fit_ms": 3.011528969, + "pred_ms": 2.421480894, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "KernelSVC_RBF", + "rows": 1000, + "features": 8, + "fit_ms": 2.6064291, + "pred_ms": 3.004882097, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "KernelSVC_RBF", + "rows": 100, + "features": 32, + "fit_ms": 0.174187005, + "pred_ms": 0.031509001, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "KernelSVC_RBF", + "rows": 100, + "features": 8, + "fit_ms": 0.088125996, + "pred_ms": 0.334419012, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "Lasso", + "rows": 10000, + "features": 32, + "fit_ms": 2.927758932, + "pred_ms": 6e-05, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "Lasso", + "rows": 10000, + "features": 8, + "fit_ms": 0.376915991, + "pred_ms": 3.4e-05, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "Lasso", + "rows": 1000, + "features": 32, + "fit_ms": 0.658141017, + "pred_ms": 5e-05, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "Lasso", + "rows": 1000, + "features": 8, + "fit_ms": 0.076453, + "pred_ms": 3.1e-05, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "Lasso", + "rows": 100, + "features": 32, + "fit_ms": 0.157155007, + "pred_ms": 3.2e-05, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "Lasso", + "rows": 100, + "features": 8, + "fit_ms": 0.029545, + "pred_ms": 5e-05, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "LinearRegression", + "rows": 10000, + "features": 32, + "fit_ms": 2.874648094, + "pred_ms": 0.017754, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "LinearRegression", + "rows": 10000, + "features": 8, + "fit_ms": 1.151620984, + "pred_ms": 0.009267, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "LinearRegression", + "rows": 1000, + "features": 32, + "fit_ms": 0.974076986, + "pred_ms": 0.002234, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "LinearRegression", + "rows": 1000, + "features": 8, + "fit_ms": 0.141497001, + "pred_ms": 0.001603, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "LinearRegression", + "rows": 100, + "features": 32, + "fit_ms": 0.075530998, + "pred_ms": 0.001424, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "LinearRegression", + "rows": 100, + "features": 8, + "fit_ms": 0.036667999, + "pred_ms": 0.006522, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "RandomForest", + "rows": 10000, + "features": 32, + "fit_ms": 32.339641571, + "pred_ms": 0.223608002, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "RandomForest", + "rows": 10000, + "features": 8, + "fit_ms": 14.628679276, + "pred_ms": 0.204684004, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "RandomForest", + "rows": 1000, + "features": 32, + "fit_ms": 5.331896782, + "pred_ms": 0.027331, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "RandomForest", + "rows": 1000, + "features": 8, + "fit_ms": 3.31953001, + "pred_ms": 0.025397999, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "RandomForest", + "rows": 100, + "features": 32, + "fit_ms": 1.247370958, + "pred_ms": 0.006071, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + }, + { + "implementation": "flow", + "algorithm": "RandomForest", + "rows": 100, + "features": 8, + "fit_ms": 1.163939953, + "pred_ms": 0.006102, + "runs_observed": 5, + "timing_unit": "ms", + "status": "ok" + } ] } diff --git a/benchmarks/summarize_scaled_ci.py b/benchmarks/summarize_scaled_ci.py new file mode 100644 index 0000000..d64f813 --- /dev/null +++ b/benchmarks/summarize_scaled_ci.py @@ -0,0 +1,189 @@ +#!/usr/bin/env python3 +"""Collect the scaled benchmark matrix across several CI runs into one artifact. + +A single run of the scaled matrix is not enough to say a row is won. The rows +at 100 and 1000 samples finish in well under a millisecond, and a shared CI +runner moves them by more than the Flow-versus-sklearn difference: on five +consecutive runs of the same code, Lasso at 1000 rows and 32 features was +measured at 3.83x and at 0.92x. + +So the published claim is the spread across runs rather than one run's number. +Feed this the scaled_comparison.json from each run: + + gh run download -n scaled-benchmark- -D /tmp/ + python benchmarks/summarize_scaled_ci.py =/tmp//scaled_comparison.json + +Runs merge by id, so re-running with an id already present replaces it and +adding a new one extends the history. +""" +from __future__ import annotations + +import argparse +import json +import statistics +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[1] +HISTORY = ROOT / "benchmarks" / "scaled_ci_history.json" +BASELINE = ROOT / "benchmarks" / "scaled_flow_baseline.json" + + +def row_key(row: dict) -> str: + return f"{row['algorithm']}/{row['rows']}x{row['features']}" + + +def load_flow(path: Path) -> dict[str, dict]: + """Flow's own fit/pred milliseconds from a run, for the self-regression gate.""" + payload = json.loads(path.read_text()) + return { + row_key(r): {"fit_ms": float(r["fit_ms"]), "pred_ms": float(r["pred_ms"])} + for r in payload["rows"] + if r.get("status") == "ok" + } + + +def write_baseline(flow_runs: dict[str, dict[str, dict]], out: Path) -> int: + """The slowest observation per row, across every run fed in. + + A regression gate should fire when the code is slower than it has ever + legitimately been, and stay quiet when a run is merely unlucky. Built from + a single run it does the opposite: RandomForest at 1000 rows and 8 features + was measured at 1.50, 1.58, 1.63, 2.31, 2.88 and 3.32 ms on identical code, + and a baseline taken from the 1.58 run failed the build on the 2.88 one. + Each row here is the slowest observation across the runs fed in. + """ + keys = sorted({k for run in flow_runs.values() for k in run}) + rows = [] + for key in keys: + seen = [run[key] for run in flow_runs.values() if key in run] + algorithm, shape = key.split("/", 1) + samples, features = shape.split("x", 1) + rows.append( + { + "implementation": "flow", + "algorithm": algorithm, + "rows": int(samples), + "features": int(features), + "fit_ms": max(v["fit_ms"] for v in seen), + "pred_ms": max(v["pred_ms"] for v in seen), + "runs_observed": len(seen), + "timing_unit": "ms", + "status": "ok", + } + ) + payload = { + "schema_version": 2, + "source": "slowest observation per row across GitHub Actions runs " + + ", ".join(sorted(flow_runs)), + "rows": rows, + } + out.write_text(json.dumps(payload, indent=2) + "\n") + return len(rows) + + +def load_run(path: Path) -> dict[str, float]: + payload = json.loads(path.read_text()) + return { + row_key(r): float(r["speedup"]) + for r in payload["rows"] + if r.get("status") == "ok" and "speedup" in r + } + + +def summarize(runs: dict[str, dict[str, float]]) -> dict: + keys = sorted({k for run in runs.values() for k in run}) + order = sorted(runs) + rows = [] + for key in keys: + observations = [runs[r][key] for r in order if key in runs[r]] + if not observations: + continue + algorithm, shape = key.split("/", 1) + samples, features = shape.split("x", 1) + rows.append( + { + "algorithm": algorithm, + "rows": int(samples), + "features": int(features), + "observations": [round(v, 4) for v in observations], + "runs_observed": len(observations), + "runs_won": sum(1 for v in observations if v >= 1.0), + "min_speedup": round(min(observations), 4), + "median_speedup": round(statistics.median(observations), 4), + "max_speedup": round(max(observations), 4), + } + ) + per_run = [ + { + "run_id": r, + "rows_compared": len(runs[r]), + "flow_wins": sum(1 for v in runs[r].values() if v >= 1.0), + } + for r in order + ] + won_every_run = sum(1 for r in rows if r["runs_won"] == r["runs_observed"]) + return { + "schema_version": 1, + "source": "GitHub Actions scaled-benchmark artifacts, godofecht/flow-scikit", + "runs": per_run, + "counts": { + "runs": len(order), + "rows": len(rows), + "rows_won_in_every_run": won_every_run, + "rows_lost_in_at_least_one_run": len(rows) - won_every_run, + }, + "rows": rows, + } + + +def main() -> int: + ap = argparse.ArgumentParser() + ap.add_argument("runs", nargs="+", metavar="RUN_ID=PATH") + ap.add_argument("--out", type=Path, default=HISTORY) + ap.add_argument( + "--baseline", + type=Path, + default=None, + help="also rewrite the self-regression baseline from the slowest observation per row; " + "each RUN_ID=PATH directory must hold scaled_flow.json beside scaled_comparison.json", + ) + args = ap.parse_args() + + existing: dict[str, dict[str, float]] = {} + if args.out.exists(): + prior = json.loads(args.out.read_text()) + for row in prior.get("rows", []): + key = f"{row['algorithm']}/{row['rows']}x{row['features']}" + for run, value in zip([r["run_id"] for r in prior["runs"]], row["observations"]): + existing.setdefault(run, {})[key] = value + + for spec in args.runs: + if "=" not in spec: + raise SystemExit(f"expected RUN_ID=PATH, got {spec!r}") + run_id, path = spec.split("=", 1) + existing[run_id] = load_run(Path(path)) + + if args.baseline is not None: + flow_runs: dict[str, dict[str, dict]] = {} + for spec in args.runs: + run_id, path = spec.split("=", 1) + flow_path = Path(path).parent / "scaled_flow.json" + if not flow_path.exists(): + raise SystemExit(f"{flow_path} is missing; the baseline needs Flow's own timings") + flow_runs[run_id] = load_flow(flow_path) + n = write_baseline(flow_runs, args.baseline) + print(f"wrote {args.baseline}: {n} rows from {len(flow_runs)} runs") + + summary = summarize(existing) + args.out.write_text(json.dumps(summary, indent=2) + "\n") + counts = summary["counts"] + print( + f"wrote {args.out}: {counts['rows']} rows over {counts['runs']} runs; " + f"{counts['rows_won_in_every_run']} won in every run, " + f"{counts['rows_lost_in_at_least_one_run']} dipped below 1x at least once" + ) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/benchmarks/validate_published_claims.py b/benchmarks/validate_published_claims.py index 98088ff..9044d16 100644 --- a/benchmarks/validate_published_claims.py +++ b/benchmarks/validate_published_claims.py @@ -26,6 +26,7 @@ ROOT = Path(__file__).resolve().parents[1] RESULT = ROOT / "benchmarks" / "headline_result_v2.json" +ARCH = ROOT / "benchmarks" / "architecture_performance_map.json" # (path, regex with one group per count, builder for the replacement text) TARGETS = [ @@ -57,6 +58,40 @@ ), ] +# The same drift, from the other artifact: both READMEs quote the mean Flow +# speedup per sklearn execution substrate, and those means move on every +# freeze. They had gone stale by 2x on one group before this was added. +SUBSTRATE_TARGETS = [ + ( + ROOT / "README.md", + re.compile( + r"at a mean of ([\d.]+)x on Python-bound rows, ([\d.]+)x on mixed rows " + r"and ([\d.]+)x on external-native-bound rows" + ), + lambda m: ( + f"at a mean of {m['python-bound']:.2f}x on Python-bound rows, " + f"{m['mixed']:.2f}x on mixed rows and " + f"{m['external-native-bound']:.2f}x on external-native-bound rows" + ), + ), + ( + ROOT / "benchmarks" / "README.md", + re.compile( + r"from a mean of ([\d.]+)x on Python-bound rows down to ([\d.]+)x on " + r"external-native-bound ones" + ), + lambda m: ( + f"from a mean of {m['python-bound']:.2f}x on Python-bound rows down to " + f"{m['external-native-bound']:.2f}x on external-native-bound ones" + ), + ), +] + + +def substrate_means() -> dict[str, float]: + groups = json.loads(ARCH.read_text())["speedup_by_execution_substrate"] + return {g["execution_class"]: float(g["mean_flow_speedup"]) for g in groups} + def main() -> int: ap = argparse.ArgumentParser() @@ -64,9 +99,12 @@ def main() -> int: args = ap.parse_args() counts = json.loads(RESULT.read_text())["counts"] + means = substrate_means() drifted: list[str] = [] - for path, pattern, build in TARGETS: + for path, pattern, build in TARGETS + [ + (p, r, (lambda b: lambda _c: b(means))(build)) for p, r, build in SUBSTRATE_TARGETS + ]: text = path.read_text() match = pattern.search(text) rel = path.relative_to(ROOT) @@ -83,7 +121,7 @@ def main() -> int: drifted.append(f"{rel}\n committed: {match.group(0)}\n artifact: {expected}") if drifted: - print("Published prose disagrees with benchmarks/headline_result_v2.json:\n") + print("Published prose disagrees with the committed artifacts:\n") for d in drifted: print(" " + d + "\n") print("Run: python benchmarks/validate_published_claims.py --fix") diff --git a/benchmarks/validate_published_timings.py b/benchmarks/validate_published_timings.py deleted file mode 100644 index 23a38b2..0000000 --- a/benchmarks/validate_published_timings.py +++ /dev/null @@ -1,45 +0,0 @@ -#!/usr/bin/env python3 -"""Fail if the published dataset benchmark contains legacy second-valued sklearn timings.""" -from __future__ import annotations - -import re -from pathlib import Path - -PATH = Path("docs/benchmarks.js") -MARKER = "SKLEARN_TIMINGS_NORMALIZED_TO_MS" -ROW_RE = re.compile(r"\{\s*algo:\s*\"([^\"]+)\"(?P[^\n]+?)\}") -FIELD_RE = re.compile(r"\b(sk_ms|fl_ms):\s*(-?\d+(?:\.\d+)?)") - - -def main() -> None: - text = PATH.read_text() - if MARKER not in text: - raise SystemExit("published benchmark timings were not normalized to milliseconds") - - usable = 0 - flow_wins = 0 - rows = 0 - for match in ROW_RE.finditer(text): - body = match.group("body") - values = {name: float(value) for name, value in FIELD_RE.findall(body)} - if "sk_ms" not in values or "fl_ms" not in values: - continue - rows += 1 - if values["sk_ms"] > 0: - usable += 1 - if values["fl_ms"] < values["sk_ms"]: - flow_wins += 1 - - if rows != 19: - raise SystemExit(f"expected 19 dataset benchmark rows, found {rows}") - if (flow_wins, usable) != (10, 17): - raise SystemExit( - "legacy published headline no longer reproduces 10/17 after unit normalization: " - f"got {flow_wins}/{usable}" - ) - - print(f"published timing audit: {rows} rows, legacy headline {flow_wins}/{usable}") - - -if __name__ == "__main__": - main() diff --git a/docs/architecture.html b/docs/architecture.html index 0f3a735..2e3ebeb 100644 --- a/docs/architecture.html +++ b/docs/architecture.html @@ -1 +1 @@ -Execution architecture — flow-scikit

execution architecture / generated appraisal

“Python versus compiled” is the wrong comparison.

scikit-learn is layered. Its public API is Python, but estimator work may execute in Python orchestration, NumPy/SciPy, BLAS/LAPACK, sklearn-owned compiled code or mature external native libraries. Flow-scikit now maps those layers, profiles representative workloads and joins that evidence to the canonical benchmark.

estimator operations—

Generated from the pinned sklearn public estimator surface.

runtime profiles—

Representative fit/inference attribution rows.

native hotspots—

Compiled/mixed paths with explicit retain/replace dispositions.

whole-estimator experiments—

Experiments separating public API overhead from underlying kernels.

Evidence status: this is no longer a provisional static table. CI regenerates the estimator-operation inventory against the pinned sklearn version, checks estimator-surface drift, profiles representative workloads, ranks optimization opportunities and joins all 19 canonical benchmark rows to substrate and speedup evidence.

substrate signal

Where Flow wins is not random.

The current grouped result is descriptive evidence, not causal proof. It is nevertheless useful enough to prioritize engineering: Flow performs best on rows classified Python-bound, is mixed on boundary-heavy workloads, and currently loses every external-native-bound headline comparison.

sklearn fit substraterowsmean Flow speedupFlow win fraction

all headline rows

Every benchmark result is attached to an execution hypothesis.

AlgorithmDatasetSubstrateParityWinnerFlow speedupPython self share

roadmap logic

Replace selectively. Reuse deliberately.

Python-bound and boundary-heavy operations can rank highly when profiling shows orchestration cost, allocation pressure or repeated crossings. sklearn-owned compiled hotspots become direct Flow replacement targets when parity and benchmark evidence justify it. BLAS/LAPACK, liblinear, libsvm and other mature native kernels receive a reuse bias unless measurements show a real reason to replace them.

The generated optimization roadmap, native hotspot audit and full estimator inventory are the detailed source artifacts.

defensible conclusion

Whole-estimator compilation is a measurable hypothesis now.

Flow-scikit no longer needs to infer opportunity from whether sklearn source files happen to be Python or Cython. The repository can connect a win or loss to the substrate underneath the public API, the Python-visible runtime share, native crossings, parity status and current Flow implementation. That makes future rewrite decisions falsifiable rather than rhetorical.

Inspect the canonical benchmark →

\ No newline at end of file +Execution architecture, flow-scikit

execution architecture / generated appraisal

"Python versus compiled" is the wrong comparison.

scikit-learn is layered. Its public API is Python, but estimator work may execute in Python orchestration, NumPy/SciPy, BLAS/LAPACK, sklearn-owned compiled code or mature external native libraries. Flow-scikit now maps those layers, profiles representative workloads and joins that evidence to the canonical benchmark.

estimator operations...

Generated from the pinned sklearn public estimator surface.

runtime profiles...

Representative fit/inference attribution rows.

native hotspots...

Compiled/mixed paths with explicit retain/replace dispositions.

whole-estimator experiments...

Experiments separating public API overhead from underlying kernels.

Evidence status: this is no longer a provisional static table. CI regenerates the estimator-operation inventory against the pinned sklearn version, checks estimator-surface drift, profiles representative workloads, ranks optimization opportunities and joins all 19 canonical benchmark rows to substrate and speedup evidence.

substrate signal

Where Flow wins follows the substrate.

The current grouped result is descriptive evidence. It does not establish cause. It is nevertheless useful enough to prioritize engineering: Flow now wins every canonical row in all three substrate classes, and the margin still follows the class: largest on rows classified Python-bound, smaller on boundary-heavy workloads, smallest where sklearn calls an external native library. The table below carries the current means. Earlier, while the Flow side was compiled unoptimized, every external-native-bound row lost, which read as a ceiling rather than as a build setting.

sklearn fit substraterowsmean Flow speedupFlow win fraction

all headline rows

Every benchmark result is attached to an execution hypothesis.

AlgorithmDatasetSubstrateParityWinnerFlow speedupPython self share

roadmap logic

Replace selectively. Reuse deliberately.

Python-bound and boundary-heavy operations can rank highly when profiling shows orchestration cost, allocation pressure or repeated crossings. sklearn-owned compiled hotspots become direct Flow replacement targets when parity and benchmark evidence justify it. BLAS/LAPACK, liblinear, libsvm and other mature native kernels receive a reuse bias unless measurements show a real reason to replace them.

The generated optimization roadmap, native hotspot audit and full estimator inventory are the detailed source artifacts.

defensible conclusion

Whole-estimator compilation is a measurable hypothesis now.

Flow-scikit no longer needs to infer opportunity from whether sklearn source files happen to be Python or Cython. The repository can connect a win or loss to the substrate underneath the public API, the Python-visible runtime share, native crossings, parity status and current Flow implementation. That makes future rewrite decisions falsifiable rather than rhetorical.

Inspect the canonical benchmark →

\ No newline at end of file diff --git a/docs/benchmark-corrections.js b/docs/benchmark-corrections.js deleted file mode 100644 index 19c4f54..0000000 --- a/docs/benchmark-corrections.js +++ /dev/null @@ -1,65 +0,0 @@ -// Timing-unit compatibility and presentation helpers for the benchmark page. -// -// The committed historical BENCH dataset stores sklearn perf_counter durations -// in seconds even though the fields are named *_ms. Pages deployment runs -// benchmarks/normalize_published_timings.py, which converts those literals to -// milliseconds before publishing. When the docs are opened directly from a -// checkout, however, the legacy second-valued literals are still present. -// -// Therefore this browser compatibility layer must be idempotent: normalize a -// legacy checkout exactly once, but never multiply an already-normalized Pages -// build by another 1000x. - -(function ensureBenchmarkMilliseconds() { - const datasets = [BENCH.iris, BENCH.digits, BENCH.diabetes]; - const rows = datasets.flat(); - - // In the committed legacy dataset every sklearn total is < 1 because those - // values are seconds; after normalization the same dataset has values up to - // hundreds of milliseconds. This discriminator is deliberately scoped to - // the historical 19-row dataset and should disappear once generated - // benchmark artifacts replace the legacy literals (#175/#176). - const maxSkTotal = Math.max(...rows.map(row => row.sk_ms)); - const legacySeconds = maxSkTotal > 0 && maxSkTotal < 1; - - if (!legacySeconds) return; - - for (const row of rows) { - row.sk_ms *= 1000.0; - row.sk_fit *= 1000.0; - row.sk_pred *= 1000.0; - } -})(); - -function formatPythonResolution(value) { - return value === 0 ? "<0.5" : value.toFixed(2); -} - -document.addEventListener("DOMContentLoaded", () => { - const accuracyRows = [ - ...BENCH.iris, - ...BENCH.digits, - ...BENCH.diabetes - ]; - - const accuracyTableRows = document.querySelectorAll("#accuracy-table-body tr"); - accuracyRows.forEach((row, i) => { - const cells = accuracyTableRows[i] && accuracyTableRows[i].children; - if (!cells) return; - cells[6].textContent = formatPythonResolution(row.sk_fit); - cells[8].textContent = formatPythonResolution(row.sk_pred); - }); - - const timeTableRows = document.querySelectorAll("#time-table-body tr"); - accuracyRows.forEach((row, i) => { - const cells = timeTableRows[i] && timeTableRows[i].children; - if (!cells) return; - cells[2].textContent = formatPythonResolution(row.sk_ms); - }); - - // Digits was previously omitted from the timing charts despite being the - // largest benchmark dataset on the page. - if (document.getElementById("digits-time-chart")) { - logBarChart("digits-time-chart", BENCH.digits, { height: 380 }); - } -}); diff --git a/docs/benchmarks.html b/docs/benchmarks.html index 491e881..f2ed530 100644 --- a/docs/benchmarks.html +++ b/docs/benchmarks.html @@ -1,4 +1,4 @@ -Benchmarks — flow-scikit

canonical v2 / parity + disparity benchmark

Eligibility never means identity.

All 19 canonical rows are measured and currently eligible for comparison, but numerical, semantic and runtime disparities remain first-class evidence. This page renders the committed benchmark and disparity artifacts directly so differences cannot disappear merely because a row passes its contract.

Flow wins—

End-to-end fit + predict comparisons won by Flow.

sklearn wins—

End-to-end comparisons won by scikit-learn.

parity eligible—

Rows admitted to the competitive denominator.

substantive disparities—

Rows whose fitted state, score, configuration or semantics genuinely diverge, above float-noise floors. Runtime differences are tracked per row but not counted here.

TIMING_UNIT|msend-to-endseed=4280/20 persisted split2% practical tie thresholddisparity retained after eligibility
KMeans note: Digits KMeans is eligible under the same declared contract as every other clustering row. Its seeded k-means++ initialization now matches scikit-learn's, so the strict diagnostic and the final eligibility decision agree. The convergence statistic, the point at which inertia is reported, empty-cluster relocation and the n_init selection rule still differ and stay visible in the disparity artifact.

runtime overview

runtime overview

The plots are generated from the canonical JSON.

Each runtime plot shows end-to-end fit + predict time on a log scale. The plots use the same rows as the table below and therefore update whenever the frozen canonical result changes.

All 19 speed ratios

scikit-learn total time divided by Flow total time. The vertical 1× line separates Flow wins from scikit-learn wins.

Iris total runtime

scikit-learnFlow

Digits total runtime

scikit-learnFlow

Diabetes total runtime

scikit-learnFlow

persistent disparity

Passing parity does not erase the gap.

The disparity plot normalizes each row's principal numerical difference against its effective tolerance where a tolerance is available. A value near 1 means the row is close to the acceptance boundary. Semantic/configuration differences are tracked in the same artifact and remain visible in the table.

Numerical disparity relative to tolerance

The dashed line is the acceptance boundary. Values can remain non-zero even for eligible rows.

all canonical rows

No selected-win table.

Every row is shown below. Speedup is sklearn_ms / flow_ms; values above 1× favor Flow. Strict diagnostic status is kept separate from final eligibility.

AlgorithmDatasetFinal parityStrict diagnosticWinnerscore |Δ|sklearn msFlow msspeedup

methodology

Correctness, disparity and timing are separate dimensions.

The benchmark consumes the same persisted train/test indices in Python and Flow. Python uses high-resolution adaptive timing and the canonical runner aggregates repeated process measurements with medians and IQR. Flow timings are emitted in milliseconds and aggregated by the same runner.

Supervised rows compare predictive metrics under declared tolerances. PCA additionally checks explained variance, singular values, reconstruction error and sign-aligned components. KMeans uses permutation-invariant clustering quality and inertia. The persistent disparity artifact preserves raw numerical gaps and known semantic/configuration differences even after the estimator-specific eligibility contract succeeds.

historical deployment evidence

Footprint and startup remain separate experiments.

The repository also contains a historical deployment comparison recording a roughly 1.4 MB Flow native executable and a roughly 65× cold-start advantage (33 ms versus 2160 ms). Those figures come from a different deployment experiment and are intentionally not mixed into the canonical estimator timing denominator.

trajectory

Flow versus Python, across freezes.

Each row's speedup at the previous freeze and at the latest one. A speedup can move because Flow changed or because scikit-learn's side changed on that runner. When a row moves by more than 10%, the last column names which side's own time moved more, from the committed absolute timings.

reproduce

Read the source artifacts.

Canonical result ↗ Disparity report ↗

\ No newline at end of file +fetch('scaled-ci-history.json').then(r=>r.json()).then(h=>{const c=h.counts,tb=document.getElementById('scaled-rows');document.getElementById('scaled-summary').textContent=`Across ${c.runs} consecutive CI runs, Flow wins ${c.rows_won_in_every_run} of the ${c.rows} rows in every one of them. The remaining ${c.rows_lost_in_at_least_one_run} dipped below 1x in at least one run, and every one of those is at 1000 samples.`;h.rows.slice().sort((a,b)=>a.median_speedup-b.median_speedup).forEach(r=>{const tr=document.createElement('tr');const cells=[[r.algorithm.replace(/_/g,' '),''],[r.rows,'num'],[r.features,'num'],[`${r.runs_won} / ${r.runs_observed}`,r.runs_won===r.runs_observed?'num win':'num loss'],[`${r.median_speedup.toFixed(2)}x`,'num'],[`${r.min_speedup.toFixed(2)} to ${r.max_speedup.toFixed(2)}`,'num']];cells.forEach(([t,cl])=>{const td=document.createElement('td');if(cl)td.className=cl;td.textContent=t;tr.appendChild(td)});tb.appendChild(tr)})}); \ No newline at end of file diff --git a/docs/benchmarks.js b/docs/benchmarks.js deleted file mode 100644 index bbf2619..0000000 --- a/docs/benchmarks.js +++ /dev/null @@ -1,607 +0,0 @@ -// SKLEARN_TIMINGS_NORMALIZED_TO_MS -// flow-scikit benchmark data and SVG chart rendering. -// All data measured with seed=42, 80/20 split, xorshift32 PRNG. -// Flow builds use -O3 -march=native, BLAS (Accelerate), SMO for kernel SVM, -// and LBFGS for logistic regression. - -const BENCH = { - iris: [ - { algo: "LogisticRegression", sk_score: 0.9333, fl_score: 0.9333, sk_ms: 4, fl_ms: 0.095, sk_fit: 4, fl_fit: 0.093, sk_pred: 0.000, fl_pred: 0.002 }, - { algo: "LinearSVC", sk_score: 0.9000, fl_score: 0.9000, sk_ms: 1, fl_ms: 0.100, sk_fit: 1, fl_fit: 0.099, sk_pred: 0.000, fl_pred: 0.001 }, - { algo: "KernelSVC_RBF", sk_score: 0.9333, fl_score: 1.0000, sk_ms: 1, fl_ms: 0.549, sk_fit: 1, fl_fit: 0.511, sk_pred: 0.000, fl_pred: 0.038 }, - { algo: "DecisionTree", sk_score: 0.9333, fl_score: 0.9333, sk_ms: 1, fl_ms: 0.082, sk_fit: 1, fl_fit: 0.081, sk_pred: 0.000, fl_pred: 0.001 }, - { algo: "RandomForest", sk_score: 0.9333, fl_score: 0.9333, sk_ms: 5, fl_ms: 0.333, sk_fit: 4, fl_fit: 0.329, sk_pred: 0.000, fl_pred: 0.004 }, - { algo: "GaussianNB", sk_score: 0.9333, fl_score: 0.9333, sk_ms: 0.000, fl_ms: 0.019, sk_fit: 0.000, fl_fit: 0.017, sk_pred: 0.000, fl_pred: 0.002 }, - { algo: "KMeans", sk_score: 0.8333, fl_score: 0.8333, sk_ms: 28, fl_ms: 0.114, sk_fit: 28, fl_fit: 0.113, sk_pred: 0.000, fl_pred: 0.001 }, - { algo: "PCA", sk_score: 0.7262, fl_score: 0.7262, sk_ms: 0.000, fl_ms: 0.020, sk_fit: 0.000, fl_fit: 0.020, sk_pred: 0.000, fl_pred: 0.000, metric: "explained_var" } - ], - digits: [ - { algo: "LogisticRegression", sk_score: 0.9749, fl_score: 0.9666, sk_ms: 7, fl_ms: 21.801, sk_fit: 7, fl_fit: 21.687, sk_pred: 0.000, fl_pred: 0.114 }, - { algo: "LinearSVC", sk_score: 0.9694, fl_score: 0.9694, sk_ms: 237, fl_ms: 100.524, sk_fit: 237, fl_fit: 100.417, sk_pred: 0.000, fl_pred: 0.107 }, - { algo: "KernelSVC_RBF", sk_score: 0.9499, fl_score: 0.9443, sk_ms: 44, fl_ms: 173.044, sk_fit: 25, fl_fit: 160.648, sk_pred: 19, fl_pred: 12.396 }, - { algo: "DecisionTree", sk_score: 0.8886, fl_score: 0.8802, sk_ms: 10, fl_ms: 15.259, sk_fit: 9, fl_fit: 15.232, sk_pred: 0.000, fl_pred: 0.027 }, - { algo: "RandomForest", sk_score: 0.9666, fl_score: 0.9499, sk_ms: 16, fl_ms: 22.483, sk_fit: 15, fl_fit: 22.236, sk_pred: 1, fl_pred: 0.247 }, - { algo: "GaussianNB", sk_score: 0.8134, fl_score: 0.8134, sk_ms: 1, fl_ms: 1.673, sk_fit: 1, fl_fit: 1.077, sk_pred: 0.000, fl_pred: 0.596 }, - { algo: "KMeans", sk_score: 0.6323, fl_score: 0.6267, sk_ms: 34, fl_ms: 103.908, sk_fit: 34, fl_fit: 103.792, sk_pred: 0.000, fl_pred: 0.116 } - ], - diabetes: [ - { algo: "Ridge", sk_score: 0.6089, fl_score: 0.6089, sk_ms: 1, fl_ms: 0.032, sk_fit: 1, fl_fit: 0.032, sk_pred: 0.000, fl_pred: 0.000 }, - { algo: "Lasso", sk_score: 0.6084, fl_score: 0.6085, sk_ms: 1, fl_ms: 4.830, sk_fit: 1, fl_fit: 4.829, sk_pred: 0.000, fl_pred: 0.001 }, - { algo: "LinearRegression", sk_score: 0.6108, fl_score: 0.6108, sk_ms: 1, fl_ms: 0.040, sk_fit: 1, fl_fit: 0.040, sk_pred: 0.000, fl_pred: 0.000 }, - { algo: "KernelRidge_RBF", sk_score: 0.4173, fl_score: 0.4173, sk_ms: 21, fl_ms: 8.562, sk_fit: 20, fl_fit: 7.787, sk_pred: 0.000, fl_pred: 0.775 } - ], - iris_combo: [ - { algo: "GaussianNB", sk_acc: 0.9333, fl_acc: 0.9333, sk_train: 0.00, fl_train: 0.01 }, - { algo: "DecisionTree", sk_acc: 0.9333, fl_acc: 0.9333, sk_train: 0.00, fl_train: 0.10 }, - { algo: "KNN_k5", sk_acc: 0.9333, fl_acc: 0.9667, sk_train: 1.84, fl_train: 0.03 }, - { algo: "LinearSVC_OVR", sk_acc: 0.9000, fl_acc: 0.9000, sk_train: 0.00, fl_train: 0.11 }, - { algo: "RandomForest_10", sk_acc: 0.9333, fl_acc: 0.9333, sk_train: 0.01, fl_train: 0.46 } - ], - // Android: scikit-learn (Flow) cross-compiled to aarch64-linux-android, - // run on the Android emulator (arm64-v8a, Android 15, API 35). Same source, - // same datasets, same seed. Uses scalar BLAS fallback (no Accelerate). - // mac_ms is the macOS arm64 Flow time from BENCH above, for comparison. - android: { - iris: [ - { algo: "LogisticRegression", score: 0.9333, mac_ms: 0.12, and_ms: 0.18 }, - { algo: "LinearSVC", score: 0.9000, mac_ms: 0.11, and_ms: 0.20 }, - { algo: "KernelSVC_RBF", score: 0.9333, mac_ms: 3.31, and_ms: 4.62 }, - { algo: "DecisionTree", score: 0.9333, mac_ms: 0.10, and_ms: 0.41 }, - { algo: "RandomForest", score: 0.9333, mac_ms: 0.46, and_ms: 1.37 }, - { algo: "GaussianNB", score: 0.9333, mac_ms: 0.01, and_ms: 0.01 }, - { algo: "KMeans", score: 0.8333, mac_ms: 0.13, and_ms: 0.16 }, - { algo: "PCA", score: 0.7262, mac_ms: 0.02, and_ms: 0.02 } - ], - digits: [ - { algo: "LogisticRegression", score: 0.9666, mac_ms: 23.91, and_ms: 175.97 }, - { algo: "LinearSVC", score: 0.9694, mac_ms: 125.50, and_ms: 182.62 }, - { algo: "KernelSVC_RBF", score: 0.9471, mac_ms: 4015.19, and_ms: 6401.20 }, - { algo: "DecisionTree", score: 0.8774, mac_ms: 23.10, and_ms: 26.84 }, - { algo: "RandomForest", score: 0.9499, mac_ms: 34.18, and_ms: 46.92 }, - { algo: "GaussianNB", score: 0.8134, mac_ms: 1.78, and_ms: 2.66 }, - { algo: "KMeans", score: 0.6267, mac_ms: 159.63, and_ms: 183.87 } - ], - diabetes: [ - { algo: "Ridge", score: 0.6089, mac_ms: 0.03, and_ms: 0.04 }, - { algo: "Lasso", score: 0.6085, mac_ms: 4.60, and_ms: 4.50 }, - { algo: "LinearRegression", score: 0.6108, mac_ms: 0.03, and_ms: 0.03 }, - { algo: "KernelRidge_RBF", score: 0.4181, mac_ms: 11.59, and_ms: 13.99 } - ] - } -}; - -const COLORS = { - flow: "#ef8068", - sklearn: "#5669e8", - ink: "#101b36", - line: "#ded8cb", - grid: "#e8e2d5", - coral: "#ef8068", - blue: "#5669e8", - green: "#2d8a4e", - red: "#c44a3a", - android: "#1f9d76", - muted: "#687493" -}; - -function el(tag, attrs, text) { - const e = document.createElementNS("http://www.w3.org/2000/svg", tag); - if (attrs) for (const k in attrs) e.setAttribute(k, attrs[k]); - if (text != null) e.textContent = text; - return e; -} - -function clearSvg(id) { - const svg = document.getElementById(id); - while (svg.firstChild) svg.removeChild(svg.firstChild); - return svg; -} - -// Grouped bar chart for accuracy comparison. -function groupedBarChart(svgId, data, opts) { - const svg = clearSvg(svgId); - const W = 900, H = opts.height || 360; - const ml = 60, mr = 30, mt = 20, mb = 70; - const cw = W - ml - mr; - const ch = H - mt - mb; - const n = data.length; - const groupW = cw / n; - const barW = groupW * 0.32; - const gap = groupW * 0.06; - const yMax = opts.yMax || 1.0; - const yMin = opts.yMin || 0.0; - - // Grid lines - for (let i = 0; i <= 5; i++) { - const y = mt + (ch * i / 5); - svg.appendChild(el("line", { x1: ml, y1: y, x2: W - mr, y2: y, stroke: COLORS.grid, "stroke-width": 1 })); - const val = yMax - (yMax - yMin) * i / 5; - svg.appendChild(el("text", { x: ml - 8, y: y + 4, "text-anchor": "end", "font-size": 11, "font-family": "DM Mono, monospace", fill: COLORS.muted }, val.toFixed(2))); - } - - // Bars - data.forEach((d, i) => { - const gx = ml + i * groupW + groupW / 2; - const skH = (d.sk_score - yMin) / (yMax - yMin) * ch; - const flH = (d.fl_score - yMin) / (yMax - yMin) * ch; - svg.appendChild(el("rect", { x: gx - barW - gap/2, y: mt + ch - skH, width: barW, height: skH, fill: COLORS.sklearn, rx: 3 })); - svg.appendChild(el("rect", { x: gx + gap/2, y: mt + ch - flH, width: barW, height: flH, fill: COLORS.flow, rx: 3 })); - - // Value labels - svg.appendChild(el("text", { x: gx - barW/2 - gap/2, y: mt + ch - skH - 5, "text-anchor": "middle", "font-size": 10, "font-family": "DM Mono, monospace", fill: COLORS.sklearn }, d.sk_score.toFixed(3))); - svg.appendChild(el("text", { x: gx + barW/2 + gap/2, y: mt + ch - flH - 5, "text-anchor": "middle", "font-size": 10, "font-family": "DM Mono, monospace", fill: COLORS.flow }, d.fl_score.toFixed(3))); - - // X label - const label = d.algo.replace(/_/g, " ").length > 14 ? d.algo.substring(0, 12) + ".." : d.algo.replace(/_/g, " "); - svg.appendChild(el("text", { x: gx, y: mt + ch + 18, "text-anchor": "middle", "font-size": 11, "font-family": "DM Sans, sans-serif", fill: COLORS.ink }, label)); - }); - - // Axis - svg.appendChild(el("line", { x1: ml, y1: mt + ch, x2: W - mr, y2: mt + ch, stroke: COLORS.ink, "stroke-width": 1.5 })); -} - -// Log-scale grouped bar chart for timing. -function logBarChart(svgId, data, opts) { - const svg = clearSvg(svgId); - const W = 900, H = opts.height || 360; - const ml = 70, mr = 30, mt = 20, mb = 70; - const cw = W - ml - mr; - const ch = H - mt - mb; - const n = data.length; - const groupW = cw / n; - const barW = groupW * 0.32; - const gap = groupW * 0.06; - - const logMin = -2; // 0.01 ms - const logMax = 5; // 100000 ms - const logRange = logMax - logMin; - - function msToY(ms) { - const l = Math.log10(Math.max(ms, 0.01)); - return mt + ch - ((l - logMin) / logRange) * ch; - } - - // Grid lines at decades - for (let i = 0; i <= logRange; i++) { - const y = mt + ch - (i / logRange) * ch; - svg.appendChild(el("line", { x1: ml, y1: y, x2: W - mr, y2: y, stroke: COLORS.grid, "stroke-width": 1 })); - const val = Math.pow(10, logMin + i); - let label; - if (val < 1) label = val.toFixed(2); - else if (val < 1000) label = val.toFixed(0); - else label = (val / 1000).toFixed(0) + "k"; - svg.appendChild(el("text", { x: ml - 8, y: y + 4, "text-anchor": "end", "font-size": 11, "font-family": "DM Mono, monospace", fill: COLORS.muted }, label + "ms")); - } - - data.forEach((d, i) => { - const gx = ml + i * groupW + groupW / 2; - const skY = msToY(d.sk_ms); - const flY = msToY(d.fl_ms); - const baseY = mt + ch; - svg.appendChild(el("rect", { x: gx - barW - gap/2, y: skY, width: barW, height: baseY - skY, fill: COLORS.sklearn, rx: 3 })); - svg.appendChild(el("rect", { x: gx + gap/2, y: flY, width: barW, height: baseY - flY, fill: COLORS.flow, rx: 3 })); - - // Value labels - const skLabel = d.sk_ms < 1 ? d.sk_ms.toFixed(2) : d.sk_ms < 100 ? d.sk_ms.toFixed(1) : d.sk_ms.toFixed(0); - const flLabel = d.fl_ms < 1 ? d.fl_ms.toFixed(2) : d.fl_ms < 100 ? d.fl_ms.toFixed(1) : d.fl_ms.toFixed(0); - svg.appendChild(el("text", { x: gx - barW/2 - gap/2, y: skY - 5, "text-anchor": "middle", "font-size": 10, "font-family": "DM Mono, monospace", fill: COLORS.sklearn }, skLabel)); - svg.appendChild(el("text", { x: gx + barW/2 + gap/2, y: flY - 5, "text-anchor": "middle", "font-size": 10, "font-family": "DM Mono, monospace", fill: COLORS.flow }, flLabel)); - - // X label - const label = d.algo.replace(/_/g, " ").length > 14 ? d.algo.substring(0, 12) + ".." : d.algo.replace(/_/g, " "); - svg.appendChild(el("text", { x: gx, y: mt + ch + 18, "text-anchor": "middle", "font-size": 11, "font-family": "DM Sans, sans-serif", fill: COLORS.ink }, label)); - }); - - svg.appendChild(el("line", { x1: ml, y1: mt + ch, x2: W - mr, y2: mt + ch, stroke: COLORS.ink, "stroke-width": 1.5 })); -} - -// Linear-scale bar chart for timing (diabetes). -function linearTimeChart(svgId, data) { - const svg = clearSvg(svgId); - const W = 900, H = 280; - const ml = 70, mr = 30, mt = 20, mb = 70; - const cw = W - ml - mr; - const ch = H - mt - mb; - const n = data.length; - const groupW = cw / n; - const barW = groupW * 0.32; - const gap = groupW * 0.06; - const yMax = 50; - - for (let i = 0; i <= 5; i++) { - const y = mt + (ch * i / 5); - svg.appendChild(el("line", { x1: ml, y1: y, x2: W - mr, y2: y, stroke: COLORS.grid, "stroke-width": 1 })); - const val = yMax - yMax * i / 5; - svg.appendChild(el("text", { x: ml - 8, y: y + 4, "text-anchor": "end", "font-size": 11, "font-family": "DM Mono, monospace", fill: COLORS.muted }, val.toFixed(0) + "ms")); - } - - data.forEach((d, i) => { - const gx = ml + i * groupW + groupW / 2; - const skH = Math.min(d.sk_ms, yMax) / yMax * ch; - const flH = Math.min(d.fl_ms, yMax) / yMax * ch; - svg.appendChild(el("rect", { x: gx - barW - gap/2, y: mt + ch - skH, width: barW, height: skH, fill: COLORS.sklearn, rx: 3 })); - svg.appendChild(el("rect", { x: gx + gap/2, y: mt + ch - flH, width: barW, height: flH, fill: COLORS.flow, rx: 3 })); - - const skLabel = d.sk_ms < 1 ? d.sk_ms.toFixed(2) : d.sk_ms.toFixed(1); - const flLabel = d.fl_ms < 1 ? d.fl_ms.toFixed(2) : d.fl_ms.toFixed(1); - svg.appendChild(el("text", { x: gx - barW/2 - gap/2, y: mt + ch - skH - 5, "text-anchor": "middle", "font-size": 10, "font-family": "DM Mono, monospace", fill: COLORS.sklearn }, skLabel)); - svg.appendChild(el("text", { x: gx + barW/2 + gap/2, y: mt + ch - flH - 5, "text-anchor": "middle", "font-size": 10, "font-family": "DM Mono, monospace", fill: COLORS.flow }, flLabel)); - - const label = d.algo.replace(/_/g, " "); - svg.appendChild(el("text", { x: gx, y: mt + ch + 18, "text-anchor": "middle", "font-size": 11, "font-family": "DM Sans, sans-serif", fill: COLORS.ink }, label)); - }); - - svg.appendChild(el("line", { x1: ml, y1: mt + ch, x2: W - mr, y2: mt + ch, stroke: COLORS.ink, "stroke-width": 1.5 })); -} - -// Startup time chart (2 bars). -function startupChart() { - const svg = clearSvg("startup-chart"); - const W = 700, H = 200; - const ml = 80, mr = 40, mt = 30, mb = 50; - const cw = W - ml - mr; - const ch = H - mt - mb; - const data = [ - { label: "scikit-learn (Flow)", ms: 33, color: COLORS.flow }, - { label: "scikit-learn (Python)", ms: 2160, color: COLORS.sklearn } - ]; - const yMax = 2400; - const barW = 120; - const gap = 80; - const startX = ml + (cw - (barW * 2 + gap)) / 2; - - for (let i = 0; i <= 4; i++) { - const y = mt + (ch * i / 4); - svg.appendChild(el("line", { x1: ml, y1: y, x2: W - mr, y2: y, stroke: COLORS.grid, "stroke-width": 1 })); - const val = yMax - yMax * i / 4; - svg.appendChild(el("text", { x: ml - 8, y: y + 4, "text-anchor": "end", "font-size": 11, "font-family": "DM Mono, monospace", fill: COLORS.muted }, val.toFixed(0) + "ms")); - } - - data.forEach((d, i) => { - const x = startX + i * (barW + gap); - const h = d.ms / yMax * ch; - svg.appendChild(el("rect", { x: x, y: mt + ch - h, width: barW, height: h, fill: d.color, rx: 5 })); - svg.appendChild(el("text", { x: x + barW/2, y: mt + ch - h - 8, "text-anchor": "middle", "font-size": 14, "font-weight": 600, "font-family": "Fraunces, serif", fill: d.color }, d.ms + " ms")); - svg.appendChild(el("text", { x: x + barW/2, y: mt + ch + 20, "text-anchor": "middle", "font-size": 13, "font-family": "DM Sans, sans-serif", fill: COLORS.ink }, d.label)); - }); - - svg.appendChild(el("line", { x1: ml, y1: mt + ch, x2: W - mr, y2: mt + ch, stroke: COLORS.ink, "stroke-width": 1.5 })); -} - -// Combo chart: accuracy bars + time dots for iris side-by-side. -function irisComboChart() { - const svg = clearSvg("iris-combo-chart"); - const W = 900, H = 380; - const ml = 60, mr = 70, mt = 30, mb = 70; - const cw = W - ml - mr; - const ch = H - mt - mb; - const data = BENCH.iris_combo; - const n = data.length; - const groupW = cw / n; - const barW = groupW * 0.28; - const gap = groupW * 0.04; - - // Left axis: accuracy (0 to 1) - for (let i = 0; i <= 5; i++) { - const y = mt + (ch * i / 5); - svg.appendChild(el("line", { x1: ml, y1: y, x2: W - mr, y2: y, stroke: COLORS.grid, "stroke-width": 1 })); - const val = 1.0 - 0.2 * i; - svg.appendChild(el("text", { x: ml - 8, y: y + 4, "text-anchor": "end", "font-size": 11, "font-family": "DM Mono, monospace", fill: COLORS.sklearn }, val.toFixed(1))); - } - // Right axis: time (log scale, 0.01 to 100) - const logMin = -2, logMax = 2; - for (let i = 0; i <= 4; i++) { - const val = Math.pow(10, logMax - (logMax - logMin) * i / 4); - const y = mt + (ch * i / 4); - svg.appendChild(el("text", { x: W - mr + 8, y: y + 4, "text-anchor": "start", "font-size": 11, "font-family": "DM Mono, monospace", fill: COLORS.flow }, val < 1 ? val.toFixed(2) + "ms" : val.toFixed(0) + "ms")); - } - - function accY(v) { return mt + ch - v * ch; } - function timeY(ms) { - const l = Math.log10(Math.max(ms, 0.01)); - return mt + ch - ((l - logMin) / (logMax - logMin)) * ch; - } - - data.forEach((d, i) => { - const gx = ml + i * groupW + groupW / 2; - - // Accuracy bars - const skH = d.sk_acc * ch; - const flH = d.fl_acc * ch; - svg.appendChild(el("rect", { x: gx - barW - gap/2, y: mt + ch - skH, width: barW, height: skH, fill: COLORS.sklearn, rx: 3, opacity: 0.85 })); - svg.appendChild(el("rect", { x: gx + gap/2, y: mt + ch - flH, width: barW, height: flH, fill: COLORS.flow, rx: 3, opacity: 0.85 })); - - // Time dots - const skTY = timeY(d.sk_train); - const flTY = timeY(d.fl_train); - svg.appendChild(el("circle", { cx: gx - barW/2 - gap/2, cy: skTY, r: 6, fill: "none", stroke: COLORS.sklearn, "stroke-width": 2 })); - svg.appendChild(el("circle", { cx: gx + barW/2 + gap/2, cy: flTY, r: 6, fill: "none", stroke: COLORS.flow, "stroke-width": 2 })); - - // Accuracy labels - svg.appendChild(el("text", { x: gx - barW/2 - gap/2, y: mt + ch - skH - 5, "text-anchor": "middle", "font-size": 10, "font-family": "DM Mono, monospace", fill: COLORS.sklearn }, (d.sk_acc * 100).toFixed(1) + "%")); - svg.appendChild(el("text", { x: gx + barW/2 + gap/2, y: mt + ch - flH - 5, "text-anchor": "middle", "font-size": 10, "font-family": "DM Mono, monospace", fill: COLORS.flow }, (d.fl_acc * 100).toFixed(1) + "%")); - - // X label - const label = d.algo.replace(/_/g, " "); - svg.appendChild(el("text", { x: gx, y: mt + ch + 18, "text-anchor": "middle", "font-size": 11, "font-family": "DM Sans, sans-serif", fill: COLORS.ink }, label)); - }); - - // Axis lines - svg.appendChild(el("line", { x1: ml, y1: mt + ch, x2: W - mr, y2: mt + ch, stroke: COLORS.ink, "stroke-width": 1.5 })); - // Axis labels - svg.appendChild(el("text", { x: ml - 50, y: mt - 8, "font-size": 11, "font-family": "DM Mono, monospace", fill: COLORS.sklearn }, "accuracy")); - svg.appendChild(el("text", { x: W - mr + 8, y: mt - 8, "font-size": 11, "font-family": "DM Mono, monospace", fill: COLORS.flow }, "train ms")); -} - -// Build accuracy table. -function buildAccuracyTable() { - const tbody = document.getElementById("accuracy-table-body"); - const all = [ - ...BENCH.iris.map(d => ({ ...d, dataset: "iris", metric: d.metric === "explained_var" ? "explained_var" : "accuracy" })), - ...BENCH.digits.map(d => ({ ...d, dataset: "digits", metric: "accuracy" })), - ...BENCH.diabetes.map(d => ({ ...d, dataset: "diabetes", metric: "r2" })) - ]; - all.forEach(d => { - const diff = d.fl_score - d.sk_score; - const cls = diff > 0.001 ? "pos" : diff < -0.001 ? "neg" : "neutral"; - const sign = diff > 0 ? "+" : ""; - const skFit = d.sk_fit != null ? d.sk_fit.toFixed(2) : "N/A"; - const flFit = d.fl_fit != null ? d.fl_fit.toFixed(2) : "N/A"; - const skPred = d.sk_pred != null ? d.sk_pred.toFixed(3) : "N/A"; - const flPred = d.fl_pred != null ? d.fl_pred.toFixed(3) : "N/A"; - const tr = document.createElement("tr"); - const tds = [ - { text: d.algo.replace(/_/g, " ") }, - { text: d.dataset }, - { text: d.metric }, - { className: "num", text: d.sk_score.toFixed(4) }, - { className: "num", text: d.fl_score.toFixed(4) }, - { className: `num ${cls}`, text: `${sign}${diff.toFixed(4)}` }, - { className: "num", text: skFit }, - { className: "num", text: flFit }, - { className: "num", text: skPred }, - { className: "num", text: flPred } - ]; - for (const cell of tds) { - const td = document.createElement("td"); - if (cell.className) td.className = cell.className; - td.textContent = cell.text; - tr.appendChild(td); - } - tbody.appendChild(tr); - }); -} - -// Build time table. -function buildTimeTable() { - const tbody = document.getElementById("time-table-body"); - const all = [ - ...BENCH.iris.map(d => ({ ...d, dataset: "iris" })), - ...BENCH.digits.map(d => ({ ...d, dataset: "digits" })), - ...BENCH.diabetes.map(d => ({ ...d, dataset: "diabetes" })) - ]; - all.forEach(d => { - const skTotal = d.sk_ms; - const flTotal = d.fl_ms; - const speedup = skTotal / flTotal; - const speedupStr = speedup > 100 ? speedup.toFixed(0) + "x" : speedup > 1 ? speedup.toFixed(1) + "x" : speedup.toFixed(2) + "x"; - const cls = speedup > 1 ? "pos" : "neg"; - const tr = document.createElement("tr"); - const tds = [ - { text: d.algo.replace(/_/g, " ") }, - { text: d.dataset }, - { className: "num", text: skTotal.toFixed(2) }, - { className: "num", text: flTotal.toFixed(2) }, - { className: `num ${cls}`, text: speedupStr } - ]; - for (const cell of tds) { - const td = document.createElement("td"); - if (cell.className) td.className = cell.className; - td.textContent = cell.text; - tr.appendChild(td); - } - tbody.appendChild(tr); - }); -} - -// Render everything on load. -document.addEventListener("DOMContentLoaded", () => { - startupChart(); - groupedBarChart("iris-accuracy-chart", BENCH.iris, { yMax: 1.0, yMin: 0.0, height: 360 }); - groupedBarChart("digits-accuracy-chart", BENCH.digits, { yMax: 1.0, yMin: 0.0, height: 360 }); - groupedBarChart("diabetes-r2-chart", BENCH.diabetes, { yMax: 0.5, yMin: 0.0, height: 280 }); - logBarChart("iris-time-chart", BENCH.iris, { height: 360 }); - linearTimeChart("diabetes-time-chart", BENCH.diabetes); - irisComboChart(); - buildAccuracyTable(); - buildTimeTable(); - paritySpeedupChart(); - buildParityTable(); - androidTimeChart("android-iris-chart", BENCH.android.iris, { height: 320 }); - androidTimeChart("android-digits-chart", BENCH.android.digits, { height: 360 }); - androidTimeChart("android-diabetes-chart", BENCH.android.diabetes, { height: 280 }); - buildAndroidTable(); -}); - -// Per-algorithm parity data. -const PARITY = [ - { algo: "StandardScaler", match: true, maxDiff: 0.000001, py_ms: 1.59, fl_ms: 0.005, speedup: 317.8 }, - { algo: "MinMaxScaler", match: true, maxDiff: 0.000001, py_ms: 0.36, fl_ms: 0.003, speedup: 118.6 }, - { algo: "MaxAbsScaler", match: true, maxDiff: 0.000000, py_ms: 0.32, fl_ms: 0.014, speedup: 23.2 }, - { algo: "GaussianNB", match: true, maxDiff: 0.000000, py_ms: 3.81, fl_ms: 0.019, speedup: 200.3 }, - { algo: "KNNClassifier_k5", match: true, maxDiff: 0.000000, py_ms: 8.32, fl_ms: 0.061, speedup: 136.4 }, - { algo: "NearestCentroid", match: true, maxDiff: 0.000000, py_ms: 72.69, fl_ms: 0.003, speedup: 24228.7 }, - { algo: "DummyClassifier_mf", match: true, maxDiff: 0.000000, py_ms: 0.29, fl_ms: 0.014, speedup: 20.8 }, - { algo: "DummyRegressor_mean", match: true, maxDiff: 0.000008, py_ms: 0.18, fl_ms: 0.001, speedup: 178.4 }, - { algo: "PCA_2comp", match: true, maxDiff: 0.000001, py_ms: 5.23, fl_ms: 0.027, speedup: 193.7 }, - { algo: "LDA", match: true, maxDiff: 0.000000, py_ms: 4.13, fl_ms: 0.016, speedup: 258.0 }, - { algo: "QDA", match: true, maxDiff: 0.000000, py_ms: 9.63, fl_ms: 0.007, speedup: 1375.3 }, - { algo: "KMeans_k3", match: true, maxDiff: 0.000000, py_ms: 15.96, fl_ms: 0.023, speedup: 693.8 }, - { algo: "KNNRegressor_k3", match: true, maxDiff: 0.000011, py_ms: 1.83, fl_ms: 0.240, speedup: 7.6 }, - { algo: "LinearRegression", match: true, maxDiff: 0.000010, py_ms: 21.15, fl_ms: 0.061, speedup: 346.7 }, - { algo: "Ridge_a1", match: true, maxDiff: 0.000015, py_ms: 12.53, fl_ms: 0.035, speedup: 357.9 }, - { algo: "Lasso_a0.1", match: true, maxDiff: 0.003916, py_ms: 5.55, fl_ms: 0.460, speedup: 12.1 } -]; - -// Parity speedup bar chart. -function paritySpeedupChart() { - const svg = clearSvg("parity-speedup-chart"); - const W = 900, H = 480; - const ml = 80, mr = 30, mt = 20, mb = 90; - const cw = W - ml - mr; - const ch = H - mt - mb; - const data = PARITY; - const n = data.length; - const barW = cw / n * 0.7; - const gap = cw / n * 0.3; - - // Log scale for speedup - const logMin = 0; - const logMax = 5; // 100000x - - function speedupToY(s) { - if (s <= 1) return mt + ch; - const l = Math.log10(Math.min(s, 100000)); - return mt + ch - (l / logMax) * ch; - } - - // Grid lines at decades - for (let i = 0; i <= 5; i++) { - const y = mt + ch - (i / 5) * ch; - svg.appendChild(el("line", { x1: ml, y1: y, x2: W - mr, y2: y, stroke: COLORS.grid, "stroke-width": 1 })); - const val = Math.pow(10, i); - svg.appendChild(el("text", { x: ml - 8, y: y + 4, "text-anchor": "end", "font-size": 11, "font-family": "DM Mono, monospace", fill: COLORS.muted }, val >= 1000 ? (val/1000).toFixed(0) + "kx" : val.toFixed(0) + "x")); - } - - data.forEach((d, i) => { - const x = ml + i * (barW + gap) + gap / 2; - const y = speedupToY(d.speedup); - const color = d.match ? COLORS.green : COLORS.coral; - svg.appendChild(el("rect", { x: x, y: y, width: barW, height: mt + ch - y, fill: color, rx: 3 })); - - // Speedup label - const label = d.speedup > 1000 ? (d.speedup/1000).toFixed(1) + "kx" : d.speedup > 10 ? d.speedup.toFixed(0) + "x" : d.speedup.toFixed(1) + "x"; - svg.appendChild(el("text", { x: x + barW/2, y: y - 5, "text-anchor": "middle", "font-size": 10, "font-family": "DM Mono, monospace", fill: color }, label)); - - // X label (rotated) - const shortName = d.algo.length > 16 ? d.algo.substring(0, 14) + ".." : d.algo; - const text = el("text", { x: x + barW/2, y: mt + ch + 15, "text-anchor": "end", "font-size": 10, "font-family": "DM Sans, sans-serif", fill: COLORS.ink, transform: `rotate(-35, ${x + barW/2}, ${mt + ch + 15})` }, shortName); - svg.appendChild(text); - }); - - svg.appendChild(el("line", { x1: ml, y1: mt + ch, x2: W - mr, y2: mt + ch, stroke: COLORS.ink, "stroke-width": 1.5 })); -} - -// Build parity table. -function buildParityTable() { - const tbody = document.getElementById("parity-table-body"); - PARITY.forEach(d => { - const matchClass = d.match ? "pos" : "neg"; - const matchText = d.match ? "EXACT" : "DIFF"; - const speedupStr = d.speedup > 1000 ? (d.speedup/1000).toFixed(1) + "kx" : d.speedup > 10 ? d.speedup.toFixed(0) + "x" : d.speedup.toFixed(1) + "x"; - const speedupClass = d.speedup > 1 ? "pos" : "neg"; - const diffStr = d.maxDiff < 0.001 ? d.maxDiff.toFixed(6) : d.maxDiff.toFixed(2); - const tr = document.createElement("tr"); - const tds = [ - { text: d.algo }, - { text: d.match ? "EXACT" : "close" }, - { className: matchClass, text: matchText }, - { className: "num", text: diffStr }, - { className: "num", text: d.py_ms.toFixed(4) }, - { className: "num", text: d.fl_ms.toFixed(4) }, - { className: `num ${speedupClass}`, text: speedupStr } - ]; - for (const cell of tds) { - const td = document.createElement("td"); - if (cell.className) td.className = cell.className; - td.textContent = cell.text; - tr.appendChild(td); - } - tbody.appendChild(tr); - }); -} - -// Android: macOS Flow vs Android Flow time, log-scale grouped bars. -function androidTimeChart(svgId, data, opts) { - const svg = clearSvg(svgId); - const W = 900, H = opts.height || 320; - const ml = 70, mr = 30, mt = 20, mb = 70; - const cw = W - ml - mr; - const ch = H - mt - mb; - const n = data.length; - const groupW = cw / n; - const barW = groupW * 0.32; - const gap = groupW * 0.06; - - const logMin = -2, logMax = 4; - const logRange = logMax - logMin; - function msToY(ms) { - const l = Math.log10(Math.max(ms, 0.01)); - return mt + ch - ((l - logMin) / logRange) * ch; - } - - for (let i = 0; i <= logRange; i++) { - const y = mt + ch - (i / logRange) * ch; - svg.appendChild(el("line", { x1: ml, y1: y, x2: W - mr, y2: y, stroke: COLORS.grid, "stroke-width": 1 })); - const val = Math.pow(10, logMin + i); - let label; - if (val < 1) label = val.toFixed(2); - else if (val < 1000) label = val.toFixed(0); - else label = (val / 1000).toFixed(0) + "k"; - svg.appendChild(el("text", { x: ml - 8, y: y + 4, "text-anchor": "end", "font-size": 11, "font-family": "DM Mono, monospace", fill: COLORS.muted }, label + "ms")); - } - - data.forEach((d, i) => { - const gx = ml + i * groupW + groupW / 2; - const baseY = mt + ch; - const flY = msToY(d.and_ms); - svg.appendChild(el("rect", { x: gx + gap/2, y: flY, width: barW, height: baseY - flY, fill: COLORS.android, rx: 3 })); - const flLabel = d.and_ms < 1 ? d.and_ms.toFixed(2) : d.and_ms < 100 ? d.and_ms.toFixed(1) : d.and_ms.toFixed(0); - svg.appendChild(el("text", { x: gx + barW/2 + gap/2, y: flY - 5, "text-anchor": "middle", "font-size": 10, "font-family": "DM Mono, monospace", fill: COLORS.android }, flLabel)); - - if (d.mac_ms != null) { - const macY = msToY(d.mac_ms); - svg.appendChild(el("rect", { x: gx - barW - gap/2, y: macY, width: barW, height: baseY - macY, fill: COLORS.flow, rx: 3 })); - const macLabel = d.mac_ms < 1 ? d.mac_ms.toFixed(2) : d.mac_ms < 100 ? d.mac_ms.toFixed(1) : d.mac_ms.toFixed(0); - svg.appendChild(el("text", { x: gx - barW/2 - gap/2, y: macY - 5, "text-anchor": "middle", "font-size": 10, "font-family": "DM Mono, monospace", fill: COLORS.flow }, macLabel)); - } - - const label = d.algo.replace(/_/g, " ").length > 14 ? d.algo.substring(0, 12) + ".." : d.algo.replace(/_/g, " "); - svg.appendChild(el("text", { x: gx, y: mt + ch + 18, "text-anchor": "middle", "font-size": 11, "font-family": "DM Sans, sans-serif", fill: COLORS.ink }, label)); - }); - - svg.appendChild(el("line", { x1: ml, y1: mt + ch, x2: W - mr, y2: mt + ch, stroke: COLORS.ink, "stroke-width": 1.5 })); -} - -// Android results table. -function buildAndroidTable() { - const tbody = document.getElementById("android-table-body"); - const all = [ - ...BENCH.android.iris.map(d => ({ ...d, dataset: "iris" })), - ...BENCH.android.digits.map(d => ({ ...d, dataset: "digits" })), - ...BENCH.android.diabetes.map(d => ({ ...d, dataset: "diabetes" })) - ]; - all.forEach(d => { - const ratio = d.mac_ms != null ? d.mac_ms / d.and_ms : null; - let ratioStr, ratioCls; - if (ratio == null) { ratioStr = "-"; ratioCls = "neutral"; } - else if (ratio > 1) { ratioStr = ratio.toFixed(2) + "x"; ratioCls = "neg"; } - else { ratioStr = (1 / ratio).toFixed(2) + "x"; ratioCls = "pos"; } - const tr = document.createElement("tr"); - const tds = [ - { text: d.algo.replace(/_/g, " ") }, - { text: d.dataset }, - { className: "num", text: d.score.toFixed(4) }, - { className: "num", text: d.mac_ms != null ? d.mac_ms.toFixed(2) : "-" }, - { className: "num", text: d.and_ms.toFixed(2) }, - { className: `num ${ratioCls}`, text: ratioStr } - ]; - for (const cell of tds) { - const td = document.createElement("td"); - if (cell.className) td.className = cell.className; - td.textContent = cell.text; - tr.appendChild(td); - } - tbody.appendChild(tr); - }); -} diff --git a/docs/demos.html b/docs/demos.html index 8e3388a..f90e0a4 100644 --- a/docs/demos.html +++ b/docs/demos.html @@ -24,7 +24,7 @@

interactive statistical machine learning

Try the whole ML process in your browser.

-

Not screenshots. Not code samples. Change the data assumptions, fit models, move thresholds, rerun validation, inspect uncertainty and watch the results change immediately.

+

Everything here runs live. Change the data assumptions, fit models, move thresholds, rerun validation, inspect uncertainty and watch the results change immediately.

Run the full workflow Browse 29 demos @@ -62,7 +62,7 @@

Try the whole ML process in your browser.

-

02–28 / ML playground

Change one assumption. Rerun the process.

+

02 to 28 / ML playground

Change one assumption. Rerun the process.

Every card includes data generation, fitting or statistical estimation, a live visualisation and quantitative diagnostics. Use the filters to jump directly to a part of the workflow.

@@ -87,13 +87,13 @@

Try the whole ML process in your browser.

-
Predicted digit—Draw a digit, then classify it.
+
Predicted digit...Draw a digit, then classify it.

keep going

-

The browser is the front door, not the limit.

+

The browser is the front door.

The repository contains the full Flow implementations, benchmark suite, tests and examples behind the project. The browser playground is being progressively switched from reference JavaScript to the direct MLIR/WASM target as the runtime surface becomes available.