Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
63 changes: 63 additions & 0 deletions .github/workflows/flow.yml
Original file line number Diff line number Diff line change
Expand Up @@ -157,6 +157,69 @@ jobs:
benchmarks/parity_diagnostics.json
benchmarks/headline_environment.json

estimator-matrix:
name: Wide estimator matrix
runs-on: ubuntu-latest
# Both sides call into OpenBLAS, so the thread count is pinned here for the
# reason the canonical job pins it.
env:
OPENBLAS_NUM_THREADS: "4"
OMP_NUM_THREADS: "4"
MKL_NUM_THREADS: "4"
steps:
- uses: actions/checkout@v4

- uses: actions/checkout@v4
with:
repository: flooooooooooow/flow
ref: 88aac5095488309813c7173d268c1f8260421c6e
path: .flow-toolchain

- uses: actions/setup-python@v5
with:
python-version: "3.12"

- name: Install benchmark dependencies
run: |
pip install -r .flow-toolchain/requirements.txt
pip install numpy scikit-learn
sudo apt-get update && sudo apt-get install -y libopenblas-dev

- name: Rebuild the registry and the generated benchmarks
run: |
python benchmarks/estimator_coverage.py
python benchmarks/generate_estimator_bench.py
git diff --exit-code benchmarks/estimator_coverage.json benchmarks/generated

- name: Time every estimator on the Flow side
env:
FLOW_BIN: ${{ github.workspace }}/.flow-toolchain/flow
FLOW_HOST: python
FLOW_OPT_LEVEL: "3"
FLOW_LDFLAGS: "-lm -lopenblas lib/scikit/flow_time.c lib/scikit/flow_parallel.c"
run: python benchmarks/run_estimator_bench.py --rounds 3 --out benchmarks/estimator_flow_raw.txt

- name: Time every estimator on the scikit-learn side
run: python benchmarks/bench_estimators_sklearn.py --repeats 5

- name: Join the two sides
run: |
python benchmarks/compare_estimators.py benchmarks/estimator_flow_raw.txt \
--note "GitHub Actions ubuntu-latest, Flow at -O3 with adaptive repeats, fastest of 3 rounds; scikit-learn wheel, fastest of 5; OpenBLAS pinned to 4 threads"

- name: Gate every ranked row against scikit-learn
run: python benchmarks/check_estimator_matrix.py benchmarks/estimator_comparison.json

- uses: actions/upload-artifact@v4
if: always()
with:
name: estimator-matrix-${{ github.run_id }}
path: |
benchmarks/estimator_flow_raw.txt
benchmarks/estimators_sklearn.json
benchmarks/estimator_comparison.json
benchmarks/estimator_coverage.json

scaled-report:
name: Scaled benchmark report
runs-on: ubuntu-latest
Expand Down
21 changes: 21 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,6 +73,27 @@ The ten rows at 10000 samples, the only sizes in the matrix large enough to meas

`benchmarks/scaled_flow_baseline.json` is a separate Flow-against-itself gate that does fail the build. Refresh it from a CI artifact rather than from a developer machine, or the gate will read a fast laptop as the standard and fail every CI run.

## The whole library

The 19 canonical rows race twelve estimators. `lib/scikit` exports 203, and a
claim about Flow against scikit-learn covers six percent of the library while
the rest go unmeasured. A registry maps every exported `fit` to its
scikit-learn counterpart, times both sides on the same data, and a CI job
fails the build when a ranked row is slower.

**Flow is faster on 188 of the 188 ranked rows.** The rows that carry no ratio
carry a reason instead: four implementations say in their own comments that
they are simplified, so they are timed and shown without being ranked; five
Flow functions do part of what their scikit-learn namesake does, such as a
voting estimator that takes models already fitted; six have no scikit-learn
counterpart at all.

These rows carry no parity contract and no declared tolerances. Each library
runs its own defaults over the same data, which answers whether an
implementation is in the same performance league and says nothing about
whether it computes the same thing. The canonical rows above are where
numerical equivalence is established.

Detailed artifacts:

- [`benchmarks/SKLEARN_EXECUTION_INVENTORY.md`](benchmarks/SKLEARN_EXECUTION_INVENTORY.md)
Expand Down
26 changes: 23 additions & 3 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -241,11 +241,19 @@ reason.
| bucket | meaning |
| --- | --- |
| `runnable` | arguments resolved and a scikit-learn counterpart exists |
| `different_shape` | takes a pipeline, a vectorizer input or a list of fitted models first, so it is not an estimator over a feature matrix |
| `shaped` | its fit does not begin with a feature matrix, so the registry carries the call written out, and the row is raced and ranked like any other |
| `different_shape` | the Flow function does part of what the scikit-learn class does, such as a voting estimator that takes models already fitted |
| `flow_only` | Flow implements it and scikit-learn has no equivalent |
| `simplified` | the implementation's own comments call it a simplified stand-in |
| `blocked` | the signature is not resolved yet, with the missing parameter named |

A `shaped` entry gives the Flow call in terms of the variables the generated
harness declares, what the scikit-learn side fits and transforms, and where
needed a preamble that builds the input: a pipeline of a scaler and a logistic
regression, a corpus for the two text vectorizers, a list of dicts for the dict
vectorizer. The corpus lives in the registry so the Flow file and the
scikit-learn harness read one copy of it.

The `simplified` bucket is detected from the source rather than listed, so it
stays true as the implementations are filled in. It matters: `spectral_biclustering`
thresholds row and column means where scikit-learn does an SVD and k-means, and
Expand All @@ -266,11 +274,23 @@ scikit-learn side of the same registry, and
```
python benchmarks/estimator_coverage.py
python benchmarks/generate_estimator_bench.py
for f in benchmarks/generated/bench_estimators_*.flow; do flow run "$f"; done | tee /tmp/flow.txt
python benchmarks/run_estimator_bench.py --rounds 3
python benchmarks/bench_estimators_sklearn.py
python benchmarks/compare_estimators.py /tmp/flow.txt
python benchmarks/compare_estimators.py benchmarks/estimator_flow_raw.txt
python benchmarks/check_estimator_matrix.py
```

[`run_estimator_bench.py`](run_estimator_bench.py) checks every chunk's exit
status, so a chunk that dies takes its own row count down with it in the
summary rather than disappearing, and reuses the binary the first round leaves
behind so later rounds pay for timing instead of for compiling the library
again. [`check_estimator_matrix.py`](check_estimator_matrix.py) fails on a
ranked row slower than scikit-learn, on a row that reported no timing, and on a
ranked count that has quietly shrunk. The `Wide estimator matrix` job in
`.github/workflows/flow.yml` runs all of it on a runner that is not competing
with anything, which is where a published number belongs. A developer machine
under load recorded the same row at 1.72x and at 0.96x in consecutive runs.

Read the result for what it is. These rows have no parity contract, no declared
tolerances and no disparity report: each library runs its own defaults over the
same data. That answers whether a Flow implementation is in the same
Expand Down
62 changes: 55 additions & 7 deletions benchmarks/bench_estimators_sklearn.py
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,8 @@
# this data. A shallow tree keeps the wrapper's own overhead visible rather than
# burying it under the base estimator's work.
def _constructors() -> dict:
from sklearn.linear_model import Ridge
from sklearn.linear_model import LogisticRegression, Ridge
from sklearn.preprocessing import FunctionTransformer, StandardScaler
from sklearn.tree import DecisionTreeClassifier, DecisionTreeRegressor
import numpy as np

Expand All @@ -57,6 +58,16 @@ def _constructors() -> dict:
"SparseCoder": lambda c: c(dictionary=np.eye(4)),
# nu=0.5 is infeasible for this class balance.
"NuSVC": lambda c: c(nu=0.1),
# A selector needs something to read importances from.
"SelectFromModel": lambda c: c(tree_c(), threshold=-np.inf, max_features=2),
# The composed rows, built to match what the Flow harness composes:
# a scaler and a logistic regression, a scaler over the columns, and a
# scaler beside a passthrough.
"Pipeline": lambda c: c([("scaler", StandardScaler()),
("classifier", LogisticRegression(max_iter=50))]),
"ColumnTransformer": lambda c: c([("scaler", StandardScaler(), [0, 1, 2, 3])]),
"FeatureUnion": lambda c: c([("scaler", StandardScaler()),
("passthrough", FunctionTransformer())]),
}


Expand Down Expand Up @@ -91,7 +102,10 @@ def main() -> int:
args = ap.parse_args()

registry = json.loads(REGISTRY.read_text())
runnable = [e for e in registry["entries"] if e["bucket"] == "runnable"]
# Matches generate_estimator_bench.py: shaped rows are raced and ranked,
# simplified rows are timed on both sides and shown without a ratio.
runnable = [e for e in registry["entries"]
if e["bucket"] in ("runnable", "shaped", "simplified")]
classes = dict(all_estimators())
constructors = _constructors()

Expand All @@ -115,19 +129,53 @@ def main() -> int:
rows.append({"flow_estimator": entry["flow_estimator"], "sklearn_estimator": name,
"status": "unavailable", "reason": "not in sklearn.utils.all_estimators()"})
continue
kind = dataset_kind(entry)
shape = entry.get("shape")
kind = shape["dataset"] if shape else dataset_kind(entry)
X, y = data[kind]
ctor_kwargs = {}
if shape and shape.get("sklearn_ctor"):
# The string is the constructor's own keyword list, kept next to the
# Flow call it matches so the two cannot drift apart.
ctor_kwargs = eval(f"dict({shape['sklearn_ctor']})") # noqa: S307
# A row whose scikit-learn side fits a target vector or a single ordered
# variable rather than a design. The Flow side of these is written out
# in the registry for the same reason.
fit_input = (shape or {}).get("sklearn_input", "X")
if fit_input == "y":
first, second = y, None
elif fit_input == "x1d":
first, second = X[:, 0], y
elif fit_input == "docs":
# The corpus travels in the registry, so the Flow file and this one
# read one copy of it.
first, second = list(shape["corpus"]), None
elif fit_input == "dicts":
first = [{f"f{j}": float(v) for j, v in enumerate(row)} for row in X]
second = None
elif fit_input == "labelsets":
first = [tuple(int(v) for v in row) for row in y]
second = None
else:
first, second = X, y
try:
with warnings.catch_warnings():
warnings.simplefilter("ignore")
build = constructors.get(name)
model = build(cls) if build else cls()
fit_ms = timed((lambda: model.fit(X)) if y is None else (lambda: model.fit(X, y)), args.repeats)
model = build(cls) if build else cls(**ctor_kwargs)
fit_ms = timed(
(lambda: model.fit(first)) if second is None
else (lambda: model.fit(first, second)),
args.repeats,
)
pred_ms = 0.0
for method in ("predict", "transform"):
# A recipe can name the method to time, for a class whose work
# is called something other than predict or transform.
methods = [shape["sklearn_work"]] if shape and shape.get("sklearn_work") \
else ["predict", "transform"]
for method in methods:
if hasattr(model, method):
try:
pred_ms = timed(lambda m=method: getattr(model, m)(X), args.repeats)
pred_ms = timed(lambda m=method: getattr(model, m)(first), args.repeats)
except Exception:
pred_ms = 0.0
break
Expand Down
66 changes: 66 additions & 0 deletions benchmarks/check_estimator_matrix.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
#!/usr/bin/env python3
"""Gate the wide estimator matrix: every ranked row has to favour Flow.

The matrix was measured by hand for a long time, which is why a published
number sat at 145 of 166 while the working tree was at 155. A gate in CI is
what keeps the claim and the code in step: a row that turns slower than
scikit-learn fails the run, and a row that stops reporting fails it too,
because a chunk that dies mid-run is indistinguishable from a chunk with
nothing to say.

Rows the registry marks simplified carry no ratio and are not ranked. Rows
under the clock's floor carry no ratio either. Both are reported and neither
fails the run.
"""
from __future__ import annotations

import argparse
import json
from pathlib import Path

ROOT = Path(__file__).resolve().parents[1]


def main() -> int:
ap = argparse.ArgumentParser()
ap.add_argument("comparison", nargs="?", type=Path,
default=ROOT / "benchmarks" / "estimator_comparison.json")
ap.add_argument("--min-compared", type=int, default=160,
help="fail if fewer rows than this produced a ratio")
ap.add_argument("--tolerance", type=float, default=1.0,
help="the ratio a row has to reach, sklearn_ms / flow_ms")
args = ap.parse_args()

payload = json.loads(args.comparison.read_text())
rows = payload["rows"]
ranked = [r for r in rows if r.get("status") == "ok" and r.get("speedup") is not None]
losers = sorted((r for r in ranked if r["speedup"] < args.tolerance),
key=lambda r: r["speedup"])
missing = [r for r in rows if r.get("status") in ("flow_missing", "sklearn_missing")]

print(f"ranked {len(ranked)} rows, {len(ranked) - len(losers)} at or above "
f"{args.tolerance:.2f}x")
for r in losers:
print(f" SLOWER {r['speedup']:6.3f}x {r['flow_estimator']:32s} "
f"flow={r['flow_ms']:9.3f} sklearn={r['sklearn_ms']:9.3f}")
for r in missing:
print(f" MISSING {r['flow_estimator']:32s} {r['status']}: {r.get('reason', '')}")

failed = False
if losers:
print(f"{len(losers)} rows are slower than scikit-learn")
failed = True
if missing:
print(f"{len(missing)} rows reported no timing")
failed = True
if len(ranked) < args.min_compared:
print(f"only {len(ranked)} rows produced a ratio, expected at least {args.min_compared}")
failed = True
if failed:
return 1
print("every ranked row favours Flow")
return 0


if __name__ == "__main__":
raise SystemExit(main())
2 changes: 1 addition & 1 deletion benchmarks/compare_estimators.py
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ def main() -> int:
rows = []
for entry in registry["entries"]:
name = entry["flow_estimator"]
if entry["bucket"] != "runnable":
if entry["bucket"] not in ("runnable", "shaped"):
row = {"flow_estimator": name, "status": entry["bucket"], "reason": entry.get("reason", "")}
if entry.get("sklearn_estimator"):
row["sklearn_estimator"] = entry["sklearn_estimator"]
Expand Down
Loading
Loading