All 19 rows are measured. All 19 are parity-eligible.
This page renders the committed headline_result_v2.json artifact directly. Competitive timings use explicit millisecond units, persisted identical train/test fixtures, repeated measurements and a separate numerical-parity contract.
Flow wins—
End-to-end fit + predict comparisons won by Flow.
sklearn wins—
End-to-end comparisons won by scikit-learn.
parity eligible—
Rows admitted to the competitive denominator.
unresolved—
Parity or measurement failures. Canonical v2 currently requires zero.
TIMING_UNIT|msend-to-endseed=4280/20 persisted split2% practical tie threshold
KMeans note: Digits KMeans is classified as approximately equivalent rather than bit-identical. The semantic audit identifies initialization as the first divergence; the final objective remains closely matched, so the row is admitted under the explicit clustering-equivalence contract rather than by loosening measurement rules.
all canonical rows
No selected-win table.
Every row is shown below. Speedup is sklearn_ms / flow_ms; values above 1× favor Flow.
Algorithm
Dataset
Parity
Winner
sklearn ms
Flow ms
speedup
methodology
Correctness and timing are separate gates.
The benchmark consumes the same persisted train/test indices in Python and Flow. Python uses high-resolution adaptive timing and the canonical runner aggregates repeated process measurements with medians and IQR. Flow timings are emitted in milliseconds and aggregated by the same runner. A row enters the headline only after its estimator-specific parity contract succeeds.
Supervised rows compare predictive metrics under declared tolerances. PCA additionally checks explained variance, singular values, reconstruction error and sign-aligned components. KMeans uses permutation-invariant clustering quality and inertia rather than label-mapped classification accuracy.
historical deployment evidence
Footprint and startup remain separate experiments.
The repository also contains a historical deployment comparison recording a roughly 1.4 MB Flow native executable and a roughly 65× cold-start advantage (33 ms versus 2160 ms). Those figures come from a different deployment experiment and are intentionally not mixed into the canonical estimator timing denominator.
\ No newline at end of file
+Benchmarks — flow-scikitflow~scikit
canonical v2 / parity + disparity benchmark
Eligibility never means identity.
All 19 canonical rows are measured and currently eligible for comparison, but numerical, semantic and runtime disparities remain first-class evidence. This page renders the committed benchmark and disparity artifacts directly so differences cannot disappear merely because a row passes its contract.
Flow wins—
End-to-end fit + predict comparisons won by Flow.
sklearn wins—
End-to-end comparisons won by scikit-learn.
parity eligible—
Rows admitted to the competitive denominator.
tracked disparities—
Rows with non-zero numerical or explicit semantic/configuration differences.
TIMING_UNIT|msend-to-endseed=4280/20 persisted split2% practical tie thresholddisparity retained after eligibility
KMeans note: Digits KMeans is eligible as approximately equivalent, not bit-identical. Its strict diagnostic disparity, different initialization trajectory and convergence semantics remain visible separately from the final eligibility decision.
runtime overview
The plots are generated from the canonical JSON.
Each runtime plot shows end-to-end fit + predict time on a log scale. The plots use the same rows as the table below and therefore update whenever the frozen canonical result changes.
All 19 speed ratios
scikit-learn total time divided by Flow total time. The vertical 1× line separates Flow wins from scikit-learn wins.
Iris total runtime
scikit-learnFlow
Digits total runtime
scikit-learnFlow
Diabetes total runtime
scikit-learnFlow
persistent disparity
Passing parity does not erase the gap.
The disparity plot normalizes each row's principal numerical difference against its effective tolerance where a tolerance is available. A value near 1 means the row is close to the acceptance boundary. Semantic/configuration differences are tracked in the same artifact and remain visible in the table.
Numerical disparity relative to tolerance
The dashed line is the acceptance boundary. Values can remain non-zero even for eligible rows.
all canonical rows
No selected-win table.
Every row is shown below. Speedup is sklearn_ms / flow_ms; values above 1× favor Flow. Strict diagnostic status is kept separate from final eligibility.
Algorithm
Dataset
Final parity
Strict diagnostic
Winner
score |Δ|
sklearn ms
Flow ms
speedup
methodology
Correctness, disparity and timing are separate dimensions.
The benchmark consumes the same persisted train/test indices in Python and Flow. Python uses high-resolution adaptive timing and the canonical runner aggregates repeated process measurements with medians and IQR. Flow timings are emitted in milliseconds and aggregated by the same runner.
Supervised rows compare predictive metrics under declared tolerances. PCA additionally checks explained variance, singular values, reconstruction error and sign-aligned components. KMeans uses permutation-invariant clustering quality and inertia. The persistent disparity artifact preserves raw numerical gaps and known semantic/configuration differences even after the estimator-specific eligibility contract succeeds.
historical deployment evidence
Footprint and startup remain separate experiments.
The repository also contains a historical deployment comparison recording a roughly 1.4 MB Flow native executable and a roughly 65× cold-start advantage (33 ms versus 2160 ms). Those figures come from a different deployment experiment and are intentionally not mixed into the canonical estimator timing denominator.
\ No newline at end of file
diff --git a/docs/disparity-history.html b/docs/disparity-history.html
new file mode 100644
index 0000000..dcfec6e
--- /dev/null
+++ b/docs/disparity-history.html
@@ -0,0 +1 @@
+Disparity history — flow-scikitflow~scikit
persistent benchmark evidence
Disparity is tracked across freezes, not only at the current commit.
Each canonical benchmark freeze stores score disparity, tolerance usage, runtime ratio and semantic/configuration counts. CI compares the new snapshot with the previous frozen report before accepting it.
Worst tolerance-normalized numerical disparity by snapshot
The dashed line is the acceptance boundary. A rising line means at least one canonical row is moving closer to its numerical tolerance even if parity still passes.