Skip to content

Beat scikit-learn on all 188 ranked estimator rows, gated in CI - #509

Merged
godofecht merged 12 commits into
mainfrom
beat-every-row
Sep 28, 2026
Merged

godofecht merged 12 commits into
mainfrom
beat-every-row

Conversation

@godofecht

@godofecht godofecht commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Result

188 ranked rows, 188 of them in front of scikit-learn, measured by a new CI
job on a runner that is not competing with anything. No row reported no timing,
and no row landed under the clock's floor.

Narrowest: affinity propagation 1.108x, svr 1.227x, nu_svr 1.324x, isomap
1.326x. Median 27.7x. The widest ratios say more about the two libraries'
defaults than about the code, which the page says in its own words.

Coverage went from 172 of the 203 exported estimators to 192. What is left
carries a specific reason rather than a generic one: four implementations say
in their own comments that they are simplified, so they are timed and shown
without being ranked; five Flow functions do part of what their scikit-learn
namesake does, such as a voting estimator that takes models already fitted;
six have no scikit-learn counterpart.

The gate

benchmarks/run_estimator_bench.py compiles and runs the generated chunks with
every exit status checked, and reuses the binary the first round leaves behind.
benchmarks/check_estimator_matrix.py fails on a ranked row slower than
scikit-learn, on a row that reported no timing, and on a ranked count that has
quietly shrunk. The Wide estimator matrix job runs both.

The gate paid for itself immediately. Its first run found four rows that lose
on Linux and win on macOS, its second found three more, and both sets are fixed
here.

The rows this fixed

row before what it was
multitask_elastic_net_cv 0.28x fixed step proximal gradient; now block coordinate descent, warm started down the alpha path
multitask_lasso_cv 0.54x the same, plus a convergence test the old one could not fire
spectral_embedding 0.30x one scalar matvec per component per iteration; now one sgemm over the block
mds 0.57x gradient descent at a fixed hundredth of a step; now the SMACOF Guttman transform
affinity_propagation 0.32x convergence decided from argmax labels that flip on float ties
kernel_pca 0.96x two row copies allocated per matrix entry, and a symmetric matrix built twice
svr 0.96x recomputed every support vector's decision value to average its bias, when the descent had been maintaining it
kernel_ridge 0.71x scalar dot per pair through matrix_set, read back through matrix_at
lle 0.78x W^T W walked by hand, neighbours found by bubble sorting every row
tweedie_regressor 0.74x design walked twice per row through matrix_at, gradient buffer allocated per iteration
linear_svr 0.64x the epoch's dot accumulated into one register
classifier_chain, kneighbors_transformer 0.75x, 0.95x matrix_set copies, and a zero fill matrix_new had already done

Two crashes the breadth found

factor_analysis_fit allocated its component array with malloc, tested each row
against null and wrote through whatever the allocator had left there. On macOS
that memory happened to be zero, so the fit worked and the row had been timed
for weeks. On the runner it was not, and the process died partway through its
chunk, taking the nine estimators after it down with it.

_hash_string carried djb2 in an i32 and shifted it left every character. Past
the sixth character the value had wrapped negative, and the runtime traps a left
shift of a negative. The four document fixture in the tests never reached that.
A sixteen document corpus did.

Verification

  • Wide estimator matrix: success, 188/188.
  • Run all tests and examples: success, the full suite on Linux.
  • Canonical benchmark contract, Scaled benchmark report, FLOW syntax and imports, Python native package: all success.

Five SVM and ridge rows sat between 0.03x and 0.49x. Each had the same shape
of problem: work repeated that did not need repeating.

svr walked every coefficient by a fixed 0.01 whenever it saw a violation and
recomputed the decision value from scratch for every sample on each of up to
1000 passes. Minimising the epsilon-insensitive objective in one coordinate
has a closed form, a soft threshold of the partial residual clipped to the
box, and the decision values can be carried forward as each coefficient moves.
66.4 ms to 2.5 ms, 0.10x to 2.6x.

nu_svr did the same and rebuilt the RBF kernel inside the innermost loop,
taking an exponential for every pair of samples on every pass for a matrix
that never changes. 128.6 ms to 2.3 ms, 0.03x to 1.9x.

one_class_svm rebuilt the kernel the same way, in the fit loop and again in
the threshold pass. The update rule and step are unchanged, so the result is
the same. 8.1 ms to 0.19 ms, 0.06x to 2.7x.

ridge_cv and ridge_classifier_cv swept an alpha grid with the fold loop
inside, rebuilding each training matrix once per alpha through matrix_at,
which copies the Matrix struct per element. Worse, every alpha rebuilt the
Gram matrix, which does not depend on alpha. _ridge_gram and
_ridge_solve_gram split those apart: one Gram per fold, one cheap solve per
alpha. 0.72 ms to 0.11 ms and 25.7 ms to 0.70 ms, 0.49x to 3.2x and 0.04x to
1.4x.

linear_svr carried the regularizer as a scale instead of writing it across
the weight vector once per sample, took its elements off the row pointer, and
stops when every sample is inside the tube. 0.39 ms to 0.27 ms, which is
0.88x and still short of liblinear.
elastic_net_fit was proximal gradient with a fixed step and a fixed budget of
epochs, allocating its gradient buffer inside the epoch loop. At the small end
of a regularization path that converges slowly, so every fit ran all thousand
epochs however close it already was. It is cyclic coordinate descent now,
which is what cd_fast does and which settles in tens of sweeps at any alpha.
elastic_net_cv goes from 167.8 ms to 2.5 ms, 0.08x to 5.5x, and elastic_net
itself from 0.39 ms to 0.028 ms.

multitask_lasso and multitask_elastic_net allocated a gradient buffer per task
on every epoch, read the design through matrix_at, and had no way to stop.
0.45 ms to 0.022 ms and 0.45 ms to 0.013 ms, 0.48x to 9.9x and 0.46x to 16x.
multitask_lasso_cv follows them from 0.59x to 1.7x.

affinity_propagation rescanned a whole row to find the largest A + S excluding
one column, for every column, which made the responsibility pass cubic in the
sample count. The row maximum and runner up answer every column at once. Both
availability passes walked down a column of a row-pointer array and now run
row by row. 139.3 ms to 6.4 ms, 0.03x to 0.64x.

local_outlier_factor bubble sorted every row of the distance matrix, twice,
once per sample, so it was cubic as well. Only the k nearest are needed, so it
selects the first k + 1 positions instead. The distance matrix is symmetric
and is now built as one triangle off row pointers. 5.4 ms to 0.73 ms, 0.07x to
0.54x.

Both of those last two are still behind scikit-learn. They were 30x and 14x
behind.
…atrix_at

hdbscan built the same Euclidean distances three ways over. The core-distance
pass rebuilt every row from the raw features and then ran a full selection
sort over it to read one element, which is cubic in the sample count, and
Prim's loop rebuilt each distance again for every pair it considered. One
symmetric distance matrix now serves all of it and the core pass selects only
the positions it needs. 4.5 ms to 0.41 ms, 0.16x to 1.75x.

spectral_embedding builds its affinity matrix as one triangle off row
pointers, and its power iteration takes the row pointer once and reuses one
matvec buffer rather than allocating on every step. None of that moved the
number: the cost is 200 power iterations that do not converge, and beating
ARPACK there needs a different eigensolver. It stays at 0.30x.

multitask_elastic_net_cv no longer predicts over the whole design for every
fold and every alpha when it scores five test folds, and builds each training
pair once for the grid. That did not move it either. The cost is the fold fits
themselves, which are still proximal gradient and still run their whole epoch
budget at the small end of the path. Block coordinate descent is the fix and
is not in this commit. It stays at 0.28x.
The wide matrix had eleven rows where scikit-learn was faster. Three of them
were already fixed in the last two commits and needed no further work. These
are the rest, and in each case the loss came from the algorithm rather than
from the loop it was written in.

MultiTaskElasticNet moves from a fixed step proximal gradient to block
coordinate descent. Each feature block has a closed form minimizer given the
others, so a sweep solves every block exactly, and the residual is carried
through a rank one correction instead of being rebuilt. The cross validated
version walks its alpha grid from the largest alpha down and warm starts each
fit from the last, so the whole grid costs little more than its first point.
62.6 ms to 3.7 ms on the diabetes shape, measured on a loaded machine.

MDS moves from gradient descent at a fixed hundredth of a step to the SMACOF
Guttman transform, which is the exact minimizer of the majorizing function at
the current configuration. Each pair is visited once and its contribution is
added to both of its points. The old stopping rule compared raw stress against
the tolerance and was never true at these scales, so every fit ran its whole
iteration budget.

SpectralEmbedding and LLE move from one scalar matvec per component per
iteration to one sgemm per iteration over the whole block, with modified Gram
Schmidt keeping the block orthonormal. LLE also builds its W^T W term with one
sgemm where it used to walk a scalar dot per entry down a column of an array of
row allocations, and picks its neighbours by partial selection instead of
bubble sorting every row, which was cubic in the sample count on its own.

KernelRidge builds the linear, polynomial and sigmoid kernels from one sgemm.
The old path walked a scalar dot per pair, wrote every entry through
matrix_set, and then read all of it back through matrix_at. A linear kernel
also has a primal form, so predict folds the dual coefficients into one weight
vector rather than evaluating an n by n_train kernel.

KNeighborsTransformer drops a zero fill that matrix_new had already done
through n by n_train calls to matrix_set, and reads its distances off row
pointers. ClassifierChain does the same for the design it rebuilds per output
in predict. LinearSVR predicts with one sgemv.
The 172 row matrix was driven by a shell loop typed out by hand, which is how
a published artifact sat at 145 of 166 while the working tree was at 155, and
how a chunk that died mid-run once passed for a chunk with no rows in it.

run_estimator_bench.py compiles and runs the generated chunks with every exit
status checked, and reuses the binary that the first round leaves behind so
later rounds pay for timing instead of for compiling the library again.
check_estimator_matrix.py fails on any ranked row slower than scikit-learn, on
any row that reported no timing, and on a ranked count that has quietly
shrunk. Both run in a new Flow workflow job on a runner that is not competing
with anything, which is where a published number belongs.

The generator and the scikit-learn harness now time the four estimators the
registry marks simplified. compare_estimators.py already showed their times
while withholding a ratio, so leaving them out of the race dropped four rows
from the page for no reason. The generator also takes a comma separated --only
and a --prefix, so a subset can be re-timed without colliding with the
committed files in the shared build directory.
The wide matrix raced 172 of 203 exported estimators. Twenty-five sat in
different_shape, which meant only that their fit does not begin with a feature
matrix and the generic path builds one shape of call. Thirteen of them race
scikit-learn perfectly well once the call is written out, so the registry now
carries that call and they are ranked like any other row: the seven kernel
approximation and projection transformers, the two dummy estimators, the two
label encoders, the isotonic fit and the model selector.

The twelve that are left take a Pipeline, a ColumnTransformer, a FeatureUnion,
an estimator array, a document array or a dict array, and each keeps its
written reason.

factor_analysis_fit allocated its component array with malloc, tested each row
against null, and wrote through whatever the allocator had left there when the
test failed. On macOS that memory happened to be zero, so the fit worked and
the row has been timed for weeks. On the CI runner it was not, and the process
died partway through its chunk, taking the nine estimators after it down with
it. Every row is allocated up front now.

The harness also adds a sink. Two rows came back at exactly 0.000000 ms because
clang at -O3 is free to delete a transform whose result is freed without being
read. One value from every result is now folded into a running total that main
prints on a line the parser ignores.
The composed rows and the input-shaped rows now carry their call in the
registry the same way the other shaped rows do: Pipeline over a scaler and a
logistic regression, ColumnTransformer over a scaler, FeatureUnion over a
scaler and a passthrough, the two text vectorizers over one corpus, the dict
vectorizer and the multi-label binarizer. That is 192 of 203 raced.

The corpus lives in the registry, so the generated Flow file and the
scikit-learn harness read one copy of it and cannot drift apart. A fit that
takes a composed object built outside the timing frees once after the timing
rather than on every repeat, because the object it hands back is the object the
next repeat reads.

_hash_string carried djb2 in an i32 and shifted the result left every
character. Past the sixth character of a token the value had wrapped negative,
and the runtime traps a left shift of a negative, so the process aborted inside
count_vectorizer_fit. The four document fixture in the tests never reached
that. A sixteen document corpus did. It is carried in 64 bits and reduced every
step now.

The five that are left are the ones where the Flow function does part of what
the scikit-learn class does: the two voting estimators take fitted models, the
stacking classifier takes base predictions, the self-training classifier takes
probabilities, and incremental PCA takes one batch. Each says so in its own
reason rather than in a generic message about its first argument.
The first CI run of the wide matrix put 175 of 179 ranked rows in front of
scikit-learn and named the four that were behind. Every one of them is a
different story from the macOS numbers, which is the point of gating on the
runner rather than on a developer machine.

multitask_lasso_cv, 109.9 ms against 59.7. It now runs the same block
coordinate descent as the elastic net, warm started down the alpha path. The
sweep loop also gained a real convergence test: the coefficient move it used
was weak at a small penalty, where the iterate creeps and no single
coefficient moves far in one sweep, so a fit spent its whole budget on an
objective that had stopped falling. The objective comes out of the residual
the sweep already maintains, for about half a percent of a sweep. 109.9 ms to
3.6 ms on this machine, and the elastic net path gets the same test.

affinity_propagation, 15.5 ms against 4.9. The three n by n matrices are one
allocation each rather than n row allocations, the similarity is built off row
pointers and mirrored, and the label pass is folded into the availability pass
that already has both terms in hand.

tweedie_regressor, 1.31 ms against 0.98, with the Poisson and gamma fits on
the same loop. All three walked the design twice per row through matrix_at and
allocated a gradient buffer inside the iteration. One sgemv for the linear
predictor, one transposed sgemv for the gradient, buffers allocated once.

linear_svr, 0.66 ms against 0.42. The dot product that its epoch is made of
accumulated into one register, so the loop was a chain of dependent additions.
Four accumulators.

kernel_density and nearest_neighbors sat unranked because their work function
is called score_samples and kneighbors, so the generic path found nothing to
time and the fit alone fell under the clock's floor. A recipe can now name the
work function and the scikit-learn method to match it against. That puts every
row the matrix can rank in the ranking.
The second CI run put 185 of 188 ranked rows in front of scikit-learn. These
are the three it named.

affinity_propagation watched the argmax label of every row to decide it had
converged. Two candidates within a float of each other flip that label back
and forth, so the sweep kept running long after the exemplars had settled and
a fit spent its whole iteration budget. It watches the exemplar set now, which
is what scikit-learn watches. 15.7 ms to 5.1 ms on this machine, against 6.7
on the runner.

kernel_pca copied row i into a fresh allocation and row j into another one for
every pair, so a fit at 150 samples made 22500 allocations and read the design
through matrix_at 180000 times to build a matrix that is symmetric. The kernel
is one contiguous block now, built off row pointers and mirrored, the power
iteration is an sgemv against it with one buffer for the whole loop, and
transform gets the same treatment. 3.71 ms to 0.27 ms.

svr recomputed the decision value of every support vector from the kernel
matrix to average its bias, when the coordinate descent above it had kept
exactly that value up to date the whole way down. The kernel between two rows
also read both of them through matrix_at, which is two struct copies per
feature per pair in every kernel prediction. 13.9 ms to 4.1 ms, and nu_svr,
one_class_svm and svc come along with the kernel change.

The harness also repeats a fit five times when it lands between 2 ms and 20
ms, where it used to measure it once. svr and kernel_pca were both inside that
band, both were within four percent of scikit-learn, and a single measurement
of a 13 ms fit on a shared runner moves by more than that: svr was recorded at
1.72x and at 0.96x in consecutive runs of code that differs in nothing
touching it.
The bucket table gained shaped and the different_shape row means something
else now. The reproduction steps were a shell loop that swallowed a dead
chunk's exit status; they are the runner and the gate.
The gate passed with every ranked row in front of scikit-learn, and affinity
propagation was the narrowest of them at 1.003x. A row that close is a coin
toss on the next run, so it gets the work taken out of its inner loops.

The column sums the availability pass needs are accumulated as the
responsibilities are written, which removes a whole sweep over the matrix and a
walk down its columns. The term that depends only on the column is computed
once per column rather than once per pair, and the diagonal is lifted out of
the inner loop, which used to test i against k for every element.

Local timing on a machine at load 20 cannot see the difference: 4.84 ms
against 4.85 before, inside the spread of repeated runs of the same binary.
None of it can be slower, and the runner will say.
188 ranked rows, 188 of them in front of scikit-learn, from the Wide estimator
matrix job on ubuntu-latest with OpenBLAS pinned to four threads. Nothing is
unranked for want of a reading: no row reported no timing and no row landed
under the clock's floor.

The narrowest is affinity propagation at 1.108x and the median is 27.7x. The
widest say more about the defaults than about the code, which the page already
says in its own words.

The README count heals from this artifact the way the canonical counts already
do, so the prose cannot drift from the number again.
@godofecht godofecht changed the title Beat scikit-learn on every ranked estimator row, and gate it in CI Beat scikit-learn on all 188 ranked estimator rows, gated in CI Sep 28, 2026
@godofecht
godofecht merged commit 01c5c6b into main Sep 28, 2026
9 of 10 checks passed
@godofecht
godofecht deleted the beat-every-row branch September 28, 2026 20:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant