Skip to content

Race the eleven rows that were excluded, by making them do the job - #511

Merged
godofecht merged 2 commits into
mainfrom
real-algorithms
Sep 29, 2026
Merged

godofecht merged 2 commits into
mainfrom
real-algorithms

Conversation

@godofecht

@godofecht godofecht commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

What this closes

Eleven of the 203 exported estimators were not ranked. Six of those have no
scikit-learn counterpart and never will. The other nine were excluded because
of what our own implementation did, and that is what this fixes.

bucket before after
raced 172 197
simplified stand-ins 4 0
doing part of the job 5 0
no counterpart 6 6

The four that were stand-ins

Each said so in its own comments, so the registry timed it and showed the
number without ranking it, and the page carried a paragraph explaining why
three of the four would otherwise have been its widest wins. They were wide
because they were not doing the work.

row was is Flow scikit-learn
spectral_coclustering row sums bucketed between min and max Dhillon: normalized adjacency, singular vectors after the trivial one, one k-means over rows and columns together 0.04 ms 4.73
spectral_biclustering row and column means bucketed the same way Kluger: Sinkhorn-Knopp, vectors ranked by how well a one dimensional k-means fits them, rows and columns clustered separately 0.15 30.60
elliptic_envelope random subsets scored by the product of the covariance diagonal FastMCD concentration steps, determinant from a Cholesky factor, full Mahalanobis distance in predict 0.80 18.36
nca a distance that squared each product before summing, and a gradient missing its expectation term the actual objective and its gradient 251.9 1167.1

Those four ratios are from the gate run on this branch, which reported 192 of
192 with the simplified bucket empty.

The five that did part of the job

A second entry point beside each one does the whole thing.

  • voting_classifier_full_fit, voting_regressor_full_fit: fit the ensemble
    from the data. Three trees on both sides, because a vote of one is not what
    either library is being asked for.
  • stacking_classifier_full_fit: cross-validates the base estimators to build
    out-of-fold meta-features, fits the meta-learner, refits the base estimators.
  • self_training_classifier_full_fit: refits its base estimator every round.
  • incremental_pca_fit: walks the design in batches, sized as scikit-learn
    sizes its own.
row Flow scikit-learn
incremental_pca 0.010 ms 1.01
voting_classifier_full 0.053 2.23
self_training_classifier_full 0.266 19.83
stacking_classifier_full 0.529 16.07
voting_regressor_full 1.184 3.74

The older functions stay, each with a reason naming the full fit that replaces
it for racing purposes.

A bug racing them found

IncrementalPCA did not work. Its partial fit accumulated the same per-feature
quantity for every component, so every component came out as the same vector,
and the step it called a power iteration multiplied a component by itself
rather than by a covariance. It keeps a running scatter matrix now, exact
across batches once the shift between two means is accounted for, and takes
its components from the leading eigenvectors of it. On the fixture the leading
component carries a spread of 640 against the second's 4, where the two were
previously identical to the digit.

Nothing caught this before because the estimator was never raced and no test
asserted on its output.

Tests

Three new files, all checking behaviour:

  • both biclusterings have to recover a planted block structure, with the row
    block and the column block it belongs with agreeing;
  • every planted outlier has to be found with no false alarms;
  • the learned projection has to separate the classes better than the identity;
  • the ensembles and the stack have to classify or regress their fixture, self
    training has to label the unlabelled rows and get them right, and the
    batched decomposition has to see every row and rank its components.

Each was verified as a real gate by inverting an assertion and confirming a
non-zero exit.

Four implementations said in their own comments that they were simplified, so
the registry timed them and showed the numbers without ranking them, and the
page carried a paragraph explaining why three of the four would otherwise have
been its widest wins. They were wide because they were not doing the work.

spectral_coclustering bucketed each row sum between the smallest and the
largest. It runs Dhillon's algorithm now: normalize by the square roots of the
row and column sums, take the singular vectors after the trivial one, and
cluster rows and columns together in the space they span. 0.03 ms against
scikit-learn's 8.09.

spectral_biclustering bucketed row means and column means the same way. It
runs Kluger's: Sinkhorn-Knopp, then rank the singular vectors by how well a one
dimensional k-means fits each one, keep the best few, and cluster rows and
columns separately. 0.09 ms against 35.17.

elliptic_envelope drew random subsets and scored each by the product of the
diagonal of its covariance, which is the determinant only when the features do
not correlate, and stopped at the draw. The concentration step is the whole
content of FastMCD, so it takes the h points closest in Mahalanobis distance,
recomputes, and repeats until the determinant stops falling. The determinant
comes from a Cholesky factor, and predict uses the full distance rather than
the diagonal. 2.03 ms against 19.46.

nca computed neither the objective nor its gradient. Its distance summed the
squares of the individual products instead of squaring the projected
difference, which is not a distance in the projected space, and its gradient
kept only the term that pulls same-class points together, with nothing pushing
back. Both are written out now. 181 ms against 856.

Two test files cover them, by behaviour rather than by shape: the block
structure has to come back out of both biclusterings, every planted outlier has
to be found with no false alarms, and the projection has to separate the
classes better than the identity did. Each was checked by inverting an
assertion and confirming a non-zero exit.
Five Flow functions covered part of what the scikit-learn class they are named
after does, so the registry kept them out of the ranking and said why. The
missing half is written now, as a second entry point beside each of them, and
the five are raced.

voting_classifier_full_fit and voting_regressor_full_fit fit the ensemble from
the data, where the functions beside them take estimators that are already
fitted and their cost is the vote alone. Three trees on both sides, because a
vote of one is not what either library is being asked for.

stacking_classifier_full_fit cross-validates its base estimators to build the
meta-features, fits the meta-learner on those, and refits the base estimators
on all of the data. Out of fold predictions are the point: a base estimator's
prediction for a row it helped fit would tell the meta-learner more than it
will know at prediction time.

self_training_classifier_full_fit refits its base estimator on every round,
which is the expensive half the function beside it leaves to its caller.

incremental_pca_fit walks the design in batches, sized as scikit-learn sizes
its own, where the partial fit beside it is one step of that.

Racing incremental PCA turned up that it did not work. Its partial fit
accumulated the same per-feature quantity for every component, so every
component came out as the same vector, and the step it called a power
iteration multiplied a component by itself rather than by a covariance. It
keeps a running scatter matrix now, which combines across batches exactly once
the shift between two means is accounted for, and takes its components from
the leading eigenvectors of that. On the test fixture the leading component
carried a spread of 640 against the second's 4, where the two were previously
identical to the digit.

Local timings against scikit-learn on the same data:

    incremental_pca                 0.010 ms  against  1.01
    voting_classifier_full          0.053     against  2.23
    self_training_classifier_full   0.266     against 19.83
    stacking_classifier_full        0.529     against 16.07
    voting_regressor_full           1.184     against  3.74

A test file covers all five by behaviour: the two ensembles and the stack have
to classify or regress the fixture, self training has to label the unlabelled
rows and get them right, and the batched decomposition has to see every row and
put more spread in the leading component than the second. It was checked as a
real gate by inverting an assertion.
@godofecht godofecht changed the title Run the algorithm in the four rows that were standing in for one Race the eleven rows that were excluded, by making them do the job Sep 29, 2026
@godofecht
godofecht merged commit 33df6df into main Sep 29, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant