Race the eleven rows that were excluded, by making them do the job - #511
Merged
Merged
Conversation
Four implementations said in their own comments that they were simplified, so the registry timed them and showed the numbers without ranking them, and the page carried a paragraph explaining why three of the four would otherwise have been its widest wins. They were wide because they were not doing the work. spectral_coclustering bucketed each row sum between the smallest and the largest. It runs Dhillon's algorithm now: normalize by the square roots of the row and column sums, take the singular vectors after the trivial one, and cluster rows and columns together in the space they span. 0.03 ms against scikit-learn's 8.09. spectral_biclustering bucketed row means and column means the same way. It runs Kluger's: Sinkhorn-Knopp, then rank the singular vectors by how well a one dimensional k-means fits each one, keep the best few, and cluster rows and columns separately. 0.09 ms against 35.17. elliptic_envelope drew random subsets and scored each by the product of the diagonal of its covariance, which is the determinant only when the features do not correlate, and stopped at the draw. The concentration step is the whole content of FastMCD, so it takes the h points closest in Mahalanobis distance, recomputes, and repeats until the determinant stops falling. The determinant comes from a Cholesky factor, and predict uses the full distance rather than the diagonal. 2.03 ms against 19.46. nca computed neither the objective nor its gradient. Its distance summed the squares of the individual products instead of squaring the projected difference, which is not a distance in the projected space, and its gradient kept only the term that pulls same-class points together, with nothing pushing back. Both are written out now. 181 ms against 856. Two test files cover them, by behaviour rather than by shape: the block structure has to come back out of both biclusterings, every planted outlier has to be found with no false alarms, and the projection has to separate the classes better than the identity did. Each was checked by inverting an assertion and confirming a non-zero exit.
Five Flow functions covered part of what the scikit-learn class they are named
after does, so the registry kept them out of the ranking and said why. The
missing half is written now, as a second entry point beside each of them, and
the five are raced.
voting_classifier_full_fit and voting_regressor_full_fit fit the ensemble from
the data, where the functions beside them take estimators that are already
fitted and their cost is the vote alone. Three trees on both sides, because a
vote of one is not what either library is being asked for.
stacking_classifier_full_fit cross-validates its base estimators to build the
meta-features, fits the meta-learner on those, and refits the base estimators
on all of the data. Out of fold predictions are the point: a base estimator's
prediction for a row it helped fit would tell the meta-learner more than it
will know at prediction time.
self_training_classifier_full_fit refits its base estimator on every round,
which is the expensive half the function beside it leaves to its caller.
incremental_pca_fit walks the design in batches, sized as scikit-learn sizes
its own, where the partial fit beside it is one step of that.
Racing incremental PCA turned up that it did not work. Its partial fit
accumulated the same per-feature quantity for every component, so every
component came out as the same vector, and the step it called a power
iteration multiplied a component by itself rather than by a covariance. It
keeps a running scatter matrix now, which combines across batches exactly once
the shift between two means is accounted for, and takes its components from
the leading eigenvectors of that. On the test fixture the leading component
carried a spread of 640 against the second's 4, where the two were previously
identical to the digit.
Local timings against scikit-learn on the same data:
incremental_pca 0.010 ms against 1.01
voting_classifier_full 0.053 against 2.23
self_training_classifier_full 0.266 against 19.83
stacking_classifier_full 0.529 against 16.07
voting_regressor_full 1.184 against 3.74
A test file covers all five by behaviour: the two ensembles and the stack have
to classify or regress the fixture, self training has to label the unlabelled
rows and get them right, and the batched decomposition has to see every row and
put more spread in the leading component than the second. It was checked as a
real gate by inverting an assertion.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this closes
Eleven of the 203 exported estimators were not ranked. Six of those have no
scikit-learn counterpart and never will. The other nine were excluded because
of what our own implementation did, and that is what this fixes.
The four that were stand-ins
Each said so in its own comments, so the registry timed it and showed the
number without ranking it, and the page carried a paragraph explaining why
three of the four would otherwise have been its widest wins. They were wide
because they were not doing the work.
Those four ratios are from the gate run on this branch, which reported 192 of
192 with the simplified bucket empty.
The five that did part of the job
A second entry point beside each one does the whole thing.
voting_classifier_full_fit,voting_regressor_full_fit: fit the ensemblefrom the data. Three trees on both sides, because a vote of one is not what
either library is being asked for.
stacking_classifier_full_fit: cross-validates the base estimators to buildout-of-fold meta-features, fits the meta-learner, refits the base estimators.
self_training_classifier_full_fit: refits its base estimator every round.incremental_pca_fit: walks the design in batches, sized as scikit-learnsizes its own.
The older functions stay, each with a reason naming the full fit that replaces
it for racing purposes.
A bug racing them found
IncrementalPCA did not work. Its partial fit accumulated the same per-feature
quantity for every component, so every component came out as the same vector,
and the step it called a power iteration multiplied a component by itself
rather than by a covariance. It keeps a running scatter matrix now, exact
across batches once the shift between two means is accounted for, and takes
its components from the leading eigenvectors of it. On the fixture the leading
component carries a spread of 640 against the second's 4, where the two were
previously identical to the digit.
Nothing caught this before because the estimator was never raced and no test
asserted on its output.
Tests
Three new files, all checking behaviour:
block and the column block it belongs with agreeing;
training has to label the unlabelled rows and get them right, and the
batched decomposition has to see every row and rank its components.
Each was verified as a real gate by inverting an assertion and confirming a
non-zero exit.