Beat scikit-learn on all 188 ranked estimator rows, gated in CI - #509
Merged
Merged
Conversation
Five SVM and ridge rows sat between 0.03x and 0.49x. Each had the same shape of problem: work repeated that did not need repeating. svr walked every coefficient by a fixed 0.01 whenever it saw a violation and recomputed the decision value from scratch for every sample on each of up to 1000 passes. Minimising the epsilon-insensitive objective in one coordinate has a closed form, a soft threshold of the partial residual clipped to the box, and the decision values can be carried forward as each coefficient moves. 66.4 ms to 2.5 ms, 0.10x to 2.6x. nu_svr did the same and rebuilt the RBF kernel inside the innermost loop, taking an exponential for every pair of samples on every pass for a matrix that never changes. 128.6 ms to 2.3 ms, 0.03x to 1.9x. one_class_svm rebuilt the kernel the same way, in the fit loop and again in the threshold pass. The update rule and step are unchanged, so the result is the same. 8.1 ms to 0.19 ms, 0.06x to 2.7x. ridge_cv and ridge_classifier_cv swept an alpha grid with the fold loop inside, rebuilding each training matrix once per alpha through matrix_at, which copies the Matrix struct per element. Worse, every alpha rebuilt the Gram matrix, which does not depend on alpha. _ridge_gram and _ridge_solve_gram split those apart: one Gram per fold, one cheap solve per alpha. 0.72 ms to 0.11 ms and 25.7 ms to 0.70 ms, 0.49x to 3.2x and 0.04x to 1.4x. linear_svr carried the regularizer as a scale instead of writing it across the weight vector once per sample, took its elements off the row pointer, and stops when every sample is inside the tube. 0.39 ms to 0.27 ms, which is 0.88x and still short of liblinear.
elastic_net_fit was proximal gradient with a fixed step and a fixed budget of epochs, allocating its gradient buffer inside the epoch loop. At the small end of a regularization path that converges slowly, so every fit ran all thousand epochs however close it already was. It is cyclic coordinate descent now, which is what cd_fast does and which settles in tens of sweeps at any alpha. elastic_net_cv goes from 167.8 ms to 2.5 ms, 0.08x to 5.5x, and elastic_net itself from 0.39 ms to 0.028 ms. multitask_lasso and multitask_elastic_net allocated a gradient buffer per task on every epoch, read the design through matrix_at, and had no way to stop. 0.45 ms to 0.022 ms and 0.45 ms to 0.013 ms, 0.48x to 9.9x and 0.46x to 16x. multitask_lasso_cv follows them from 0.59x to 1.7x. affinity_propagation rescanned a whole row to find the largest A + S excluding one column, for every column, which made the responsibility pass cubic in the sample count. The row maximum and runner up answer every column at once. Both availability passes walked down a column of a row-pointer array and now run row by row. 139.3 ms to 6.4 ms, 0.03x to 0.64x. local_outlier_factor bubble sorted every row of the distance matrix, twice, once per sample, so it was cubic as well. Only the k nearest are needed, so it selects the first k + 1 positions instead. The distance matrix is symmetric and is now built as one triangle off row pointers. 5.4 ms to 0.73 ms, 0.07x to 0.54x. Both of those last two are still behind scikit-learn. They were 30x and 14x behind.
…atrix_at hdbscan built the same Euclidean distances three ways over. The core-distance pass rebuilt every row from the raw features and then ran a full selection sort over it to read one element, which is cubic in the sample count, and Prim's loop rebuilt each distance again for every pair it considered. One symmetric distance matrix now serves all of it and the core pass selects only the positions it needs. 4.5 ms to 0.41 ms, 0.16x to 1.75x. spectral_embedding builds its affinity matrix as one triangle off row pointers, and its power iteration takes the row pointer once and reuses one matvec buffer rather than allocating on every step. None of that moved the number: the cost is 200 power iterations that do not converge, and beating ARPACK there needs a different eigensolver. It stays at 0.30x. multitask_elastic_net_cv no longer predicts over the whole design for every fold and every alpha when it scores five test folds, and builds each training pair once for the grid. That did not move it either. The cost is the fold fits themselves, which are still proximal gradient and still run their whole epoch budget at the small end of the path. Block coordinate descent is the fix and is not in this commit. It stays at 0.28x.
The wide matrix had eleven rows where scikit-learn was faster. Three of them were already fixed in the last two commits and needed no further work. These are the rest, and in each case the loss came from the algorithm rather than from the loop it was written in. MultiTaskElasticNet moves from a fixed step proximal gradient to block coordinate descent. Each feature block has a closed form minimizer given the others, so a sweep solves every block exactly, and the residual is carried through a rank one correction instead of being rebuilt. The cross validated version walks its alpha grid from the largest alpha down and warm starts each fit from the last, so the whole grid costs little more than its first point. 62.6 ms to 3.7 ms on the diabetes shape, measured on a loaded machine. MDS moves from gradient descent at a fixed hundredth of a step to the SMACOF Guttman transform, which is the exact minimizer of the majorizing function at the current configuration. Each pair is visited once and its contribution is added to both of its points. The old stopping rule compared raw stress against the tolerance and was never true at these scales, so every fit ran its whole iteration budget. SpectralEmbedding and LLE move from one scalar matvec per component per iteration to one sgemm per iteration over the whole block, with modified Gram Schmidt keeping the block orthonormal. LLE also builds its W^T W term with one sgemm where it used to walk a scalar dot per entry down a column of an array of row allocations, and picks its neighbours by partial selection instead of bubble sorting every row, which was cubic in the sample count on its own. KernelRidge builds the linear, polynomial and sigmoid kernels from one sgemm. The old path walked a scalar dot per pair, wrote every entry through matrix_set, and then read all of it back through matrix_at. A linear kernel also has a primal form, so predict folds the dual coefficients into one weight vector rather than evaluating an n by n_train kernel. KNeighborsTransformer drops a zero fill that matrix_new had already done through n by n_train calls to matrix_set, and reads its distances off row pointers. ClassifierChain does the same for the design it rebuilds per output in predict. LinearSVR predicts with one sgemv.
The 172 row matrix was driven by a shell loop typed out by hand, which is how a published artifact sat at 145 of 166 while the working tree was at 155, and how a chunk that died mid-run once passed for a chunk with no rows in it. run_estimator_bench.py compiles and runs the generated chunks with every exit status checked, and reuses the binary that the first round leaves behind so later rounds pay for timing instead of for compiling the library again. check_estimator_matrix.py fails on any ranked row slower than scikit-learn, on any row that reported no timing, and on a ranked count that has quietly shrunk. Both run in a new Flow workflow job on a runner that is not competing with anything, which is where a published number belongs. The generator and the scikit-learn harness now time the four estimators the registry marks simplified. compare_estimators.py already showed their times while withholding a ratio, so leaving them out of the race dropped four rows from the page for no reason. The generator also takes a comma separated --only and a --prefix, so a subset can be re-timed without colliding with the committed files in the shared build directory.
The wide matrix raced 172 of 203 exported estimators. Twenty-five sat in different_shape, which meant only that their fit does not begin with a feature matrix and the generic path builds one shape of call. Thirteen of them race scikit-learn perfectly well once the call is written out, so the registry now carries that call and they are ranked like any other row: the seven kernel approximation and projection transformers, the two dummy estimators, the two label encoders, the isotonic fit and the model selector. The twelve that are left take a Pipeline, a ColumnTransformer, a FeatureUnion, an estimator array, a document array or a dict array, and each keeps its written reason. factor_analysis_fit allocated its component array with malloc, tested each row against null, and wrote through whatever the allocator had left there when the test failed. On macOS that memory happened to be zero, so the fit worked and the row has been timed for weeks. On the CI runner it was not, and the process died partway through its chunk, taking the nine estimators after it down with it. Every row is allocated up front now. The harness also adds a sink. Two rows came back at exactly 0.000000 ms because clang at -O3 is free to delete a transform whose result is freed without being read. One value from every result is now folded into a running total that main prints on a line the parser ignores.
The composed rows and the input-shaped rows now carry their call in the registry the same way the other shaped rows do: Pipeline over a scaler and a logistic regression, ColumnTransformer over a scaler, FeatureUnion over a scaler and a passthrough, the two text vectorizers over one corpus, the dict vectorizer and the multi-label binarizer. That is 192 of 203 raced. The corpus lives in the registry, so the generated Flow file and the scikit-learn harness read one copy of it and cannot drift apart. A fit that takes a composed object built outside the timing frees once after the timing rather than on every repeat, because the object it hands back is the object the next repeat reads. _hash_string carried djb2 in an i32 and shifted the result left every character. Past the sixth character of a token the value had wrapped negative, and the runtime traps a left shift of a negative, so the process aborted inside count_vectorizer_fit. The four document fixture in the tests never reached that. A sixteen document corpus did. It is carried in 64 bits and reduced every step now. The five that are left are the ones where the Flow function does part of what the scikit-learn class does: the two voting estimators take fitted models, the stacking classifier takes base predictions, the self-training classifier takes probabilities, and incremental PCA takes one batch. Each says so in its own reason rather than in a generic message about its first argument.
The first CI run of the wide matrix put 175 of 179 ranked rows in front of scikit-learn and named the four that were behind. Every one of them is a different story from the macOS numbers, which is the point of gating on the runner rather than on a developer machine. multitask_lasso_cv, 109.9 ms against 59.7. It now runs the same block coordinate descent as the elastic net, warm started down the alpha path. The sweep loop also gained a real convergence test: the coefficient move it used was weak at a small penalty, where the iterate creeps and no single coefficient moves far in one sweep, so a fit spent its whole budget on an objective that had stopped falling. The objective comes out of the residual the sweep already maintains, for about half a percent of a sweep. 109.9 ms to 3.6 ms on this machine, and the elastic net path gets the same test. affinity_propagation, 15.5 ms against 4.9. The three n by n matrices are one allocation each rather than n row allocations, the similarity is built off row pointers and mirrored, and the label pass is folded into the availability pass that already has both terms in hand. tweedie_regressor, 1.31 ms against 0.98, with the Poisson and gamma fits on the same loop. All three walked the design twice per row through matrix_at and allocated a gradient buffer inside the iteration. One sgemv for the linear predictor, one transposed sgemv for the gradient, buffers allocated once. linear_svr, 0.66 ms against 0.42. The dot product that its epoch is made of accumulated into one register, so the loop was a chain of dependent additions. Four accumulators. kernel_density and nearest_neighbors sat unranked because their work function is called score_samples and kneighbors, so the generic path found nothing to time and the fit alone fell under the clock's floor. A recipe can now name the work function and the scikit-learn method to match it against. That puts every row the matrix can rank in the ranking.
The second CI run put 185 of 188 ranked rows in front of scikit-learn. These are the three it named. affinity_propagation watched the argmax label of every row to decide it had converged. Two candidates within a float of each other flip that label back and forth, so the sweep kept running long after the exemplars had settled and a fit spent its whole iteration budget. It watches the exemplar set now, which is what scikit-learn watches. 15.7 ms to 5.1 ms on this machine, against 6.7 on the runner. kernel_pca copied row i into a fresh allocation and row j into another one for every pair, so a fit at 150 samples made 22500 allocations and read the design through matrix_at 180000 times to build a matrix that is symmetric. The kernel is one contiguous block now, built off row pointers and mirrored, the power iteration is an sgemv against it with one buffer for the whole loop, and transform gets the same treatment. 3.71 ms to 0.27 ms. svr recomputed the decision value of every support vector from the kernel matrix to average its bias, when the coordinate descent above it had kept exactly that value up to date the whole way down. The kernel between two rows also read both of them through matrix_at, which is two struct copies per feature per pair in every kernel prediction. 13.9 ms to 4.1 ms, and nu_svr, one_class_svm and svc come along with the kernel change. The harness also repeats a fit five times when it lands between 2 ms and 20 ms, where it used to measure it once. svr and kernel_pca were both inside that band, both were within four percent of scikit-learn, and a single measurement of a 13 ms fit on a shared runner moves by more than that: svr was recorded at 1.72x and at 0.96x in consecutive runs of code that differs in nothing touching it.
The bucket table gained shaped and the different_shape row means something else now. The reproduction steps were a shell loop that swallowed a dead chunk's exit status; they are the runner and the gate.
The gate passed with every ranked row in front of scikit-learn, and affinity propagation was the narrowest of them at 1.003x. A row that close is a coin toss on the next run, so it gets the work taken out of its inner loops. The column sums the availability pass needs are accumulated as the responsibilities are written, which removes a whole sweep over the matrix and a walk down its columns. The term that depends only on the column is computed once per column rather than once per pair, and the diagonal is lifted out of the inner loop, which used to test i against k for every element. Local timing on a machine at load 20 cannot see the difference: 4.84 ms against 4.85 before, inside the spread of repeated runs of the same binary. None of it can be slower, and the runner will say.
188 ranked rows, 188 of them in front of scikit-learn, from the Wide estimator matrix job on ubuntu-latest with OpenBLAS pinned to four threads. Nothing is unranked for want of a reading: no row reported no timing and no row landed under the clock's floor. The narrowest is affinity propagation at 1.108x and the median is 27.7x. The widest say more about the defaults than about the code, which the page already says in its own words. The README count heals from this artifact the way the canonical counts already do, so the prose cannot drift from the number again.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Result
188 ranked rows, 188 of them in front of scikit-learn, measured by a new CI
job on a runner that is not competing with anything. No row reported no timing,
and no row landed under the clock's floor.
Narrowest: affinity propagation 1.108x, svr 1.227x, nu_svr 1.324x, isomap
1.326x. Median 27.7x. The widest ratios say more about the two libraries'
defaults than about the code, which the page says in its own words.
Coverage went from 172 of the 203 exported estimators to 192. What is left
carries a specific reason rather than a generic one: four implementations say
in their own comments that they are simplified, so they are timed and shown
without being ranked; five Flow functions do part of what their scikit-learn
namesake does, such as a voting estimator that takes models already fitted;
six have no scikit-learn counterpart.
The gate
benchmarks/run_estimator_bench.pycompiles and runs the generated chunks withevery exit status checked, and reuses the binary the first round leaves behind.
benchmarks/check_estimator_matrix.pyfails on a ranked row slower thanscikit-learn, on a row that reported no timing, and on a ranked count that has
quietly shrunk. The
Wide estimator matrixjob runs both.The gate paid for itself immediately. Its first run found four rows that lose
on Linux and win on macOS, its second found three more, and both sets are fixed
here.
The rows this fixed
Two crashes the breadth found
factor_analysis_fitallocated its component array with malloc, tested each rowagainst null and wrote through whatever the allocator had left there. On macOS
that memory happened to be zero, so the fit worked and the row had been timed
for weeks. On the runner it was not, and the process died partway through its
chunk, taking the nine estimators after it down with it.
_hash_stringcarried djb2 in an i32 and shifted it left every character. Pastthe sixth character the value had wrapped negative, and the runtime traps a left
shift of a negative. The four document fixture in the tests never reached that.
A sixteen document corpus did.
Verification
Wide estimator matrix: success, 188/188.Run all tests and examples: success, the full suite on Linux.Canonical benchmark contract,Scaled benchmark report,FLOW syntax and imports,Python native package: all success.