From 9ad67b5fbdc503ed097894bd6be39f21fe98150a Mon Sep 17 00:00:00 2001 From: godofecht Date: Sun, 27 Sep 2026 20:05:36 +0100 Subject: [PATCH 01/12] Fix the six worst solvers the wide benchmark exposed Five SVM and ridge rows sat between 0.03x and 0.49x. Each had the same shape of problem: work repeated that did not need repeating. svr walked every coefficient by a fixed 0.01 whenever it saw a violation and recomputed the decision value from scratch for every sample on each of up to 1000 passes. Minimising the epsilon-insensitive objective in one coordinate has a closed form, a soft threshold of the partial residual clipped to the box, and the decision values can be carried forward as each coefficient moves. 66.4 ms to 2.5 ms, 0.10x to 2.6x. nu_svr did the same and rebuilt the RBF kernel inside the innermost loop, taking an exponential for every pair of samples on every pass for a matrix that never changes. 128.6 ms to 2.3 ms, 0.03x to 1.9x. one_class_svm rebuilt the kernel the same way, in the fit loop and again in the threshold pass. The update rule and step are unchanged, so the result is the same. 8.1 ms to 0.19 ms, 0.06x to 2.7x. ridge_cv and ridge_classifier_cv swept an alpha grid with the fold loop inside, rebuilding each training matrix once per alpha through matrix_at, which copies the Matrix struct per element. Worse, every alpha rebuilt the Gram matrix, which does not depend on alpha. _ridge_gram and _ridge_solve_gram split those apart: one Gram per fold, one cheap solve per alpha. 0.72 ms to 0.11 ms and 25.7 ms to 0.70 ms, 0.49x to 3.2x and 0.04x to 1.4x. linear_svr carried the regularizer as a scale instead of writing it across the weight vector once per sample, took its elements off the row pointer, and stops when every sample is inside the tube. 0.39 ms to 0.27 ms, which is 0.88x and still short of liblinear. --- lib/scikit/linear.flow | 283 ++++++++++++++++++++++++++--------------- lib/scikit/svm.flow | 210 +++++++++++++++++++++--------- 2 files changed, 324 insertions(+), 169 deletions(-) diff --git a/lib/scikit/linear.flow b/lib/scikit/linear.flow index 3795fad..85a6fbe 100644 --- a/lib/scikit/linear.flow +++ b/lib/scikit/linear.flow @@ -5096,85 +5096,155 @@ export struct RidgeCV { fitted: bool } +# One Gram matrix, many alphas. +# +# ridge_fit rebuilds XtX and Xty from the design on every call, which is the +# whole cost of a ridge solve at these shapes. Neither depends on alpha, so a +# cross-validation that sweeps a grid was doing that work once per alpha when +# once per fold is enough. This builds the two of them, and _ridge_solve_gram +# below turns a copy plus one alpha into coefficients. +function _ridge_gram(X: Matrix, y: ptr, XtX: ptr, Xty: ptr, x_mean: ptr) -> f64 { + let n: i32 = X.cols + let m: i32 = X.rows + let xdata: ptr = X.data + + let mut y_mean: f64 = 0.0 + for j in 0 to n { x_mean[j] = 0.0 } + for i in 0 to m { + y_mean = y_mean + (y[i] as f64) + let base: i32 = i * n + for j in 0 to n { + x_mean[j] = x_mean[j] + (xdata[base + j] as f64) + } + } + y_mean = y_mean / (m as f64) + for j in 0 to n { x_mean[j] = x_mean[j] / (m as f64) } + + for j in 0 to n * n { XtX[j] = 0.0 } + for j in 0 to n { Xty[j] = 0.0 } + + let xrow: ptr = array_new_f64(n) + for i in 0 to m { + let base: i32 = i * n + for j in 0 to n { + xrow[j] = (xdata[base + j] as f64) - x_mean[j] + } + let yi: f64 = (y[i] as f64) - y_mean + for j in 0 to n { + Xty[j] = Xty[j] + xrow[j] * yi + } + blas_ger_ld_f64(XtX, n, n, n, xrow, 1, xrow, 1.0) + } + array_free_f64(xrow) + return y_mean +} + +# Solves (XtX + alpha I) w = Xty without disturbing either input, so the same +# Gram matrix serves every alpha on the grid. +function _ridge_solve_gram(XtX: ptr, Xty: ptr, x_mean: ptr, y_mean: f64, + n: i32, alpha: f32, weights: ptr) -> f32 { + let A: ptr = array_new_f64(n * n) + let rhs: ptr = array_new_f64(n) + for j in 0 to n * n { A[j] = XtX[j] } + for j in 0 to n { rhs[j] = Xty[j] } + let alpha_d: f64 = alpha as f64 + for j in 0 to n { + A[j * n + j] = A[j * n + j] + alpha_d + } + _solve_f64(A, rhs, n) + + let mut bias_d: f64 = y_mean + for j in 0 to n { + bias_d = bias_d - x_mean[j] * rhs[j] + weights[j] = rhs[j] as f32 + } + array_free_f64(A) + array_free_f64(rhs) + return bias_d as f32 +} + +# The fold loop is outside the alpha loop, so each fold's training matrix is +# built once rather than once per alpha. The old order rebuilt it n_alphas +# times over, copying every element through matrix_at and matrix_set, which +# each copy the Matrix struct by value. On ten alphas that is nine tenths of +# the copying thrown away. export function ridge_cv_fit(X: Matrix, y: ptr, n_alphas: i32) -> RidgeCV { let n: i32 = X.rows let n_features: i32 = X.cols + let xdata: ptr = X.data - # Generate alpha grid (logspace from 0.001 to 10) let alphas: ptr = array_new_f32(n_alphas) - let mut i: i32 = 0 - while i < n_alphas { + for i in 0 to n_alphas { let t: f32 = (i as f32) / ((n_alphas - 1) as f32) alphas[i] = 0.001 * pow(10.0 as f64, (t * 4.0) as f64) as f32 - i = i + 1 } - # Try each alpha, compute 5-fold CV MSE - let mut best_alpha: f32 = alphas[0] - let mut best_mse: f32 = 999999.0 + # Squared error per alpha, summed over folds. + let sse: ptr = array_new_f32(n_alphas) let fold_size: i32 = n / 5 - let mut a_idx: i32 = 0 - while a_idx < n_alphas { - let alpha: f32 = alphas[a_idx] - let mut total_mse: f32 = 0.0 - let mut fold: i32 = 0 - while fold < 5 { - let test_start: i32 = fold * fold_size - let test_end: i32 = test_start + fold_size - if fold == 4 { test_end = n } - let n_train: i32 = n - (test_end - test_start) - - let X_train: Matrix = matrix_new(n_train, n_features) - let y_train: ptr = array_new_f32(n_train) - let mut idx: i32 = 0 - let mut j: i32 = 0 - while j < n { - if j < test_start || j >= test_end { - let mut k: i32 = 0 - while k < n_features { - matrix_set(X_train, idx, k, matrix_at(X, j, k)) - k = k + 1 - } - y_train[idx] = y[j] - idx = idx + 1 - } - j = j + 1 + for fold in 0 to 5 { + let test_start: i32 = fold * fold_size + let mut test_end: i32 = test_start + fold_size + if fold == 4 { test_end = n } + let n_test: i32 = test_end - test_start + let n_train: i32 = n - n_test + + let X_train: Matrix = matrix_new(n_train, n_features) + let y_train: ptr = array_new_f32(n_train) + let train_data: ptr = X_train.data + let mut idx: i32 = 0 + for j in 0 to n { + if j < test_start || j >= test_end { + let src: ptr = xdata + j * n_features + let dst: ptr = train_data + idx * n_features + for k in 0 to n_features { dst[k] = src[k] } + y_train[idx] = y[j] + idx = idx + 1 } + } - let model: Ridge = ridge_fit(X_train, y_train, alpha, 1000, 0.01) + # One Gram matrix for the whole alpha grid on this fold. + let XtX: ptr = array_new_f64(n_features * n_features) + let Xty: ptr = array_new_f64(n_features) + let x_mean: ptr = array_new_f64(n_features) + let y_mean: f64 = _ridge_gram(X_train, y_train, XtX, Xty, x_mean) + let w: ptr = array_new_f32(n_features) + for a_idx in 0 to n_alphas { + let bias: f32 = _ridge_solve_gram(XtX, Xty, x_mean, y_mean, n_features, alphas[a_idx], w) let mut fold_mse: f32 = 0.0 - let n_test: i32 = test_end - test_start - let mut t2: i32 = test_start - while t2 < test_end { - let mut pred: f32 = model.bias - let mut k2: i32 = 0 - while k2 < n_features { - pred = pred + model.weights[k2] * matrix_at(X, t2, k2) - k2 = k2 + 1 + for t2 in test_start to test_end { + let trow: ptr = xdata + t2 * n_features + let mut pred: f32 = bias + for k2 in 0 to n_features { + pred = pred + w[k2] * trow[k2] } let diff: f32 = y[t2] - pred fold_mse = fold_mse + diff * diff - t2 = t2 + 1 } - total_mse = total_mse + fold_mse / (n_test as f32) - - ridge_free(model) - matrix_free(X_train) - array_free_f32(y_train) - fold = fold + 1 + sse[a_idx] = sse[a_idx] + fold_mse / (n_test as f32) } - let avg_mse: f32 = total_mse / 5.0 + array_free_f32(w) + array_free_f64(XtX) + array_free_f64(Xty) + array_free_f64(x_mean) + matrix_free(X_train) + array_free_f32(y_train) + } + + let mut best_alpha: f32 = alphas[0] + let mut best_mse: f32 = 999999.0 + for a_idx in 0 to n_alphas { + let avg_mse: f32 = sse[a_idx] / 5.0 if avg_mse < best_mse { best_mse = avg_mse - best_alpha = alpha + best_alpha = alphas[a_idx] } - a_idx = a_idx + 1 } + array_free_f32(sse) - # Refit with best alpha on all data let best_model: Ridge = ridge_fit(X, y, best_alpha, 1000, 0.01) let coef_copy: ptr = array_copy_f32(best_model.weights, n_features) let intercept_copy: f32 = best_model.bias @@ -5478,84 +5548,85 @@ export struct RidgeClassifierCV { fitted: bool } +# Same restructuring as ridge_cv_fit: the fold loop is outside the alpha loop +# so each training matrix is built once, and the Gram matrix each fold shares +# across the whole grid is built once rather than once per alpha. export function ridge_classifier_cv_fit(X: Matrix, y: ptr, n_classes: i32, n_alphas: i32) -> RidgeClassifierCV { let n: i32 = X.rows let n_features: i32 = X.cols + let xdata: ptr = X.data let alphas: ptr = array_new_f32(n_alphas) - let mut i: i32 = 0 - while i < n_alphas { + for i in 0 to n_alphas { let t: f32 = (i as f32) / ((n_alphas - 1) as f32) alphas[i] = 0.001 * pow(10.0 as f64, (t * 4.0) as f64) as f32 - i = i + 1 } - let mut best_alpha: f32 = alphas[0] - let mut best_acc: f32 = -1.0 + let acc_sum: ptr = array_new_f32(n_alphas) let fold_size: i32 = n / 5 - let mut a_idx: i32 = 0 - while a_idx < n_alphas { - let alpha: f32 = alphas[a_idx] - let mut total_acc: f32 = 0.0 - let mut fold: i32 = 0 - while fold < 5 { - let test_start: i32 = fold * fold_size - let test_end: i32 = test_start + fold_size - if fold == 4 { test_end = n } - let n_train: i32 = n - (test_end - test_start) - - let X_train: Matrix = matrix_new(n_train, n_features) - let y_train: ptr = array_new_f32(n_train) - let mut idx: i32 = 0 - let mut j: i32 = 0 - while j < n { - if j < test_start || j >= test_end { - let mut k: i32 = 0 - while k < n_features { - matrix_set(X_train, idx, k, matrix_at(X, j, k)) - k = k + 1 - } - y_train[idx] = y[j] - idx = idx + 1 - } - j = j + 1 + for fold in 0 to 5 { + let test_start: i32 = fold * fold_size + let mut test_end: i32 = test_start + fold_size + if fold == 4 { test_end = n } + let n_test: i32 = test_end - test_start + let n_train: i32 = n - n_test + + let X_train: Matrix = matrix_new(n_train, n_features) + let y_train: ptr = array_new_f32(n_train) + let train_data: ptr = X_train.data + let mut idx: i32 = 0 + for j in 0 to n { + if j < test_start || j >= test_end { + let src: ptr = xdata + j * n_features + let dst: ptr = train_data + idx * n_features + for k in 0 to n_features { dst[k] = src[k] } + y_train[idx] = y[j] + idx = idx + 1 } + } - let model: RidgeClassifier = ridge_classifier_fit(X_train, y_train, alpha, 1000, 0.01) + let XtX: ptr = array_new_f64(n_features * n_features) + let Xty: ptr = array_new_f64(n_features) + let x_mean: ptr = array_new_f64(n_features) + let y_mean: f64 = _ridge_gram(X_train, y_train, XtX, Xty, x_mean) + let w: ptr = array_new_f32(n_features) + for a_idx in 0 to n_alphas { + let bias: f32 = _ridge_solve_gram(XtX, Xty, x_mean, y_mean, n_features, alphas[a_idx], w) let mut correct: i32 = 0 - let mut t2: i32 = test_start - while t2 < test_end { - let mut pred_val: f32 = model.bias - let mut k2: i32 = 0 - while k2 < n_features { - pred_val = pred_val + model.weights[k2] * matrix_at(X, t2, k2) - k2 = k2 + 1 + for t2 in test_start to test_end { + let trow: ptr = xdata + t2 * n_features + let mut pred_val: f32 = bias + for k2 in 0 to n_features { + pred_val = pred_val + w[k2] * trow[k2] } - let pred_cls: f32 = 0.0 + let mut pred_cls: f32 = 0.0 if pred_val > 0.5 { pred_cls = 1.0 } if pred_val > 1.5 { pred_cls = 2.0 } - if y[t2] == pred_cls { - correct = correct + 1 - } - t2 = t2 + 1 + if y[t2] == pred_cls { correct = correct + 1 } } - total_acc = total_acc + (correct as f32) / ((test_end - test_start) as f32) - - ridge_classifier_free(model) - matrix_free(X_train) - array_free_f32(y_train) - fold = fold + 1 + acc_sum[a_idx] = acc_sum[a_idx] + (correct as f32) / (n_test as f32) } - let avg_acc: f32 = total_acc / 5.0 + array_free_f32(w) + array_free_f64(XtX) + array_free_f64(Xty) + array_free_f64(x_mean) + matrix_free(X_train) + array_free_f32(y_train) + } + + let mut best_alpha: f32 = alphas[0] + let mut best_acc: f32 = -1.0 + for a_idx in 0 to n_alphas { + let avg_acc: f32 = acc_sum[a_idx] / 5.0 if avg_acc > best_acc { best_acc = avg_acc - best_alpha = alpha + best_alpha = alphas[a_idx] } - a_idx = a_idx + 1 } + array_free_f32(acc_sum) # Refit with best alpha let best_model: RidgeClassifier = ridge_classifier_fit(X, y, best_alpha, 1000, 0.01) diff --git a/lib/scikit/svm.flow b/lib/scikit/svm.flow index e1d1e00..947444b 100644 --- a/lib/scikit/svm.flow +++ b/lib/scikit/svm.flow @@ -589,32 +589,65 @@ export function linear_svr_fit(X: Matrix, y: ptr, C: f32, epsilon: f32, epo let weights: ptr = array_new_f32(n) let bias: f32 = 0.0 + # The regularization term shrinks every weight by the same factor on every + # sample, whichever branch is taken, so it is carried as a scale instead of + # being written across the weight vector n_rows * epochs times. The weights + # are stored as scale * w_hat and folded back once at the end. The hinge + # update only touches the vector when the sample is outside the tube. + # + # Element access goes through the row pointer. matrix_at copies the Matrix + # struct by value for every element, which this loop did twice per sample. + let rows: i32 = X.rows + let xdata: ptr = X.data + let decay: f32 = 1.0 - lr / (rows as f32) + let w_hat: ptr = array_new_f32(n) + let mut scale: f32 = 1.0 + for epoch in 0 to epochs { - for i in 0 to X.rows { - let mut pred: f32 = bias + let mut violations: i32 = 0 + for i in 0 to rows { + let row: ptr = xdata + i * n + let mut dot: f32 = 0.0 for j in 0 to n { - pred = pred + weights[j] * matrix_at(X, i, j) + dot = dot + w_hat[j] * row[j] } + let pred: f32 = scale * dot + bias let err: f32 = pred - y[i] + scale = scale * decay + # A scale that has decayed far enough loses precision in w_hat, so + # it is folded back into the vector before that happens. + if scale < 0.00001 { + for j in 0 to n { w_hat[j] = w_hat[j] * scale } + scale = 1.0 + } + if err > epsilon { + let step: f32 = lr * C / scale for j in 0 to n { - weights[j] = weights[j] - lr * (C * matrix_at(X, i, j) + weights[j] / (X.rows as f32)) + w_hat[j] = w_hat[j] - step * row[j] } bias = bias - lr * C - } elif err < -epsilon { + violations = violations + 1 + } elif err < 0.0 - epsilon { + let step2: f32 = lr * C / scale for j in 0 to n { - weights[j] = weights[j] - lr * (C * -matrix_at(X, i, j) + weights[j] / (X.rows as f32)) + w_hat[j] = w_hat[j] + step2 * row[j] } bias = bias + lr * C - } else { - for j in 0 to n { - weights[j] = weights[j] - lr * weights[j] / (X.rows as f32) - } + violations = violations + 1 } } + # Every sample sits inside the tube, so the hinge term is satisfied and + # further epochs only shrink the weights toward zero. + if violations == 0 { break } } + for j in 0 to n { + weights[j] = w_hat[j] * scale + } + array_free_f32(w_hat) + return LinearSVR { weights: weights, bias: bias, @@ -1874,28 +1907,50 @@ export function nu_svr_fit(X: Matrix, y: ptr, nu: f32, C: f32, gamma: f32, let alpha_max: f32 = C / (nu * (n as f32)) + # The kernel is built once here. It used to be rebuilt inside the + # innermost loop, recomputing the squared distance between two samples + # across every feature and taking an exponential, for every pair, on every + # one of max_iter passes. That is O(max_iter * n^2 * p) arithmetic and + # O(max_iter * n^2) calls to exp for a matrix that never changes. + let K: Matrix = _svm_kernel_matrix(X, 0, gamma, 0, 0.0) + let kdata: ptr = K.data + + # Exact coordinate descent with the decision values kept up to date, in + # place of walking each coefficient by a fixed 0.01 per violation. The + # box here is symmetric and the loss has no insensitive band, so the + # coordinate minimiser is the partial residual over the diagonal. + let f: ptr = array_new_f32(n) + let tol: f32 = 0.0001 + for iter in 0 to max_iter { + let mut max_change: f32 = 0.0 for i in 0 to n { - let mut g_i: f32 = y[i] - b - for j in 0 to n { - if (fabs((alphas[j]) as f64) as f32) < 0.0000000001 { continue } - let mut s: f32 = 0.0 - for k in 0 to p { - let d: f32 = matrix_at(X, i, k) - matrix_at(X, j, k) - s = s + d * d - } - g_i = g_i - alphas[j] * exp((-gamma * s) as f64) as f32 - } - - if g_i > 0 && alphas[i] < alpha_max { - alphas[i] = alphas[i] + 0.01 - } else { - if g_i < 0 && alphas[i] > -alpha_max { - alphas[i] = alphas[i] - 0.01 + let row: ptr = kdata + i * n + let kii: f32 = row[i] + if kii <= 0.0000001 { continue } + + let a_old: f32 = alphas[i] + let partial: f32 = f[i] - kii * a_old + let g: f32 = y[i] - b - partial + + let mut a_new: f32 = g / kii + if a_new > alpha_max { a_new = alpha_max } + if a_new < 0.0 - alpha_max { a_new = 0.0 - alpha_max } + + let delta: f32 = a_new - a_old + let adelta: f32 = fabs((delta) as f64) as f32 + if adelta > 0.0 { + alphas[i] = a_new + for k in 0 to n { + f[k] = f[k] + row[k] * delta } + if adelta > max_change { max_change = adelta } } } + if max_change < tol { break } } + array_free_f32(f) + matrix_free(K) let mut n_support: i32 = 0 for i in 0 to n { @@ -1974,39 +2029,45 @@ export function one_class_svm_fit(X: Matrix, nu: f32, gamma: f32, max_iter: i32) let mut b: f32 = 0.0 + # The kernel is fixed, so it is built once rather than rebuilt from the + # raw features inside the innermost loop. That loop recomputed a squared + # distance across every feature and took an exponential for every pair of + # samples, on every one of max_iter passes, and the threshold pass below + # did the whole thing again. The update rule and the step are unchanged, + # so the result is the same. + let K: Matrix = _svm_kernel_matrix(X, 0, gamma, 0, 0.0) + let kdata: ptr = K.data + for iter in 0 to max_iter { + let mut changed: bool = false for i in 0 to n { - let mut g_i: f32 = -b + let row: ptr = kdata + i * n + let mut g_i: f32 = 0.0 - b for j in 0 to n { - let mut s: f32 = 0.0 - for k in 0 to p { - let d: f32 = matrix_at(X, i, k) - matrix_at(X, j, k) - s = s + d * d - } - g_i = g_i + alphas[j] * exp((-gamma * s) as f64) as f32 + g_i = g_i + alphas[j] * row[j] } if g_i > 0 && alphas[i] < alpha_max { alphas[i] = alphas[i] + 0.001 + changed = true } else { if g_i < 0 && alphas[i] > 0 { alphas[i] = alphas[i] - 0.001 if alphas[i] < 0 { alphas[i] = 0.0 } + changed = true } } } + # Every coefficient is clamped, so another pass moves nothing. + if not changed { break } } let mut threshold: f32 = 0.0 for i in 0 to n { - let mut g_i: f32 = -b + let row2: ptr = kdata + i * n + let mut g_i: f32 = 0.0 - b for j in 0 to n { - let mut s: f32 = 0.0 - for k in 0 to p { - let d: f32 = matrix_at(X, i, k) - matrix_at(X, j, k) - s = s + d * d - } - g_i = g_i + alphas[j] * exp((-gamma * s) as f64) as f32 + g_i = g_i + alphas[j] * row2[j] } threshold = threshold + g_i } @@ -2017,6 +2078,8 @@ export function one_class_svm_fit(X: Matrix, nu: f32, gamma: f32, max_iter: i32) if alphas[i] > 0.000001 { n_support = n_support + 1 } } + matrix_free(K) + return OneClassSVM { # Freed by this model's free function, so it must be this model's # own memory. Aliasing the caller's made fit and free destroy the @@ -2554,37 +2617,58 @@ export function svr_fit(X: Matrix, y: ptr, C: f32, epsilon: f32, kernel: i3 let alpha_diff: ptr = array_new_f32(n) let mut b: f32 = 0.0 + # Exact coordinate descent on the epsilon-insensitive dual. + # + # This used to walk every coefficient by a fixed 0.01 whenever it saw a + # violation, recomputing the whole decision value from scratch for each + # sample on each of up to 1000 passes. That is O(max_iter * n^2) with no + # early exit, and the fixed step means it stops wherever the budget runs + # out rather than at the optimum. + # + # Minimising the objective in one coordinate has a closed form. With + # K_ii a - g + epsilon * sign(a) = 0 the solution is a soft threshold of + # the partial residual, clipped to the box. The decision values are kept + # up to date as each coefficient moves, so a pass costs what it did + # before while converging in a handful of passes instead of a thousand. + let kdata: ptr = K.data + let f: ptr = array_new_f32(n) let max_iter: i32 = 1000 - let step: f32 = 0.01 + let tol: f32 = 0.0001 for iter in 0 to max_iter { - let mut num_changed: i32 = 0 + let mut max_change: f32 = 0.0 for i in 0 to n { - let mut f_i: f32 = b - for j in 0 to n { - f_i = f_i + alpha_diff[j] * matrix_at(K, i, j) - } - - let g_i: f32 = y[i] - f_i - - if g_i > epsilon && alpha_diff[i] < C { - alpha_diff[i] = alpha_diff[i] + step - if alpha_diff[i] > C { - alpha_diff[i] = C - } - num_changed = num_changed + 1 - } elif g_i < -epsilon && alpha_diff[i] > -C { - alpha_diff[i] = alpha_diff[i] - step - if alpha_diff[i] < -C { - alpha_diff[i] = -C + let row: ptr = kdata + i * n + let kii: f32 = row[i] + if kii <= 0.0000001 { continue } + + let a_old: f32 = alpha_diff[i] + # The decision value this coefficient is not responsible for. + let partial: f32 = f[i] - kii * a_old + let g: f32 = y[i] - partial + + let mut a_new: f32 = 0.0 + if g > epsilon { + a_new = (g - epsilon) / kii + } elif g < 0.0 - epsilon { + a_new = (g + epsilon) / kii + } + if a_new > C { a_new = C } + if a_new < 0.0 - C { a_new = 0.0 - C } + + let delta: f32 = a_new - a_old + let adelta: f32 = fabs((delta) as f64) as f32 + if adelta > 0.0 { + alpha_diff[i] = a_new + for k in 0 to n { + f[k] = f[k] + row[k] * delta } - num_changed = num_changed + 1 + if adelta > max_change { max_change = adelta } } } - if num_changed == 0 { - break - } + if max_change < tol { break } } + array_free_f32(f) let mut b_sum: f32 = 0.0 let mut b_count: i32 = 0 From 74f02934d1e3d49a1c2a1ccc5c4d3a8b96cbfe76 Mon Sep 17 00:00:00 2001 From: godofecht Date: Sun, 27 Sep 2026 20:22:13 +0100 Subject: [PATCH 02/12] Fix seven more solvers, including two that were quadratic or worse elastic_net_fit was proximal gradient with a fixed step and a fixed budget of epochs, allocating its gradient buffer inside the epoch loop. At the small end of a regularization path that converges slowly, so every fit ran all thousand epochs however close it already was. It is cyclic coordinate descent now, which is what cd_fast does and which settles in tens of sweeps at any alpha. elastic_net_cv goes from 167.8 ms to 2.5 ms, 0.08x to 5.5x, and elastic_net itself from 0.39 ms to 0.028 ms. multitask_lasso and multitask_elastic_net allocated a gradient buffer per task on every epoch, read the design through matrix_at, and had no way to stop. 0.45 ms to 0.022 ms and 0.45 ms to 0.013 ms, 0.48x to 9.9x and 0.46x to 16x. multitask_lasso_cv follows them from 0.59x to 1.7x. affinity_propagation rescanned a whole row to find the largest A + S excluding one column, for every column, which made the responsibility pass cubic in the sample count. The row maximum and runner up answer every column at once. Both availability passes walked down a column of a row-pointer array and now run row by row. 139.3 ms to 6.4 ms, 0.03x to 0.64x. local_outlier_factor bubble sorted every row of the distance matrix, twice, once per sample, so it was cubic as well. Only the k nearest are needed, so it selects the first k + 1 positions instead. The distance matrix is symmetric and is now built as one triangle off row pointers. 5.4 ms to 0.73 ms, 0.07x to 0.54x. Both of those last two are still behind scikit-learn. They were 30x and 14x behind. --- lib/scikit/cluster.flow | 76 ++++++--- lib/scikit/linear.flow | 317 +++++++++++++++++++++++++++----------- lib/scikit/neighbors.flow | 85 ++++++---- 3 files changed, 339 insertions(+), 139 deletions(-) diff --git a/lib/scikit/cluster.flow b/lib/scikit/cluster.flow index d9bf174..d697eec 100644 --- a/lib/scikit/cluster.flow +++ b/lib/scikit/cluster.flow @@ -1957,41 +1957,80 @@ export function affinity_propagation_fit(X: Matrix, damping: f32, max_iter: i32, let mut old_labels: ptr = malloc((n as i64) * 4) as ptr for i in 0 to n { old_labels[i] = -1 } + # One label buffer for the whole sweep. It used to be allocated and freed + # on every iteration. + let labels: ptr = malloc((n as i64) * 4) as ptr + let col_pos: ptr = array_new_f32(n) + let rkk: ptr = array_new_f32(n) + for iter in 0 to max_iter { n_iter = iter + 1 + # The largest A + S in row i, excluding column k, is the row maximum + # unless k is where that maximum sits, in which case it is the runner + # up. Finding both once per row makes this pass quadratic. Rescanning + # the row for every k made it cubic, which is the whole cost of a fit + # at these sizes. for i in 0 to n { - for k in 0 to n { - let mut max_other: f32 = -9999999999.0 - for kk in 0 to n { - if kk == k { continue } - let val: f32 = A[i][kk] + S[i][kk] - if val > max_other { max_other = val } + let ai: ptr = A[i] + let si: ptr = S[i] + let ri: ptr = R[i] + + let mut best: f32 = -9999999999.0 + let mut second: f32 = -9999999999.0 + let mut best_k: i32 = 0 + for kk in 0 to n { + let val: f32 = ai[kk] + si[kk] + if val > best { + second = best + best = val + best_k = kk + } elif val > second { + second = val } - R[i][k] = (1.0 - damping) * (S[i][k] - max_other) + damping * R[i][k] + } + + for k in 0 to n { + let mut max_other: f32 = best + if k == best_k { max_other = second } + ri[k] = (1.0 - damping) * (si[k] - max_other) + damping * ri[k] } } + # Both passes below used to walk down a column of a row-pointer array, + # touching a different cache line for every element. They accumulate + # and write row by row instead, which is the order the rows are + # actually laid out in. for k in 0 to n { - let mut sum_pos: f32 = 0.0 - for i in 0 to n { - if i == k { continue } - if R[i][k] > 0.0 { sum_pos = sum_pos + R[i][k] } + col_pos[k] = 0.0 + rkk[k] = R[k][k] + } + for i in 0 to n { + let ri2: ptr = R[i] + for k in 0 to n { + if i != k { + let v: f32 = ri2[k] + if v > 0.0 { col_pos[k] = col_pos[k] + v } + } } - for i in 0 to n { + } + for i in 0 to n { + let ai2: ptr = A[i] + let ri3: ptr = R[i] + for k in 0 to n { + let base_val: f32 = col_pos[k] + rkk[k] if i == k { - A[i][k] = (1.0 - damping) * (sum_pos + R[k][k]) + damping * A[i][k] + ai2[k] = (1.0 - damping) * base_val + damping * ai2[k] } else { - let r_ik: f32 = R[i][k] - let val: f32 = sum_pos + R[k][k] + let r_ik: f32 = ri3[k] + let mut val: f32 = base_val if r_ik > 0.0 { val = val - r_ik } if val > 0.0 { val = 0.0 } - A[i][k] = (1.0 - damping) * val + damping * A[i][k] + ai2[k] = (1.0 - damping) * val + damping * ai2[k] } } } - let labels: ptr = malloc((n as i64) * 4) as ptr for i in 0 to n { let mut best_k: i32 = 0 let mut best_val: f32 = -9999999999.0 @@ -2010,8 +2049,6 @@ export function affinity_propagation_fit(X: Matrix, damping: f32, max_iter: i32, if labels[i] != old_labels[i] { changed = changed + 1 } old_labels[i] = labels[i] } - free(labels as ptr) - if changed == 0 { conv_count = conv_count + 1 if conv_count >= convergence_iter { break } @@ -2020,7 +2057,6 @@ export function affinity_propagation_fit(X: Matrix, damping: f32, max_iter: i32, } } - let labels: ptr = malloc((n as i64) * 4) as ptr let exemplars: ptr = malloc((n as i64) * 4) as ptr let mut n_exemplars: i32 = 0 diff --git a/lib/scikit/linear.flow b/lib/scikit/linear.flow index 85a6fbe..08272fa 100644 --- a/lib/scikit/linear.flow +++ b/lib/scikit/linear.flow @@ -3526,46 +3526,114 @@ export struct ElasticNet { fitted: bool } +# Cyclic coordinate descent, which is what scikit-learn's cd_fast does. +# +# This was proximal gradient with a fixed learning rate and a fixed budget of +# epochs. At the small end of a regularization path that converges slowly, so +# every fit ran all thousand epochs however close it already was: the +# cross-validated wrapper spent 61 ms sweeping ten alphas over five folds. +# Coordinate descent has a closed form per coordinate and settles in tens of +# sweeps at any alpha on the path. +# +# The objective is unchanged: (1/2n)||y - Xw||^2 + l1 |w|_1 + (l2/2)||w||^2, +# so l1 and l2 keep their meaning and the fitted coefficients agree with what +# the gradient loop was converging toward. export function elastic_net_fit(X: Matrix, y: ptr, alpha: f32, l1_ratio: f32, epochs: i32, lr: f32) -> ElasticNet { let n: i32 = X.cols + let m: i32 = X.rows let weights: ptr = array_new_f32(n) - let bias: f32 = 0.0 + let mut bias: f32 = 0.0 let l1: f32 = alpha * l1_ratio let l2: f32 = alpha * (1.0 - l1_ratio) + if m <= 0 || n <= 0 { + return ElasticNet { weights: weights, bias: 0.0, n_features: n, alpha: alpha, l1_ratio: l1_ratio, fitted: true } + } + + let xdata: ptr = X.data + let inv_m: f64 = 1.0 / (m as f64) + + # Centre both sides so the intercept stays out of the penalty. + let x_mean: ptr = array_new_f64(n) + let mut y_mean: f64 = 0.0 + for i in 0 to m { + y_mean = y_mean + (y[i] as f64) + let base: i32 = i * n + for j in 0 to n { x_mean[j] = x_mean[j] + (xdata[base + j] as f64) } + } + y_mean = y_mean * inv_m + for j in 0 to n { x_mean[j] = x_mean[j] * inv_m } + + let Xc: ptr = array_new_f64(m * n) + let resid: ptr = array_new_f64(m) + let col_sq: ptr = array_new_f64(n) + for i in 0 to m { + let base: i32 = i * n + resid[i] = (y[i] as f64) - y_mean + for j in 0 to n { + let v: f64 = (xdata[base + j] as f64) - x_mean[j] + Xc[base + j] = v + col_sq[j] = col_sq[j] + v * v + } + } + + let w: ptr = array_new_f64(n) + let l1_d: f64 = l1 as f64 + let l2_d: f64 = l2 as f64 + let tol: f64 = 0.0001 for epoch in 0 to epochs { - let grad_w: ptr = array_new_f32(n) - let mut grad_b: f32 = 0.0 + let mut max_move: f64 = 0.0 + let mut w_max: f64 = 0.0 + for j in 0 to n { + let cj: f64 = col_sq[j] * inv_m + if cj <= 0.0 { continue } + let w_old: f64 = w[j] - for i in 0 to X.rows { - let mut pred: f32 = bias - for j in 0 to n { - pred = pred + weights[j] * matrix_at(X, i, j) + # Correlation of this column with the residual, with this + # coordinate's own contribution added back in. + let mut rho: f64 = 0.0 + for i in 0 to m { + rho = rho + Xc[i * n + j] * resid[i] } - let err: f32 = pred - y[i] - for j in 0 to n { - grad_w[j] = grad_w[j] + err * matrix_at(X, i, j) + rho = rho * inv_m + cj * w_old + + let mut w_new: f64 = 0.0 + if rho > l1_d { + w_new = (rho - l1_d) / (cj + l2_d) + } elif rho < 0.0 - l1_d { + w_new = (rho + l1_d) / (cj + l2_d) } - grad_b = grad_b + err - } - let inv_n: f32 = 1.0 / (X.rows as f32) - for j in 0 to n { - grad_w[j] = grad_w[j] * inv_n + l2 * weights[j] - weights[j] = weights[j] - lr * grad_w[j] - if weights[j] > l1 { - weights[j] = weights[j] - l1 - } elif weights[j] < -l1 { - weights[j] = weights[j] + l1 - } else { - weights[j] = 0.0 + let d: f64 = w_new - w_old + if d != 0.0 { + w[j] = w_new + for i in 0 to m { + resid[i] = resid[i] - d * Xc[i * n + j] + } + let ad: f64 = fabs(d) + if ad > max_move { max_move = ad } } + let aw: f64 = fabs(w_new) + if aw > w_max { w_max = aw } } - bias = bias - lr * grad_b * inv_n - array_free_f32(grad_w) + if w_max == 0.0 { break } + if max_move / w_max < tol { break } } + let mut bias_d: f64 = y_mean + for j in 0 to n { + weights[j] = w[j] as f32 + bias_d = bias_d - x_mean[j] * w[j] + } + bias = bias_d as f32 + + array_free_f64(x_mean) + array_free_f64(Xc) + array_free_f64(resid) + array_free_f64(col_sq) + array_free_f64(w) + return ElasticNet { weights: weights, bias: bias, @@ -4872,22 +4940,39 @@ export function multitask_lasso_fit(X: Matrix, Y: Matrix, alpha: f32, epochs: i3 weights[c] = array_new_f32(p) } + # The gradient buffers are allocated once. They used to be allocated and + # freed on every epoch, which is epochs * (t + 1) allocator round trips + # for buffers whose size never changes, and the design was read through + # matrix_at, which copies the Matrix struct for every element. + let xdata: ptr = X.data + let ydata: ptr = Y.data + let grads: ptr > = malloc((t as i64) * 8) as ptr > + let prev: ptr > = malloc((t as i64) * 8) as ptr > + let grad_b: ptr = array_new_f32(t) + for c in 0 to t { + grads[c] = array_new_f32(p) + prev[c] = array_new_f32(p) + } + for epoch in 0 to epochs { - let grads: ptr > = malloc((t as i64) * 8) as ptr > - let grad_b: ptr = array_new_f32(t) for c in 0 to t { - grads[c] = array_new_f32(p) + grad_b[c] = 0.0 + for j in 0 to p { grads[c][j] = 0.0 } } for i in 0 to n { + let xrow: ptr = xdata + i * p + let yrow: ptr = ydata + i * t for c in 0 to t { + let wc: ptr = weights[c] + let gc: ptr = grads[c] let mut pred: f32 = bias[c] for j in 0 to p { - pred = pred + weights[c][j] * matrix_at(X, i, j) + pred = pred + wc[j] * xrow[j] } - let err: f32 = pred - matrix_at(Y, i, c) + let err: f32 = pred - yrow[c] for j in 0 to p { - grads[c][j] = grads[c][j] + err * matrix_at(X, i, j) + gc[j] = gc[j] + err * xrow[j] } grad_b[c] = grad_b[c] + err } @@ -4925,12 +5010,33 @@ export function multitask_lasso_fit(X: Matrix, Y: Matrix, alpha: f32, epochs: i3 } } + # Largest group move against the largest group norm, so the sweep + # ends when the iterate has settled rather than always running its + # whole epoch budget. + let mut max_move: f32 = 0.0 + let mut w_max: f32 = 0.0 for c in 0 to t { - array_free_f32(grads[c]) + let wc2: ptr = weights[c] + let pc: ptr = prev[c] + for j in 0 to p { + let mv: f32 = fabs((wc2[j] - pc[j]) as f64) as f32 + if mv > max_move { max_move = mv } + let aw: f32 = fabs((wc2[j]) as f64) as f32 + if aw > w_max { w_max = aw } + pc[j] = wc2[j] + } } - free(grads as ptr) - array_free_f32(grad_b) + if w_max == 0.0 { break } + if max_move / w_max < 0.0001 { break } + } + + for c in 0 to t { + array_free_f32(grads[c]) + array_free_f32(prev[c]) } + free(grads as ptr) + free(prev as ptr) + array_free_f32(grad_b) return MultiTaskLasso { weights: weights, @@ -4993,22 +5099,39 @@ export function multitask_elastic_net_fit(X: Matrix, Y: Matrix, alpha: f32, l1_r weights[c] = array_new_f32(p) } + # The gradient buffers are allocated once. They used to be allocated and + # freed on every epoch, which is epochs * (t + 1) allocator round trips + # for buffers whose size never changes, and the design was read through + # matrix_at, which copies the Matrix struct for every element. + let xdata: ptr = X.data + let ydata: ptr = Y.data + let grads: ptr > = malloc((t as i64) * 8) as ptr > + let prev: ptr > = malloc((t as i64) * 8) as ptr > + let grad_b: ptr = array_new_f32(t) + for c in 0 to t { + grads[c] = array_new_f32(p) + prev[c] = array_new_f32(p) + } + for epoch in 0 to epochs { - let grads: ptr > = malloc((t as i64) * 8) as ptr > - let grad_b: ptr = array_new_f32(t) for c in 0 to t { - grads[c] = array_new_f32(p) + grad_b[c] = 0.0 + for j in 0 to p { grads[c][j] = 0.0 } } for i in 0 to n { + let xrow: ptr = xdata + i * p + let yrow: ptr = ydata + i * t for c in 0 to t { + let wc: ptr = weights[c] + let gc: ptr = grads[c] let mut pred: f32 = bias[c] for j in 0 to p { - pred = pred + weights[c][j] * matrix_at(X, i, j) + pred = pred + wc[j] * xrow[j] } - let err: f32 = pred - matrix_at(Y, i, c) + let err: f32 = pred - yrow[c] for j in 0 to p { - grads[c][j] = grads[c][j] + err * matrix_at(X, i, j) + gc[j] = gc[j] + err * xrow[j] } grad_b[c] = grad_b[c] + err } @@ -5044,13 +5167,34 @@ export function multitask_elastic_net_fit(X: Matrix, Y: Matrix, alpha: f32, l1_r } } + # Largest group move against the largest group norm, so the sweep + # ends when the iterate has settled rather than always running its + # whole epoch budget. + let mut max_move: f32 = 0.0 + let mut w_max: f32 = 0.0 for c in 0 to t { - array_free_f32(grads[c]) + let wc2: ptr = weights[c] + let pc: ptr = prev[c] + for j in 0 to p { + let mv: f32 = fabs((wc2[j] - pc[j]) as f64) as f32 + if mv > max_move { max_move = mv } + let aw: f32 = fabs((wc2[j]) as f64) as f32 + if aw > w_max { w_max = aw } + pc[j] = wc2[j] + } } - free(grads as ptr) - array_free_f32(grad_b) + if w_max == 0.0 { break } + if max_move / w_max < 0.0001 { break } } + for c in 0 to t { + array_free_f32(grads[c]) + array_free_f32(prev[c]) + } + free(grads as ptr) + free(prev as ptr) + array_free_f32(grad_b) + return MultiTaskElasticNet { weights: weights, bias: bias, @@ -5420,82 +5564,75 @@ export struct ElasticNetCV { fitted: bool } +# Fold loop outside the alpha loop, as in ridge_cv_fit, so each training +# matrix is built once for the whole grid rather than once per alpha, and the +# copy runs off row pointers rather than through matrix_at. export function elastic_net_cv_fit(X: Matrix, y: ptr, n_alphas: i32) -> ElasticNetCV { let n: i32 = X.rows let n_features: i32 = X.cols + let xdata: ptr = X.data let alphas: ptr = array_new_f32(n_alphas) - let mut i: i32 = 0 - while i < n_alphas { + for i in 0 to n_alphas { let t: f32 = (i as f32) / ((n_alphas - 1) as f32) alphas[i] = 0.001 * pow(10.0 as f64, (t * 4.0) as f64) as f32 - i = i + 1 } let l1_ratio: f32 = 0.5 - let mut best_alpha: f32 = alphas[0] - let mut best_mse: f32 = 999999.0 + let sse: ptr = array_new_f32(n_alphas) let fold_size: i32 = n / 5 - let mut a_idx: i32 = 0 - while a_idx < n_alphas { - let alpha: f32 = alphas[a_idx] - let mut total_mse: f32 = 0.0 - let mut fold: i32 = 0 - while fold < 5 { - let test_start: i32 = fold * fold_size - let test_end: i32 = test_start + fold_size - if fold == 4 { test_end = n } - let n_train: i32 = n - (test_end - test_start) + for fold in 0 to 5 { + let test_start: i32 = fold * fold_size + let mut test_end: i32 = test_start + fold_size + if fold == 4 { test_end = n } + let n_test: i32 = test_end - test_start + let n_train: i32 = n - n_test - let X_train: Matrix = matrix_new(n_train, n_features) - let y_train: ptr = array_new_f32(n_train) - let mut idx: i32 = 0 - let mut j: i32 = 0 - while j < n { - if j < test_start || j >= test_end { - let mut k: i32 = 0 - while k < n_features { - matrix_set(X_train, idx, k, matrix_at(X, j, k)) - k = k + 1 - } - y_train[idx] = y[j] - idx = idx + 1 - } - j = j + 1 + let X_train: Matrix = matrix_new(n_train, n_features) + let y_train: ptr = array_new_f32(n_train) + let train_data: ptr = X_train.data + let mut idx: i32 = 0 + for j in 0 to n { + if j < test_start || j >= test_end { + let src: ptr = xdata + j * n_features + let dst: ptr = train_data + idx * n_features + for k in 0 to n_features { dst[k] = src[k] } + y_train[idx] = y[j] + idx = idx + 1 } + } - let model: ElasticNet = elastic_net_fit(X_train, y_train, alpha, l1_ratio, 1000, 0.01) - + for a_idx in 0 to n_alphas { + let model: ElasticNet = elastic_net_fit(X_train, y_train, alphas[a_idx], l1_ratio, 1000, 0.01) let mut fold_mse: f32 = 0.0 - let n_test: i32 = test_end - test_start - let mut t2: i32 = test_start - while t2 < test_end { + for t2 in test_start to test_end { + let trow: ptr = xdata + t2 * n_features let mut pred: f32 = model.bias - let mut k2: i32 = 0 - while k2 < n_features { - pred = pred + model.weights[k2] * matrix_at(X, t2, k2) - k2 = k2 + 1 + for k2 in 0 to n_features { + pred = pred + model.weights[k2] * trow[k2] } let diff: f32 = y[t2] - pred fold_mse = fold_mse + diff * diff - t2 = t2 + 1 } - total_mse = total_mse + fold_mse / (n_test as f32) - + sse[a_idx] = sse[a_idx] + fold_mse / (n_test as f32) elastic_net_free(model) - matrix_free(X_train) - array_free_f32(y_train) - fold = fold + 1 } - let avg_mse: f32 = total_mse / 5.0 + matrix_free(X_train) + array_free_f32(y_train) + } + + let mut best_alpha: f32 = alphas[0] + let mut best_mse: f32 = 999999.0 + for a_idx in 0 to n_alphas { + let avg_mse: f32 = sse[a_idx] / 5.0 if avg_mse < best_mse { best_mse = avg_mse - best_alpha = alpha + best_alpha = alphas[a_idx] } - a_idx = a_idx + 1 } + array_free_f32(sse) let best_model: ElasticNet = elastic_net_fit(X, y, best_alpha, l1_ratio, 1000, 0.01) let coef_copy: ptr = array_copy_f32(best_model.weights, n_features) diff --git a/lib/scikit/neighbors.flow b/lib/scikit/neighbors.flow index 2743acb..bc35f5c 100644 --- a/lib/scikit/neighbors.flow +++ b/lib/scikit/neighbors.flow @@ -582,31 +582,51 @@ export function local_outlier_factor_fit(X: Matrix, n_neighbors: i32) -> LocalOu let k: i32 = n_neighbors if k > n - 1 { k = n - 1 } + # The distance matrix is symmetric, so only the upper triangle is computed + # and the lower half is mirrored. Elements come off row pointers rather + # than through matrix_at, which copies the Matrix struct per element and + # was called twice for every feature of every pair. + let xdata: ptr = X.data let dists: ptr > = malloc((n as i64) * 8) as ptr > for i in 0 to n { dists[i] = array_new_f32(n) - for j in 0 to n { - let mut s: f32 = 0.0 + } + for i in 0 to n { + let ri: ptr = xdata + i * p + let di: ptr = dists[i] + di[i] = 0.0 + for j in i + 1 to n { + let rj: ptr = xdata + j * p + let mut acc: f32 = 0.0 for m in 0 to p { - let d: f32 = matrix_at(X, i, m) - matrix_at(X, j, m) - s = s + d * d + let d: f32 = ri[m] - rj[m] + acc = acc + d * d } - dists[i][j] = sqrt((s as f64)) as f32 + let dist: f32 = sqrt((acc as f64)) as f32 + di[j] = dist + dists[j][i] = dist } } let k_dist: ptr = array_new_f32(n) let lrd: ptr = array_new_f32(n) + # Only the k nearest matter, so the first k + 1 positions are selected + # rather than the whole row being ordered. This was a bubble sort over n + # elements for each of n samples, which is cubic in the sample count and + # was the whole cost of a fit. + let sorted: ptr = array_new_f32(n) for i in 0 to n { - let sorted: ptr = array_copy_f32(dists[i], n) - for a in 0 to n { - for b in 0 to n - 1 - a { - if sorted[b] > sorted[b + 1] { - let t: f32 = sorted[b] - sorted[b] = sorted[b + 1] - sorted[b + 1] = t - } + for j in 0 to n { sorted[j] = dists[i][j] } + for a in 0 to k + 1 { + let mut m_idx: i32 = a + for b in a + 1 to n { + if sorted[b] < sorted[m_idx] { m_idx = b } + } + if m_idx != a { + let t: f32 = sorted[a] + sorted[a] = sorted[m_idx] + sorted[m_idx] = t } } k_dist[i] = sorted[k] @@ -619,34 +639,41 @@ export function local_outlier_factor_fit(X: Matrix, n_neighbors: i32) -> LocalOu reach_sum = reach_sum + reach } lrd[i] = (k as f32) / (reach_sum + 0.0000000001) - array_free_f32(sorted) } + array_free_f32(sorted) let lof: ptr = array_new_f32(n) + let sorted2: ptr = array_new_f32(n) + let indices: ptr = malloc((n as i64) * 4) as ptr for i in 0 to n { let mut sum_lrd: f32 = 0.0 - let sorted: ptr = array_copy_f32(dists[i], n) - let indices: ptr = malloc((n as i64) * 4) as ptr - for j in 0 to n { indices[j] = j } - for a in 0 to n { - for b in 0 to n - 1 - a { - if sorted[b] > sorted[b + 1] { - let t: f32 = sorted[b] - sorted[b] = sorted[b + 1] - sorted[b + 1] = t - let ti: i32 = indices[b] - indices[b] = indices[b + 1] - indices[b + 1] = ti - } + for j in 0 to n { + sorted2[j] = dists[i][j] + indices[j] = j + } + # Same partial selection, carrying the indices so the k neighbours can + # be named afterwards. + for a in 0 to k + 1 { + let mut m_idx: i32 = a + for b in a + 1 to n { + if sorted2[b] < sorted2[m_idx] { m_idx = b } + } + if m_idx != a { + let t: f32 = sorted2[a] + sorted2[a] = sorted2[m_idx] + sorted2[m_idx] = t + let ti: i32 = indices[a] + indices[a] = indices[m_idx] + indices[m_idx] = ti } } for j in 1 to k + 1 { sum_lrd = sum_lrd + lrd[indices[j]] } lof[i] = sum_lrd / (k as f32) / (lrd[i] + 0.0000000001) - array_free_f32(sorted) - free(indices as ptr) } + array_free_f32(sorted2) + free(indices as ptr) let mut mean_lof: f32 = 0.0 for i in 0 to n { mean_lof = mean_lof + lof[i] } From df35c5a64d9962f61fde3721bfcc659e5eded0ee Mon Sep 17 00:00:00 2001 From: godofecht Date: Sun, 27 Sep 2026 20:34:27 +0100 Subject: [PATCH 03/12] Take hdbscan off its cubic path, and two more distance matrices off matrix_at hdbscan built the same Euclidean distances three ways over. The core-distance pass rebuilt every row from the raw features and then ran a full selection sort over it to read one element, which is cubic in the sample count, and Prim's loop rebuilt each distance again for every pair it considered. One symmetric distance matrix now serves all of it and the core pass selects only the positions it needs. 4.5 ms to 0.41 ms, 0.16x to 1.75x. spectral_embedding builds its affinity matrix as one triangle off row pointers, and its power iteration takes the row pointer once and reuses one matvec buffer rather than allocating on every step. None of that moved the number: the cost is 200 power iterations that do not converge, and beating ARPACK there needs a different eigensolver. It stays at 0.30x. multitask_elastic_net_cv no longer predicts over the whole design for every fold and every alpha when it scores five test folds, and builds each training pair once for the grid. That did not move it either. The cost is the fold fits themselves, which are still proximal gradient and still run their whole epoch budget at the small end of the path. Block coordinate descent is the fix and is not in this commit. It stays at 0.28x. --- lib/scikit/cluster.flow | 87 ++++++++++++++++---------------- lib/scikit/linear.flow | 105 +++++++++++++++++++-------------------- lib/scikit/manifold.flow | 40 +++++++++++---- 3 files changed, 126 insertions(+), 106 deletions(-) diff --git a/lib/scikit/cluster.flow b/lib/scikit/cluster.flow index d697eec..229c723 100644 --- a/lib/scikit/cluster.flow +++ b/lib/scikit/cluster.flow @@ -2407,48 +2407,54 @@ export function hdbscan_fit(X: Matrix, min_cluster_size: i32, min_samples: i32) let labels: ptr = array_new_f32(n) # Compute core distances (distance to min_samples-th nearest neighbor) + # One distance matrix for the whole fit. The core-distance pass rebuilt the + # distances from the raw features for every point, and Prim's loop below + # rebuilt each one again for every pair it considered, both through + # matrix_at, which copies the Matrix struct per element. The matrix is + # symmetric, so only one triangle is computed. + let xdata: ptr = X.data + let D: ptr = array_new_f32(n * n) + for a0 in 0 to n { + let ra: ptr = xdata + a0 * n_features + D[a0 * n + a0] = 0.0 + for b0 in a0 + 1 to n { + let rb: ptr = xdata + b0 * n_features + let mut acc: f32 = 0.0 + for k0 in 0 to n_features { + let d0: f32 = ra[k0] - rb[k0] + acc = acc + d0 * d0 + } + let dv: f32 = sqrt((acc) as f64) as f32 + D[a0 * n + b0] = dv + D[b0 * n + a0] = dv + } + } + + # Only the min_samples nearest matter, so the row is partially selected + # rather than fully ordered. A whole selection sort per point made this + # cubic in the sample count. let core_dist: ptr = array_new_f32(n) + let dists: ptr = array_new_f32(n) + let mut need: i32 = min_samples + if need >= n { need = n - 1 } let mut i: i32 = 0 while i < n { - # Compute distances from point i to all others - let dists: ptr = array_new_f32(n) - let mut j: i32 = 0 - while j < n { - let mut s: f32 = 0.0 - let mut k: i32 = 0 - while k < n_features { - let d: f32 = matrix_at(X, i, k) - matrix_at(X, j, k) - s = s + d * d - k = k + 1 - } - dists[j] = sqrt((s) as f64) as f32 - j = j + 1 - } - # Sort to find min_samples-th nearest - let mut a: i32 = 0 - while a < n { + for j0 in 0 to n { dists[j0] = D[i * n + j0] } + for a in 0 to need + 1 { let mut min_idx: i32 = a - let mut b: i32 = a + 1 - while b < n { - if dists[b] < dists[min_idx] { - min_idx = b - } - b = b + 1 + for b in a + 1 to n { + if dists[b] < dists[min_idx] { min_idx = b } + } + if min_idx != a { + let tmp: f32 = dists[a] + dists[a] = dists[min_idx] + dists[min_idx] = tmp } - let tmp: f32 = dists[a] - dists[a] = dists[min_idx] - dists[min_idx] = tmp - a = a + 1 - } - # dists[0] is self (0), dists[min_samples] is the core distance - if min_samples < n { - core_dist[i] = dists[min_samples] - } else { - core_dist[i] = dists[n - 1] } - array_free_f32(dists) + core_dist[i] = dists[need] i = i + 1 } + array_free_f32(dists) # Build mutual reachability distance matrix # d_mreach(a,b) = max(core_dist[a], core_dist[b], dist(a,b)) @@ -2498,16 +2504,7 @@ export function hdbscan_fit(X: Matrix, min_cluster_size: i32, min_samples: i32) let mut v: i32 = 0 while v < n { if in_tree[v] == 0 { - # Compute mutual reachability distance - let mut s: f32 = 0.0 - let mut k: i32 = 0 - while k < n_features { - let d: f32 = matrix_at(X, u, k) - matrix_at(X, v, k) - s = s + d * d - k = k + 1 - } - let dist_uv: f32 = sqrt((s) as f64) as f32 - let mut mr: f32 = dist_uv + let mut mr: f32 = D[u * n + v] if core_dist[u] > mr { mr = core_dist[u] } if core_dist[v] > mr { mr = core_dist[v] } if mr < min_edge[v] { @@ -2639,6 +2636,8 @@ export function hdbscan_fit(X: Matrix, min_cluster_size: i32, min_samples: i32) n_samples: n, fitted: true } + array_free_f32(D) + return result } diff --git a/lib/scikit/linear.flow b/lib/scikit/linear.flow index 08272fa..c670738 100644 --- a/lib/scikit/linear.flow +++ b/lib/scikit/linear.flow @@ -6961,86 +6961,85 @@ export struct MultiTaskElasticNetCV { fitted: bool } +# Fold loop outside the alpha loop, and the score comes from the test rows +# only. This used to call predict on the whole design for every fold and every +# alpha, throwing away four fifths of the result each time, and rebuild both +# training matrices once per alpha through matrix_at. export function multitask_elastic_net_cv_fit(X: Matrix, Y: Matrix, n_alphas: i32) -> MultiTaskElasticNetCV { let n: i32 = X.rows let n_features: i32 = X.cols let n_tasks: i32 = Y.cols let l1_ratio: f32 = 0.5 + let xdata: ptr = X.data + let ydata: ptr = Y.data let alphas: ptr = array_new_f32(n_alphas) - let mut i: i32 = 0 - while i < n_alphas { + for i in 0 to n_alphas { let t: f32 = (i as f32) / ((n_alphas - 1) as f32) alphas[i] = 0.001 * pow(10.0 as f64, (t * 4.0) as f64) as f32 - i = i + 1 } - let mut best_alpha: f32 = alphas[0] - let mut best_mse: f32 = 999999.0 + let sse: ptr = array_new_f32(n_alphas) let fold_size: i32 = n / 5 - let mut a_idx: i32 = 0 - while a_idx < n_alphas { - let alpha: f32 = alphas[a_idx] - let mut total_mse: f32 = 0.0 - let mut fold: i32 = 0 - while fold < 5 { - let test_start: i32 = fold * fold_size - let test_end: i32 = test_start + fold_size - if fold == 4 { test_end = n } - let n_train: i32 = n - (test_end - test_start) + for fold in 0 to 5 { + let test_start: i32 = fold * fold_size + let mut test_end: i32 = test_start + fold_size + if fold == 4 { test_end = n } + let n_test: i32 = test_end - test_start + let n_train: i32 = n - n_test - let X_train: Matrix = matrix_new(n_train, n_features) - let Y_train: Matrix = matrix_new(n_train, n_tasks) - let mut idx: i32 = 0 - let mut j: i32 = 0 - while j < n { - if j < test_start || j >= test_end { - let mut k: i32 = 0 - while k < n_features { - matrix_set(X_train, idx, k, matrix_at(X, j, k)) - k = k + 1 - } - let mut k2: i32 = 0 - while k2 < n_tasks { - matrix_set(Y_train, idx, k2, matrix_at(Y, j, k2)) - k2 = k2 + 1 - } - idx = idx + 1 - } - j = j + 1 + let X_train: Matrix = matrix_new(n_train, n_features) + let Y_train: Matrix = matrix_new(n_train, n_tasks) + let xt: ptr = X_train.data + let yt: ptr = Y_train.data + let mut idx: i32 = 0 + for j in 0 to n { + if j < test_start || j >= test_end { + let src: ptr = xdata + j * n_features + let dst: ptr = xt + idx * n_features + for k in 0 to n_features { dst[k] = src[k] } + let ysrc: ptr = ydata + j * n_tasks + let ydst: ptr = yt + idx * n_tasks + for k2 in 0 to n_tasks { ydst[k2] = ysrc[k2] } + idx = idx + 1 } + } - let model: MultiTaskElasticNet = multitask_elastic_net_fit(X_train, Y_train, alpha, l1_ratio, 500, 0.01) - let preds: Matrix = multitask_elastic_net_predict(model, X) - + for a_idx in 0 to n_alphas { + let model: MultiTaskElasticNet = multitask_elastic_net_fit(X_train, Y_train, alphas[a_idx], l1_ratio, 500, 0.01) let mut fold_mse: f32 = 0.0 - let mut t2: i32 = test_start - while t2 < test_end { - let mut k3: i32 = 0 - while k3 < n_tasks { - let diff: f32 = matrix_at(Y, t2, k3) - matrix_at(preds, t2, k3) + for t2 in test_start to test_end { + let xrow: ptr = xdata + t2 * n_features + let yrow: ptr = ydata + t2 * n_tasks + for k3 in 0 to n_tasks { + let wk: ptr = model.weights[k3] + let mut pred: f32 = model.bias[k3] + for m in 0 to n_features { + pred = pred + wk[m] * xrow[m] + } + let diff: f32 = yrow[k3] - pred fold_mse = fold_mse + diff * diff - k3 = k3 + 1 } - t2 = t2 + 1 } - total_mse = total_mse + fold_mse / ((test_end - test_start) as f32 * (n_tasks as f32)) - - matrix_free(preds) + sse[a_idx] = sse[a_idx] + fold_mse / ((n_test as f32) * (n_tasks as f32)) multitask_elastic_net_free(model) - matrix_free(X_train) - matrix_free(Y_train) - fold = fold + 1 } - let avg_mse: f32 = total_mse / 5.0 + matrix_free(X_train) + matrix_free(Y_train) + } + + let mut best_alpha: f32 = alphas[0] + let mut best_mse: f32 = 999999.0 + for a_idx in 0 to n_alphas { + let avg_mse: f32 = sse[a_idx] / 5.0 if avg_mse < best_mse { best_mse = avg_mse - best_alpha = alpha + best_alpha = alphas[a_idx] } - a_idx = a_idx + 1 } + array_free_f32(sse) let best_model: MultiTaskElasticNet = multitask_elastic_net_fit(X, Y, best_alpha, l1_ratio, 500, 0.01) let weights_copy: ptr > = malloc((n_tasks as i64) * 8) as ptr > diff --git a/lib/scikit/manifold.flow b/lib/scikit/manifold.flow index f24a1e9..d0defcb 100644 --- a/lib/scikit/manifold.flow +++ b/lib/scikit/manifold.flow @@ -613,16 +613,29 @@ export function spectral_embedding_fit(X: Matrix, n_components: i32, gamma: f32) let n: i32 = X.rows let p: i32 = X.cols + # The affinity matrix is symmetric, so one triangle is computed and + # mirrored, which halves both the squared distances and the exponentials. + # Elements come off row pointers rather than through matrix_at, which + # copies the Matrix struct for every element and was called twice per + # feature of every pair. + let xdata: ptr = X.data let W: ptr > = malloc((n as i64) * 8) as ptr > for i in 0 to n { W[i] = array_new_f32(n) - for j in 0 to n { - let mut s: f32 = 0.0 + } + for i in 0 to n { + let ri: ptr = xdata + i * p + W[i][i] = 1.0 + for j in i + 1 to n { + let rj: ptr = xdata + j * p + let mut acc: f32 = 0.0 for k in 0 to p { - let d: f32 = matrix_at(X, i, k) - matrix_at(X, j, k) - s = s + d * d + let d: f32 = ri[k] - rj[k] + acc = acc + d * d } - W[i][j] = exp((-gamma * s) as f64) as f32 + let wv: f32 = exp((0.0 - gamma * acc) as f64) as f32 + W[i][j] = wv + W[j][i] = wv } } @@ -643,6 +656,12 @@ export function spectral_embedding_fit(X: Matrix, n_components: i32, gamma: f32) let embedding: ptr = array_new_f32(n * n_components) let found_vecs: ptr > = malloc((n_components as i64) * 8) as ptr > + # One matvec buffer for every component and every iteration. It used to be + # allocated and freed inside the iteration loop, which is up to 200 + # allocator round trips per component for a vector whose size never + # changes. + let Lv: ptr = array_new_f32(n) + for c in 0 to n_components { let v: ptr = array_new_f32(n) srand(42 + c) @@ -657,10 +676,13 @@ export function spectral_embedding_fit(X: Matrix, n_components: i32, gamma: f32) let mut iter: i32 = 0 while iter < 200 { - let Lv: ptr = array_new_f32(n) for i in 0 to n { + # The row pointer is taken once. Indexing L[i][j] in the inner + # loop is two loads per element, one of them the same row + # pointer over and over. + let Li: ptr = L[i] let mut s: f32 = 0.0 - for j in 0 to n { s = s + L[i][j] * v[j] } + for j in 0 to n { s = s + Li[j] * v[j] } Lv[i] = s } @@ -683,8 +705,6 @@ export function spectral_embedding_fit(X: Matrix, n_components: i32, gamma: f32) if diff > max_change { max_change = diff } v[i] = new_v } - array_free_f32(Lv) - if max_change < 0.000001 { break } iter = iter + 1 } @@ -703,6 +723,8 @@ export function spectral_embedding_fit(X: Matrix, n_components: i32, gamma: f32) free(L as ptr) array_free_f32(deg) + array_free_f32(Lv) + return SpectralEmbedding { embedding: embedding, n_samples: n, From d3c5dd94a11df813d5e0f453df722a62fe451a27 Mon Sep 17 00:00:00 2001 From: godofecht Date: Sun, 27 Sep 2026 21:50:42 +0100 Subject: [PATCH 04/12] Replace the solvers behind the rows that still lost The wide matrix had eleven rows where scikit-learn was faster. Three of them were already fixed in the last two commits and needed no further work. These are the rest, and in each case the loss came from the algorithm rather than from the loop it was written in. MultiTaskElasticNet moves from a fixed step proximal gradient to block coordinate descent. Each feature block has a closed form minimizer given the others, so a sweep solves every block exactly, and the residual is carried through a rank one correction instead of being rebuilt. The cross validated version walks its alpha grid from the largest alpha down and warm starts each fit from the last, so the whole grid costs little more than its first point. 62.6 ms to 3.7 ms on the diabetes shape, measured on a loaded machine. MDS moves from gradient descent at a fixed hundredth of a step to the SMACOF Guttman transform, which is the exact minimizer of the majorizing function at the current configuration. Each pair is visited once and its contribution is added to both of its points. The old stopping rule compared raw stress against the tolerance and was never true at these scales, so every fit ran its whole iteration budget. SpectralEmbedding and LLE move from one scalar matvec per component per iteration to one sgemm per iteration over the whole block, with modified Gram Schmidt keeping the block orthonormal. LLE also builds its W^T W term with one sgemm where it used to walk a scalar dot per entry down a column of an array of row allocations, and picks its neighbours by partial selection instead of bubble sorting every row, which was cubic in the sample count on its own. KernelRidge builds the linear, polynomial and sigmoid kernels from one sgemm. The old path walked a scalar dot per pair, wrote every entry through matrix_set, and then read all of it back through matrix_at. A linear kernel also has a primal form, so predict folds the dual coefficients into one weight vector rather than evaluating an n by n_train kernel. KNeighborsTransformer drops a zero fill that matrix_new had already done through n by n_train calls to matrix_set, and reads its distances off row pointers. ClassifierChain does the same for the design it rebuilds per output in predict. LinearSVR predicts with one sgemv. --- lib/scikit/kernel_ridge.flow | 55 ++++- lib/scikit/linear.flow | 283 +++++++++++++++-------- lib/scikit/manifold.flow | 421 ++++++++++++++++++----------------- lib/scikit/multioutput.flow | 45 +++- lib/scikit/neighbors.flow | 69 +++--- lib/scikit/svm.flow | 15 +- 6 files changed, 517 insertions(+), 371 deletions(-) diff --git a/lib/scikit/kernel_ridge.flow b/lib/scikit/kernel_ridge.flow index c147d43..0255953 100644 --- a/lib/scikit/kernel_ridge.flow +++ b/lib/scikit/kernel_ridge.flow @@ -70,6 +70,15 @@ function tanh_f32(x: f32) -> f32 { return ((e1 - e2) / (e1 + e2)) as f32 } +# A linear, polynomial or sigmoid kernel entry is a function of the dot +# product alone, so the dot products can come from one sgemm and this turns +# each one into its kernel entry. +function _kr_from_dot(dot: f32, kernel_type: i32, gamma: f32, degree: i32, coef0: f32) -> f32 { + if kernel_type == KR_KERNEL_POLY { return pow_f32(gamma * dot + coef0, degree) } + if kernel_type == KR_KERNEL_SIGMOID { return tanh_f32(gamma * dot + coef0) } + return dot +} + function compute_kernel_matrix(X: Matrix, kernel_type: i32, gamma: f32, degree: i32, coef0: f32) -> Matrix { let n: i32 = X.rows let p: i32 = X.cols @@ -124,14 +133,24 @@ export function kernel_ridge_fit(X: Matrix, y: ptr, alpha: f32, kernel_type free(cross as ptr) array_free_f32(x_sq) } else { - let K: Matrix = compute_kernel_matrix(X, kernel_type, gamma, degree, coef0) + # The dot products behind these kernels are one sgemm. + # compute_kernel_matrix walked them with a scalar dot per pair and + # wrote every entry through matrix_set, which copies the Matrix struct + # per element, and the loop below then read all of it back through + # matrix_at. The matrix is symmetric, so one triangle is enough. + let x_data2: ptr = X.data + let cross2: ptr = malloc((n as i64) * (n as i64) * 4) as ptr + cblas_sgemm(101, 111, 112, n, n, n_features, 1.0, x_data2, n_features, x_data2, n_features, 0.0, cross2, n) for i in 0 to n { - for j in 0 to n { - K_reg[i * n + j] = (matrix_at(K, i, j)) as f64 + let crow: ptr = cross2 + i * n + for j in i to n { + let kv: f64 = (_kr_from_dot(crow[j], kernel_type, gamma, degree, coef0)) as f64 + K_reg[i * n + j] = kv + K_reg[j * n + i] = kv } K_reg[i * n + i] = K_reg[i * n + i] + (alpha as f64) } - matrix_free(K) + free(cross2 as ptr) } # Cholesky decomposition in f64 via LAPACK dpotrf_: L * L^T = K + alpha*I @@ -258,15 +277,27 @@ export function kernel_ridge_predict(model: KernelRidge, X: Matrix) -> ptr array_free_f32(x_sq) array_free_f32(xt_sq) } else { - for i in 0 to n { - let xi: ptr = X.data + i * model.n_features - let mut sum: f32 = 0.0 - for j in 0 to model.n_samples { - let xj: ptr = model.X_train.data + j * model.n_features - let k: f32 = kernel_compute(xi, xj, model.n_features, model.kernel_type, model.gamma, model.degree, model.coef0) - sum = sum + model.dual_coef[j] * k + let xt_data2: ptr = model.X_train.data + if model.kernel_type == KR_KERNEL_LINEAR { + # A linear kernel has a primal form. Folding the dual + # coefficients into one weight vector replaces an n by n_train + # kernel with a single matrix vector product. + let w: ptr = array_new_f32(nf) + blas_matvec_trans(xt_data2, nt, nf, model.dual_coef, w, 1.0, 0.0) + blas_matvec(X.data, n, nf, w, result, 1.0, 0.0) + array_free_f32(w) + } else { + let cross3: ptr = malloc((n as i64) * (nt as i64) * 4) as ptr + cblas_sgemm(101, 111, 112, n, nt, nf, 1.0, X.data, nf, xt_data2, nf, 0.0, cross3, nt) + for i in 0 to n { + let crow2: ptr = cross3 + i * nt + let mut sum: f32 = 0.0 + for j in 0 to nt { + sum = sum + model.dual_coef[j] * _kr_from_dot(crow2[j], model.kernel_type, model.gamma, model.degree, model.coef0) + } + result[i] = sum } - result[i] = sum + free(cross3 as ptr) } } diff --git a/lib/scikit/linear.flow b/lib/scikit/linear.flow index c670738..da65fc7 100644 --- a/lib/scikit/linear.flow +++ b/lib/scikit/linear.flow @@ -5085,115 +5085,171 @@ export struct MultiTaskElasticNet { fitted: bool } -export function multitask_elastic_net_fit(X: Matrix, Y: Matrix, alpha: f32, l1_ratio: f32, epochs: i32, lr: f32) -> MultiTaskElasticNet { +# Multi-task elastic net by block coordinate descent. +# +# Each feature block has a closed form minimizer given the others, so one +# sweep over the p blocks solves every block exactly. The old solver took +# full gradient steps of a fixed size and then applied a group soft +# threshold, which crawls toward the optimum and spends its whole epoch +# budget doing it. scikit-learn's enet_coordinate_descent_multi_task solves +# the same subproblem the same way. +# +# The design is held column major and centered once, because a block update +# reads one column and writes one rank one correction into the residual. +function _mtl_columns(X: Matrix, xcol: ptr, x_mean: ptr, xsq: ptr) -> void { let n: i32 = X.rows let p: i32 = X.cols - let t: i32 = Y.cols - - let l1: f32 = alpha * l1_ratio - let l2: f32 = alpha * (1.0 - l1_ratio) - - let weights: ptr > = malloc((t as i64) * 8) as ptr > - let bias: ptr = array_new_f32(t) - for c in 0 to t { - weights[c] = array_new_f32(p) + let xdata: ptr = X.data + for j in 0 to p { x_mean[j] = 0.0 } + for i in 0 to n { + let row: ptr = xdata + i * p + for j in 0 to p { x_mean[j] = x_mean[j] + row[j] } + } + let inv_n: f32 = 1.0 / (n as f32) + for j in 0 to p { x_mean[j] = x_mean[j] * inv_n } + for j in 0 to p { + let col: ptr = xcol + j * n + let mu: f32 = x_mean[j] + let mut s: f32 = 0.0 + for i in 0 to n { + let v: f32 = xdata[i * p + j] - mu + col[i] = v + s = s + v * v + } + xsq[j] = s } +} - # The gradient buffers are allocated once. They used to be allocated and - # freed on every epoch, which is epochs * (t + 1) allocator round trips - # for buffers whose size never changes, and the design was read through - # matrix_at, which copies the Matrix struct for every element. - let xdata: ptr = X.data +# Column means of Y, and the residual for a zero coefficient matrix. +function _mtl_center_y(Y: Matrix, y_mean: ptr, resid: ptr) -> void { + let n: i32 = Y.rows + let t: i32 = Y.cols let ydata: ptr = Y.data - let grads: ptr > = malloc((t as i64) * 8) as ptr > - let prev: ptr > = malloc((t as i64) * 8) as ptr > - let grad_b: ptr = array_new_f32(t) - for c in 0 to t { - grads[c] = array_new_f32(p) - prev[c] = array_new_f32(p) + for c in 0 to t { y_mean[c] = 0.0 } + for i in 0 to n { + let yrow: ptr = ydata + i * t + for c in 0 to t { y_mean[c] = y_mean[c] + yrow[c] } + } + let inv_n: f32 = 1.0 / (n as f32) + for c in 0 to t { y_mean[c] = y_mean[c] * inv_n } + for i in 0 to n { + let yrow2: ptr = ydata + i * t + let rrow: ptr = resid + i * t + for c in 0 to t { rrow[c] = yrow2[c] - y_mean[c] } } +} - for epoch in 0 to epochs { - for c in 0 to t { - grad_b[c] = 0.0 - for j in 0 to p { grads[c][j] = 0.0 } - } +# Sweeps until the largest coefficient move is small against the largest +# coefficient. W and resid are both in and out, so a caller walking an alpha +# path can warm start the next fit from the last one. +function _mtl_bcd(xcol: ptr, xsq: ptr, resid: ptr, W: ptr >, n: i32, p: i32, t: i32, l1_reg: f32, l2_reg: f32, max_sweeps: i32, tol: f32) -> void { + let tmp: ptr = array_new_f32(t) + let delta: ptr = array_new_f32(t) - for i in 0 to n { - let xrow: ptr = xdata + i * p - let yrow: ptr = ydata + i * t - for c in 0 to t { - let wc: ptr = weights[c] - let gc: ptr = grads[c] - let mut pred: f32 = bias[c] - for j in 0 to p { - pred = pred + wc[j] * xrow[j] - } - let err: f32 = pred - yrow[c] - for j in 0 to p { - gc[j] = gc[j] + err * xrow[j] - } - grad_b[c] = grad_b[c] + err + for sweep in 0 to max_sweeps { + let mut d_w_max: f32 = 0.0 + let mut w_max: f32 = 0.0 + + for j in 0 to p { + let xs: f32 = xsq[j] + if xs <= 0.0 { continue } + let col: ptr = xcol + j * n + + for c in 0 to t { tmp[c] = 0.0 } + for i in 0 to n { + let xij: f32 = col[i] + let ri: ptr = resid + i * t + for c in 0 to t { tmp[c] = tmp[c] + xij * ri[c] } } - } - let inv_n: f32 = 1.0 / (n as f32) - for c in 0 to t { - for j in 0 to p { - grads[c][j] = grads[c][j] * inv_n + l2 * weights[c][j] - weights[c][j] = weights[c][j] - lr * grads[c][j] + let mut nrm: f32 = 0.0 + for c in 0 to t { + let v: f32 = tmp[c] + xs * W[c][j] + tmp[c] = v + nrm = nrm + v * v } - bias[c] = bias[c] - lr * grad_b[c] * inv_n - } + nrm = sqrt((nrm) as f64) as f32 - # Group soft-thresholding with L1 portion - for j in 0 to p { - let mut norm: f32 = 0.0 + let mut scale: f32 = 0.0 + if nrm > l1_reg { scale = (1.0 - l1_reg / nrm) / (xs + l2_reg) } + + let mut moved: bool = false for c in 0 to t { - norm = norm + weights[c][j] * weights[c][j] + let old_w: f32 = W[c][j] + let new_w: f32 = tmp[c] * scale + W[c][j] = new_w + let d: f32 = old_w - new_w + delta[c] = d + if d != 0.0 { moved = true } + let mv: f32 = fabs((d) as f64) as f32 + if mv > d_w_max { d_w_max = mv } + let aw: f32 = fabs((new_w) as f64) as f32 + if aw > w_max { w_max = aw } } - norm = sqrt((norm) as f64) as f32 - if norm > 0.0 { - if norm <= l1 { - for c in 0 to t { - weights[c][j] = 0.0 - } - } else { - let scale: f32 = 1.0 - l1 / norm - for c in 0 to t { - weights[c][j] = weights[c][j] * scale - } + + # Rank one correction keeps the residual exact without a pass + # over the whole design. + if moved { + for i in 0 to n { + let xij2: f32 = col[i] + let ri2: ptr = resid + i * t + for c in 0 to t { ri2[c] = ri2[c] + xij2 * delta[c] } } } } - # Largest group move against the largest group norm, so the sweep - # ends when the iterate has settled rather than always running its - # whole epoch budget. - let mut max_move: f32 = 0.0 - let mut w_max: f32 = 0.0 - for c in 0 to t { - let wc2: ptr = weights[c] - let pc: ptr = prev[c] - for j in 0 to p { - let mv: f32 = fabs((wc2[j] - pc[j]) as f64) as f32 - if mv > max_move { max_move = mv } - let aw: f32 = fabs((wc2[j]) as f64) as f32 - if aw > w_max { w_max = aw } - pc[j] = wc2[j] - } - } if w_max == 0.0 { break } - if max_move / w_max < 0.0001 { break } + if d_w_max / w_max < tol { break } } + array_free_f32(tmp) + array_free_f32(delta) +} + +# Intercepts from the two sets of column means. +function _mtl_bias(W: ptr >, x_mean: ptr, y_mean: ptr, bias: ptr, p: i32, t: i32) -> void { for c in 0 to t { - array_free_f32(grads[c]) - array_free_f32(prev[c]) + let wc: ptr = W[c] + let mut b: f32 = y_mean[c] + for j in 0 to p { b = b - wc[j] * x_mean[j] } + bias[c] = b } - free(grads as ptr) - free(prev as ptr) - array_free_f32(grad_b) +} + +export function multitask_elastic_net_fit(X: Matrix, Y: Matrix, alpha: f32, l1_ratio: f32, epochs: i32, lr: f32) -> MultiTaskElasticNet { + let n: i32 = X.rows + let p: i32 = X.cols + let t: i32 = Y.cols + + # The penalties carry the sample count, because the loss below is + # halved and divided by n while the subproblem above is not. + let nf: f32 = n as f32 + let l1_reg: f32 = alpha * l1_ratio * nf + let l2_reg: f32 = alpha * (1.0 - l1_ratio) * nf + + let weights: ptr > = malloc((t as i64) * 8) as ptr > + for c in 0 to t { weights[c] = array_new_f32(p) } + let bias: ptr = array_new_f32(t) + + let xcol: ptr = array_new_f32(p * n) + let x_mean: ptr = array_new_f32(p) + let xsq: ptr = array_new_f32(p) + _mtl_columns(X, xcol, x_mean, xsq) + + let y_mean: ptr = array_new_f32(t) + let resid: ptr = array_new_f32(n * t) + _mtl_center_y(Y, y_mean, resid) + + let mut sweeps: i32 = epochs + if sweeps < 1 { sweeps = 1 } + _mtl_bcd(xcol, xsq, resid, weights, n, p, t, l1_reg, l2_reg, sweeps, 0.0001) + _mtl_bias(weights, x_mean, y_mean, bias, p, t) + + array_free_f32(xcol) + array_free_f32(x_mean) + array_free_f32(xsq) + array_free_f32(y_mean) + array_free_f32(resid) return MultiTaskElasticNet { weights: weights, @@ -6965,6 +7021,13 @@ export struct MultiTaskElasticNetCV { # only. This used to call predict on the whole design for every fold and every # alpha, throwing away four fifths of the result each time, and rebuild both # training matrices once per alpha through matrix_at. +# One centering per fold, and a warm start down the alpha path. +# +# The fit for a smaller alpha starts from the solution for the larger one, +# which is where sklearn's path solvers start too, and the residual that +# _mtl_bcd leaves behind is already the residual the next alpha needs. So a +# whole grid costs little more than its first point. The grid is walked from +# the largest alpha down, because that is the direction the warm start helps. export function multitask_elastic_net_cv_fit(X: Matrix, Y: Matrix, n_alphas: i32) -> MultiTaskElasticNetCV { let n: i32 = X.rows let n_features: i32 = X.cols @@ -6982,6 +7045,15 @@ export function multitask_elastic_net_cv_fit(X: Matrix, Y: Matrix, n_alphas: i32 let sse: ptr = array_new_f32(n_alphas) let fold_size: i32 = n / 5 + let W: ptr > = malloc((n_tasks as i64) * 8) as ptr > + for c in 0 to n_tasks { W[c] = array_new_f32(n_features) } + let bias: ptr = array_new_f32(n_tasks) + let x_mean: ptr = array_new_f32(n_features) + let xsq: ptr = array_new_f32(n_features) + let y_mean: ptr = array_new_f32(n_tasks) + let xcol: ptr = array_new_f32(n_features * n) + let resid: ptr = array_new_f32(n * n_tasks) + for fold in 0 to 5 { let test_start: i32 = fold * fold_size let mut test_end: i32 = test_start + fold_size @@ -7006,15 +7078,29 @@ export function multitask_elastic_net_cv_fit(X: Matrix, Y: Matrix, n_alphas: i32 } } - for a_idx in 0 to n_alphas { - let model: MultiTaskElasticNet = multitask_elastic_net_fit(X_train, Y_train, alphas[a_idx], l1_ratio, 500, 0.01) + _mtl_columns(X_train, xcol, x_mean, xsq) + _mtl_center_y(Y_train, y_mean, resid) + for c in 0 to n_tasks { + let wc0: ptr = W[c] + for j in 0 to n_features { wc0[j] = 0.0 } + } + + let nf: f32 = n_train as f32 + let mut a_idx: i32 = n_alphas - 1 + while a_idx >= 0 { + let alpha: f32 = alphas[a_idx] + let l1_reg: f32 = alpha * l1_ratio * nf + let l2_reg: f32 = alpha * (1.0 - l1_ratio) * nf + _mtl_bcd(xcol, xsq, resid, W, n_train, n_features, n_tasks, l1_reg, l2_reg, 500, 0.0001) + _mtl_bias(W, x_mean, y_mean, bias, n_features, n_tasks) + let mut fold_mse: f32 = 0.0 for t2 in test_start to test_end { let xrow: ptr = xdata + t2 * n_features let yrow: ptr = ydata + t2 * n_tasks for k3 in 0 to n_tasks { - let wk: ptr = model.weights[k3] - let mut pred: f32 = model.bias[k3] + let wk: ptr = W[k3] + let mut pred: f32 = bias[k3] for m in 0 to n_features { pred = pred + wk[m] * xrow[m] } @@ -7023,20 +7109,29 @@ export function multitask_elastic_net_cv_fit(X: Matrix, Y: Matrix, n_alphas: i32 } } sse[a_idx] = sse[a_idx] + fold_mse / ((n_test as f32) * (n_tasks as f32)) - multitask_elastic_net_free(model) + a_idx = a_idx - 1 } matrix_free(X_train) matrix_free(Y_train) } + array_free_f32(xcol) + array_free_f32(resid) + array_free_f32(x_mean) + array_free_f32(xsq) + array_free_f32(y_mean) + array_free_f32(bias) + for c in 0 to n_tasks { array_free_f32(W[c]) } + free(W as ptr) + let mut best_alpha: f32 = alphas[0] let mut best_mse: f32 = 999999.0 - for a_idx in 0 to n_alphas { - let avg_mse: f32 = sse[a_idx] / 5.0 + for a_idx2 in 0 to n_alphas { + let avg_mse: f32 = sse[a_idx2] / 5.0 if avg_mse < best_mse { best_mse = avg_mse - best_alpha = alphas[a_idx] + best_alpha = alphas[a_idx2] } } array_free_f32(sse) diff --git a/lib/scikit/manifold.flow b/lib/scikit/manifold.flow index d0defcb..d323c36 100644 --- a/lib/scikit/manifold.flow +++ b/lib/scikit/manifold.flow @@ -3,6 +3,7 @@ # Uses Barnes-Hut approximation for O(n log n) complexity. import "lib/scikit/matrix.flow" +import "lib/scikit/blas.flow" extern { function rand() -> i32 @@ -455,128 +456,126 @@ export struct LLE { fitted: bool } +# Modified Gram-Schmidt on the columns of a row-major n by k block. +# A column that collapses is left at zero, and the remaining columns stay +# orthonormal. +function _orthonormalize_block(V: ptr, n: i32, k: i32) -> void { + for c in 0 to k { + for prev in 0 to c { + let mut dot: f32 = 0.0 + for i in 0 to n { dot = dot + V[i * k + c] * V[i * k + prev] } + for i in 0 to n { V[i * k + c] = V[i * k + c] - dot * V[i * k + prev] } + } + let mut norm: f32 = 0.0 + for i in 0 to n { + let v: f32 = V[i * k + c] + norm = norm + v * v + } + norm = sqrt((norm) as f64) as f32 + if norm > 0.0000000001 { + let inv: f32 = 1.0 / norm + for i in 0 to n { V[i * k + c] = V[i * k + c] * inv } + } + } +} + export function lle_fit(X: Matrix, n_components: i32, n_neighbors: i32) -> LLE { let n: i32 = X.rows let p: i32 = X.cols + let kc: i32 = n_components - let W: ptr > = malloc((n as i64) * 8) as ptr > + # Squared distances off row pointers rather than through matrix_at, one + # triangle computed and mirrored. + let xdata: ptr = X.data + let dist: ptr = array_new_f32(n * n) for i in 0 to n { - W[i] = array_new_f32(n) - - let dists: ptr = array_new_f32(n) - for j in 0 to n { + let ri: ptr = xdata + i * p + let di: ptr = dist + i * n + for j in i + 1 to n { + let rj: ptr = xdata + j * p let mut s: f32 = 0.0 - for k in 0 to p { - let d: f32 = matrix_at(X, i, k) - matrix_at(X, j, k) + for f in 0 to p { + let d: f32 = ri[f] - rj[f] s = s + d * d } - dists[j] = s + di[j] = s + dist[j * n + i] = s } + } - let indices: ptr = malloc((n as i64) * 4) as ptr - for j in 0 to n { indices[j] = j } - for a in 0 to n { - for b in 0 to n - 1 - a { - if dists[indices[b]] > dists[indices[b + 1]] { - let t: i32 = indices[b] - indices[b] = indices[b + 1] - indices[b + 1] = t - } + # The neighbours of a row are the k + 1 smallest entries of its distance + # row, the first of which is the row itself. Partial selection stops after + # those, where the old loop bubble sorted the whole row, which is cubic in + # the sample count for a result that only needs a handful of positions. + let mut k: i32 = n_neighbors + if k > n - 1 { k = n - 1 } + let W: ptr = array_new_f32(n * n) + let idx: ptr = malloc((n as i64) * 4) as ptr + let wv: f32 = 1.0 / (k as f32) + for i in 0 to n { + let di2: ptr = dist + i * n + for j in 0 to n { idx[j] = j } + for a in 0 to k + 1 { + let mut best: i32 = a + for b in a + 1 to n { + if di2[idx[b]] < di2[idx[best]] { best = b } } + let t: i32 = idx[a] + idx[a] = idx[best] + idx[best] = t } - - let k: i32 = n_neighbors - if k > n - 1 { k = n - 1 } - let mut sum_w: f32 = 0.0 - for j in 1 to k + 1 { - W[i][indices[j]] = 1.0 / (k as f32) - sum_w = sum_w + W[i][indices[j]] - } - if sum_w > 0.0000000001 { - for j in 0 to n { W[i][j] = W[i][j] / sum_w } - } - - array_free_f32(dists) - free(indices as ptr) + let wi: ptr = W + i * n + for a2 in 1 to k + 1 { wi[idx[a2]] = wv } } + free(idx as ptr) + array_free_f32(dist) - let M: ptr > = malloc((n as i64) * 8) as ptr > + # M = (I - W)^T (I - W), which expands to I - W - W^T + W^T W. The cubic + # term is one sgemm. The old loop walked it with a scalar dot per entry and + # read one operand down a column of an array of row allocations, which is a + # cache miss per element. + let M: ptr = array_new_f32(n * n) + cblas_sgemm(101, 112, 111, n, n, n, 1.0, W, n, W, n, 0.0, M, n) for i in 0 to n { - M[i] = array_new_f32(n) + let Mi: ptr = M + i * n + let Wi: ptr = W + i * n for j in 0 to n { - let mut s: f32 = 0.0 - if i == j { s = 1.0 } - s = s - W[i][j] - W[j][i] - for k in 0 to n { - s = s + W[k][i] * W[k][j] - } - M[i][j] = s + Mi[j] = Mi[j] - Wi[j] - W[j * n + i] } + Mi[i] = Mi[i] + 1.0 } + array_free_f32(W) - let embedding: ptr = array_new_f32(n * n_components) - let found_vecs: ptr > = malloc((n_components as i64) * 8) as ptr > - for c in 0 to n_components { - let v: ptr = array_new_f32(n) - srand(42 + c) - for i in 0 to n { v[i] = ((rand() % 1000) as f32) / 500.0 - 1.0 } + # Block power iteration. One sgemm per iteration advances every component, + # where the old loop ran a scalar matvec per component per iteration and + # allocated its result buffer inside the iteration. + let V: ptr = array_new_f32(n * kc) + let Z: ptr = array_new_f32(n * kc) + srand(42) + for i in 0 to n * kc { V[i] = ((rand() % 1000) as f32) / 500.0 - 1.0 } + _orthonormalize_block(V, n, kc) - # Orthogonalize against previously found eigenvectors - for prev in 0 to c { - let mut dot: f32 = 0.0 - for i in 0 to n { dot = dot + v[i] * found_vecs[prev][i] } - for i in 0 to n { v[i] = v[i] - dot * found_vecs[prev][i] } - } - - let mut iter: i32 = 0 - while iter < 200 { - let Mv: ptr = array_new_f32(n) - for i in 0 to n { - let mut s: f32 = 0.0 - for j in 0 to n { s = s + M[i][j] * v[j] } - Mv[i] = s - } - - for prev in 0 to c { - let mut dot: f32 = 0.0 - for i in 0 to n { dot = dot + Mv[i] * found_vecs[prev][i] } - for i in 0 to n { Mv[i] = Mv[i] - dot * found_vecs[prev][i] } - } - - let mut norm: f32 = 0.0 - for i in 0 to n { norm = norm + Mv[i] * Mv[i] } - norm = sqrt((norm as f64)) as f32 - if norm < 0.0000000001 { norm = 0.0000000001 } - - let mut max_change: f32 = 0.0 - for i in 0 to n { - let new_v: f32 = Mv[i] / norm - let diff: f32 = new_v - v[i] - if diff < 0.0 { diff = -diff } - if diff > max_change { max_change = diff } - v[i] = new_v - } - array_free_f32(Mv) - - if max_change < 0.000001 { break } - iter = iter + 1 - } + let mut iter: i32 = 0 + while iter < 200 { + blas_matmat(M, V, Z, n, kc, n, 1.0, 0.0) + _orthonormalize_block(Z, n, kc) - for i in 0 to n { - embedding[i * n_components + c] = v[i] + let mut max_change: f32 = 0.0 + for i in 0 to n * kc { + let mut diff: f32 = Z[i] - V[i] + if diff < 0.0 { diff = 0.0 - diff } + if diff > max_change { max_change = diff } + V[i] = Z[i] } - found_vecs[c] = v + if max_change < 0.000001 { break } + iter = iter + 1 } - for c in 0 to n_components { array_free_f32(found_vecs[c]) } - free(found_vecs as ptr) - - for i in 0 to n { array_free_f32(W[i]); array_free_f32(M[i]) } - free(W as ptr) - free(M as ptr) + array_free_f32(M) + array_free_f32(Z) return LLE { - embedding: embedding, + embedding: V, n_samples: n, n_components: n_components, n_neighbors: n_neighbors, @@ -612,123 +611,86 @@ export struct SpectralEmbedding { export function spectral_embedding_fit(X: Matrix, n_components: i32, gamma: f32) -> SpectralEmbedding { let n: i32 = X.rows let p: i32 = X.cols + let k: i32 = n_components # The affinity matrix is symmetric, so one triangle is computed and # mirrored, which halves both the squared distances and the exponentials. - # Elements come off row pointers rather than through matrix_at, which - # copies the Matrix struct for every element and was called twice per - # feature of every pair. + # It is held as one contiguous block rather than an array of row pointers, + # because the iteration below hands it to BLAS. let xdata: ptr = X.data - let W: ptr > = malloc((n as i64) * 8) as ptr > - for i in 0 to n { - W[i] = array_new_f32(n) - } + let A: ptr = array_new_f32(n * n) for i in 0 to n { let ri: ptr = xdata + i * p - W[i][i] = 1.0 + let Ai: ptr = A + i * n + Ai[i] = 1.0 for j in i + 1 to n { let rj: ptr = xdata + j * p let mut acc: f32 = 0.0 - for k in 0 to p { - let d: f32 = ri[k] - rj[k] + for kk in 0 to p { + let d: f32 = ri[kk] - rj[kk] acc = acc + d * d } let wv: f32 = exp((0.0 - gamma * acc) as f64) as f32 - W[i][j] = wv - W[j][i] = wv + Ai[j] = wv + A[j * n + i] = wv } } - let deg: ptr = array_new_f32(n) + # Degrees, then the normalized Laplacian in place. Two n by n buffers used + # to exist, one holding the affinity and one holding the Laplacian, when + # the second is a transform of the first. + let isd: ptr = array_new_f32(n) for i in 0 to n { - for j in 0 to n { deg[i] = deg[i] + W[i][j] } - if deg[i] < 0.0000000001 { deg[i] = 0.0000000001 } + let Ai2: ptr = A + i * n + let mut s: f32 = 0.0 + for j in 0 to n { s = s + Ai2[j] } + if s < 0.0000000001 { s = 0.0000000001 } + isd[i] = 1.0 / (sqrt((s) as f64) as f32) } - - let L: ptr > = malloc((n as i64) * 8) as ptr > for i in 0 to n { - L[i] = array_new_f32(n) + let Ai3: ptr = A + i * n + let si: f32 = isd[i] for j in 0 to n { - if i == j { L[i][j] = 1.0 } - else { L[i][j] = -W[i][j] / sqrt((deg[i] * deg[j]) as f64) as f32 } + if i == j { Ai3[j] = 1.0 } + else { Ai3[j] = 0.0 - Ai3[j] * si * isd[j] } } } - let embedding: ptr = array_new_f32(n * n_components) - let found_vecs: ptr > = malloc((n_components as i64) * 8) as ptr > - # One matvec buffer for every component and every iteration. It used to be - # allocated and freed inside the iteration loop, which is up to 200 - # allocator round trips per component for a vector whose size never - # changes. - let Lv: ptr = array_new_f32(n) + # Block power iteration. One sgemm per iteration advances every component + # at once, where the old loop ran a separate scalar matvec per component + # per iteration and re-orthogonalized against the components already + # found. Orthonormalizing the block each iteration does the same job and + # keeps the whole iteration in BLAS. + let V: ptr = array_new_f32(n * k) + let Z: ptr = array_new_f32(n * k) + srand(42) + for i in 0 to n * k { V[i] = ((rand() % 1000) as f32) / 500.0 - 1.0 } + _orthonormalize_block(V, n, k) - for c in 0 to n_components { - let v: ptr = array_new_f32(n) - srand(42 + c) - for i in 0 to n { v[i] = ((rand() % 1000) as f32) / 500.0 - 1.0 } + let mut iter: i32 = 0 + while iter < 200 { + blas_matmat(A, V, Z, n, k, n, 1.0, 0.0) + _orthonormalize_block(Z, n, k) - # Orthogonalize against previously found eigenvectors - for prev in 0 to c { - let mut dot: f32 = 0.0 - for i in 0 to n { dot = dot + v[i] * found_vecs[prev][i] } - for i in 0 to n { v[i] = v[i] - dot * found_vecs[prev][i] } + let mut max_change: f32 = 0.0 + for i in 0 to n * k { + let mut diff: f32 = Z[i] - V[i] + if diff < 0.0 { diff = 0.0 - diff } + if diff > max_change { max_change = diff } + V[i] = Z[i] } - - let mut iter: i32 = 0 - while iter < 200 { - for i in 0 to n { - # The row pointer is taken once. Indexing L[i][j] in the inner - # loop is two loads per element, one of them the same row - # pointer over and over. - let Li: ptr = L[i] - let mut s: f32 = 0.0 - for j in 0 to n { s = s + Li[j] * v[j] } - Lv[i] = s - } - - for prev in 0 to c { - let mut dot: f32 = 0.0 - for i in 0 to n { dot = dot + Lv[i] * found_vecs[prev][i] } - for i in 0 to n { Lv[i] = Lv[i] - dot * found_vecs[prev][i] } - } - - let mut norm: f32 = 0.0 - for i in 0 to n { norm = norm + Lv[i] * Lv[i] } - norm = sqrt((norm as f64)) as f32 - if norm < 0.0000000001 { norm = 0.0000000001 } - - let mut max_change: f32 = 0.0 - for i in 0 to n { - let new_v: f32 = Lv[i] / norm - let diff: f32 = new_v - v[i] - if diff < 0.0 { diff = -diff } - if diff > max_change { max_change = diff } - v[i] = new_v - } - if max_change < 0.000001 { break } - iter = iter + 1 - } - - for i in 0 to n { - embedding[i * n_components + c] = v[i] - } - found_vecs[c] = v + if max_change < 0.000001 { break } + iter = iter + 1 } - for c in 0 to n_components { array_free_f32(found_vecs[c]) } - free(found_vecs as ptr) - - for i in 0 to n { array_free_f32(W[i]); array_free_f32(L[i]) } - free(W as ptr) - free(L as ptr) - array_free_f32(deg) - - array_free_f32(Lv) + array_free_f32(A) + array_free_f32(isd) + array_free_f32(Z) return SpectralEmbedding { - embedding: embedding, + embedding: V, n_samples: n, - n_components: n_components, + n_components: k, fitted: true } } @@ -763,59 +725,100 @@ export struct MDS { export function mds_fit(X: Matrix, n_components: i32, max_iter: i32, tol: f32, seed: i32) -> MDS { let n: i32 = X.rows let p: i32 = X.cols + let k: i32 = n_components - let D: ptr > = malloc((n as i64) * 8) as ptr > + # Target distances. Symmetric, so one triangle is computed and mirrored, + # off row pointers rather than through matrix_at, and held contiguous + # instead of as an array of row allocations. + let xdata: ptr = X.data + let D: ptr = array_new_f32(n * n) for i in 0 to n { - D[i] = array_new_f32(n) - for j in 0 to n { - let mut s: f32 = 0.0 - for k in 0 to p { - let d: f32 = matrix_at(X, i, k) - matrix_at(X, j, k) - s = s + d * d + let ri: ptr = xdata + i * p + let Di: ptr = D + i * n + for j in i + 1 to n { + let rj: ptr = xdata + j * p + let mut acc: f32 = 0.0 + for c in 0 to p { + let d: f32 = ri[c] - rj[c] + acc = acc + d * d } - D[i][j] = sqrt((s as f64)) as f32 + let dv: f32 = sqrt((acc) as f64) as f32 + Di[j] = dv + D[j * n + i] = dv } } srand(seed) - let embedding: ptr = array_new_f32(n * n_components) - for i in 0 to n * n_components { + let embedding: ptr = array_new_f32(n * k) + for i in 0 to n * k { embedding[i] = ((rand() % 1000) as f32) / 500.0 - 1.0 } + # SMACOF. The Guttman transform is the exact minimizer of the function + # that majorizes the stress at the current configuration, so one step + # moves as far as the majorization allows. The old loop took a fixed + # hundredth of the gradient instead, which needs many more sweeps to + # cover the same ground. + # + # A pair is visited once and its contribution is added to both of its + # points, which halves the square roots and the divisions, and the stress + # for the sweep falls out of the same pass. + let acc_pos: ptr = array_new_f32(n * k) + let inv_n: f32 = 1.0 / (n as f32) let mut n_iter: i32 = 0 let mut stress: f32 = 0.0 + let mut prev_stress: f32 = 0.0 + for iter in 0 to max_iter { n_iter = iter + 1 stress = 0.0 + for i in 0 to n * k { acc_pos[i] = 0.0 } for i in 0 to n { - for j in 0 to n { - if i == j { continue } - let mut d_hat: f32 = 0.0 - for c in 0 to n_components { - let d: f32 = embedding[i * n_components + c] - embedding[j * n_components + c] - d_hat = d_hat + d * d + let Di2: ptr = D + i * n + let ei: ptr = embedding + i * k + let ai: ptr = acc_pos + i * k + for j in i + 1 to n { + let ej: ptr = embedding + j * k + let aj: ptr = acc_pos + j * k + let mut sq: f32 = 0.0 + for c in 0 to k { + let d: f32 = ei[c] - ej[c] + sq = sq + d * d } - d_hat = sqrt((d_hat as f64)) as f32 + let mut d_hat: f32 = sqrt((sq) as f64) as f32 if d_hat < 0.0000000001 { d_hat = 0.0000000001 } - let delta: f32 = D[i][j] - d_hat - stress = stress + delta * delta + let target: f32 = Di2[j] + let diff: f32 = target - d_hat + stress = stress + diff * diff - let grad_scale: f32 = delta / d_hat - for c in 0 to n_components { - let g: f32 = grad_scale * (embedding[i * n_components + c] - embedding[j * n_components + c]) - embedding[i * n_components + c] = embedding[i * n_components + c] + 0.01 * g + let ratio: f32 = target / d_hat + for c in 0 to k { + let contrib: f32 = ratio * (ei[c] - ej[c]) + ai[c] = ai[c] + contrib + aj[c] = aj[c] - contrib } } } - if stress / (n as f32) < tol { break } + for i in 0 to n * k { embedding[i] = acc_pos[i] * inv_n } + + # The sweep that stops lowering the stress is the last one worth + # running. The old test compared the raw stress against the tolerance, + # which at these scales is never true, so every fit ran its whole + # iteration budget. + if iter > 0 { + if prev_stress > 0.0 { + let rel: f32 = (prev_stress - stress) / prev_stress + if rel < tol { break } + } + } + prev_stress = stress } - for i in 0 to n { array_free_f32(D[i]) } - free(D as ptr) + array_free_f32(D) + array_free_f32(acc_pos) return MDS { embedding: embedding, diff --git a/lib/scikit/multioutput.flow b/lib/scikit/multioutput.flow index 3e6cf90..34a7e5c 100644 --- a/lib/scikit/multioutput.flow +++ b/lib/scikit/multioutput.flow @@ -131,10 +131,17 @@ export function classifier_chain_fit(X: Matrix, Y: ptr >, n_samples: i3 let order: ptr = malloc((n_outputs as i64) * 4) as ptr for o in 0 to n_outputs { order[o] = o } + # Row-pointer copies. matrix_at and matrix_set each copy the Matrix struct + # by value for every element, and the chain rebuilds a wider design for + # every output. let mut current_X: Matrix = matrix_new(n_samples, n_features) + let xsrc: ptr = X.data + let xdst: ptr = current_X.data for i in 0 to n_samples { + let srow: ptr = xsrc + i * n_features + let drow: ptr = xdst + i * n_features for j in 0 to n_features { - matrix_set(current_X, i, j, matrix_at(X, i, j)) + drow[j] = srow[j] } } @@ -146,10 +153,15 @@ export function classifier_chain_fit(X: Matrix, Y: ptr >, n_samples: i3 } models[o] = logistic_regression_fit(current_X, y, 2, n_iter, lr, penalty_none()) - let augmented: Matrix = matrix_new(n_samples, current_X.cols + 1) + let cur_cols: i32 = current_X.cols + let augmented: Matrix = matrix_new(n_samples, cur_cols + 1) + let csrc: ptr = current_X.data + let adst: ptr = augmented.data for i in 0 to n_samples { - for j in 0 to current_X.cols { - matrix_set(augmented, i, j, matrix_at(current_X, i, j)) + let crow: ptr = csrc + i * cur_cols + let arow: ptr = adst + i * (cur_cols + 1) + for j in 0 to cur_cols { + arow[j] = crow[j] } matrix_set(augmented, i, current_X.cols, y[i]) } @@ -173,10 +185,16 @@ export function classifier_chain_predict(model: ClassifierChain, X: Matrix) -> p let result: ptr > = malloc((n as i64) * 8) as ptr > for i in 0 to n { result[i] = array_new_f32(model.n_outputs) } - let mut current_X: Matrix = matrix_new(n, model.n_features) + # Row-pointer copies, for the reason the fit above gives. + let nf: i32 = model.n_features + let mut current_X: Matrix = matrix_new(n, nf) + let xsrc: ptr = X.data + let xdst: ptr = current_X.data for i in 0 to n { - for j in 0 to model.n_features { - matrix_set(current_X, i, j, matrix_at(X, i, j)) + let srow: ptr = xsrc + i * nf + let drow: ptr = xdst + i * nf + for j in 0 to nf { + drow[j] = srow[j] } } @@ -186,12 +204,17 @@ export function classifier_chain_predict(model: ClassifierChain, X: Matrix) -> p result[i][o] = preds[i] } - let augmented: Matrix = matrix_new(n, current_X.cols + 1) + let cur_cols: i32 = current_X.cols + let augmented: Matrix = matrix_new(n, cur_cols + 1) + let csrc: ptr = current_X.data + let adst: ptr = augmented.data for i in 0 to n { - for j in 0 to current_X.cols { - matrix_set(augmented, i, j, matrix_at(current_X, i, j)) + let crow: ptr = csrc + i * cur_cols + let arow: ptr = adst + i * (cur_cols + 1) + for j in 0 to cur_cols { + arow[j] = crow[j] } - matrix_set(augmented, i, current_X.cols, preds[i]) + arow[cur_cols] = preds[i] } matrix_free(current_X) current_X = augmented diff --git a/lib/scikit/neighbors.flow b/lib/scikit/neighbors.flow index bc35f5c..ceeb059 100644 --- a/lib/scikit/neighbors.flow +++ b/lib/scikit/neighbors.flow @@ -947,58 +947,53 @@ export function kneighbors_transformer_fit(X: Matrix, n_neighbors: i32, mode: i3 export function kneighbors_transformer_transform(model: KNeighborsTransformer, X: Matrix) -> Matrix { let n: i32 = X.rows let n_train: i32 = model.X_train.rows - let k: i32 = model.n_neighbors + let f_cols: i32 = model.X_train.cols + let mut k: i32 = model.n_neighbors if k > n_train { k = n_train } + + # matrix_new returns zeroed storage, so the explicit zero fill it used to + # run was n by n_train calls to matrix_set, each copying the Matrix struct + # by value. The distances come off row pointers for the same reason. let result: Matrix = matrix_new(n, n_train) - let mut i: i32 = 0 - while i < n { - let mut j: i32 = 0 - while j < n_train { - matrix_set(result, i, j, 0.0) - j = j + 1 - } - i = i + 1 - } - i = 0 - while i < n { - let dists: ptr = array_new_f32(n_train) - let mut j: i32 = 0 - while j < n_train { + let rdata: ptr = result.data + let xdata: ptr = X.data + let tdata: ptr = model.X_train.data + let dists: ptr = array_new_f32(n_train) + + for i in 0 to n { + let xrow: ptr = xdata + i * f_cols + for j in 0 to n_train { + let trow: ptr = tdata + j * f_cols let mut d: f32 = 0.0 - let mut f: i32 = 0 - while f < model.X_train.cols { - let diff: f32 = matrix_at(X, i, f) - matrix_at(model.X_train, j, f) + for f in 0 to f_cols { + let diff: f32 = xrow[f] - trow[f] d = d + diff * diff - f = f + 1 } dists[j] = sqrt((d) as f64) as f32 - j = j + 1 } - let mut kk: i32 = 0 - while kk < k { + + let out_row: ptr = rdata + i * n_train + for kk in 0 to k { let mut best_j: i32 = -1 let mut best_d: f32 = 9999999999.0 - let mut j2: i32 = 0 - while j2 < n_train { - if dists[j2] < best_d && dists[j2] >= 0.0 { - best_d = dists[j2] - best_j = j2 + for j2 in 0 to n_train { + let dj: f32 = dists[j2] + if dj >= 0.0 { + if dj < best_d { + best_d = dj + best_j = j2 + } } - j2 = j2 + 1 } if best_j >= 0 { - if model.mode == 1 { - matrix_set(result, i, best_j, best_d) - } else { - matrix_set(result, i, best_j, 1.0) - } - dists[best_j] = -1.0 + if model.mode == 1 { out_row[best_j] = best_d } + else { out_row[best_j] = 1.0 } + dists[best_j] = 0.0 - 1.0 } - kk = kk + 1 } - array_free_f32(dists) - i = i + 1 } + + array_free_f32(dists) return result } diff --git a/lib/scikit/svm.flow b/lib/scikit/svm.flow index 947444b..6ae1c84 100644 --- a/lib/scikit/svm.flow +++ b/lib/scikit/svm.flow @@ -659,14 +659,13 @@ export function linear_svr_fit(X: Matrix, y: ptr, C: f32, epsilon: f32, epo } export function linear_svr_predict(model: LinearSVR, X: Matrix) -> ptr { - let preds: ptr = array_new_f32(X.rows) - for i in 0 to X.rows { - let mut pred: f32 = model.bias - for j in 0 to model.n_features { - pred = pred + model.weights[j] * matrix_at(X, i, j) - } - preds[i] = pred - } + # One sgemv for the whole design. The loop it replaces read every element + # through matrix_at, which copies the Matrix struct per element. + let rows: i32 = X.rows + let preds: ptr = array_new_f32(rows) + blas_matvec(X.data, rows, model.n_features, model.weights, preds, 1.0, 0.0) + let b: f32 = model.bias + for i in 0 to rows { preds[i] = preds[i] + b } return preds } From f2f54397a3eb264989301c4a8b0ef917362ad36f Mon Sep 17 00:00:00 2001 From: godofecht Date: Sun, 27 Sep 2026 21:50:47 +0100 Subject: [PATCH 05/12] Measure the wide estimator matrix in CI, and gate every row The 172 row matrix was driven by a shell loop typed out by hand, which is how a published artifact sat at 145 of 166 while the working tree was at 155, and how a chunk that died mid-run once passed for a chunk with no rows in it. run_estimator_bench.py compiles and runs the generated chunks with every exit status checked, and reuses the binary that the first round leaves behind so later rounds pay for timing instead of for compiling the library again. check_estimator_matrix.py fails on any ranked row slower than scikit-learn, on any row that reported no timing, and on a ranked count that has quietly shrunk. Both run in a new Flow workflow job on a runner that is not competing with anything, which is where a published number belongs. The generator and the scikit-learn harness now time the four estimators the registry marks simplified. compare_estimators.py already showed their times while withholding a ratio, so leaving them out of the race dropped four rows from the page for no reason. The generator also takes a comma separated --only and a --prefix, so a subset can be re-timed without colliding with the committed files in the shared build directory. --- .github/workflows/flow.yml | 63 +++++++++++++ benchmarks/bench_estimators_sklearn.py | 4 +- benchmarks/check_estimator_matrix.py | 66 ++++++++++++++ benchmarks/generate_estimator_bench.py | 25 +++-- benchmarks/run_estimator_bench.py | 121 +++++++++++++++++++++++++ 5 files changed, 271 insertions(+), 8 deletions(-) create mode 100755 benchmarks/check_estimator_matrix.py create mode 100755 benchmarks/run_estimator_bench.py diff --git a/.github/workflows/flow.yml b/.github/workflows/flow.yml index 0b27b5f..c87d677 100644 --- a/.github/workflows/flow.yml +++ b/.github/workflows/flow.yml @@ -157,6 +157,69 @@ jobs: benchmarks/parity_diagnostics.json benchmarks/headline_environment.json + estimator-matrix: + name: Wide estimator matrix + runs-on: ubuntu-latest + # Both sides call into OpenBLAS, so the thread count is pinned here for the + # reason the canonical job pins it. + env: + OPENBLAS_NUM_THREADS: "4" + OMP_NUM_THREADS: "4" + MKL_NUM_THREADS: "4" + steps: + - uses: actions/checkout@v4 + + - uses: actions/checkout@v4 + with: + repository: flooooooooooow/flow + ref: 88aac5095488309813c7173d268c1f8260421c6e + path: .flow-toolchain + + - uses: actions/setup-python@v5 + with: + python-version: "3.12" + + - name: Install benchmark dependencies + run: | + pip install -r .flow-toolchain/requirements.txt + pip install numpy scikit-learn + sudo apt-get update && sudo apt-get install -y libopenblas-dev + + - name: Rebuild the registry and the generated benchmarks + run: | + python benchmarks/estimator_coverage.py + python benchmarks/generate_estimator_bench.py + git diff --exit-code benchmarks/estimator_coverage.json benchmarks/generated + + - name: Time every estimator on the Flow side + env: + FLOW_BIN: ${{ github.workspace }}/.flow-toolchain/flow + FLOW_HOST: python + FLOW_OPT_LEVEL: "3" + FLOW_LDFLAGS: "-lm -lopenblas lib/scikit/flow_time.c lib/scikit/flow_parallel.c" + run: python benchmarks/run_estimator_bench.py --rounds 3 --out benchmarks/estimator_flow_raw.txt + + - name: Time every estimator on the scikit-learn side + run: python benchmarks/bench_estimators_sklearn.py --repeats 5 + + - name: Join the two sides + run: | + python benchmarks/compare_estimators.py benchmarks/estimator_flow_raw.txt \ + --note "GitHub Actions ubuntu-latest, Flow at -O3 with adaptive repeats, fastest of 3 rounds; scikit-learn wheel, fastest of 5; OpenBLAS pinned to 4 threads" + + - name: Gate every ranked row against scikit-learn + run: python benchmarks/check_estimator_matrix.py benchmarks/estimator_comparison.json + + - uses: actions/upload-artifact@v4 + if: always() + with: + name: estimator-matrix-${{ github.run_id }} + path: | + benchmarks/estimator_flow_raw.txt + benchmarks/estimators_sklearn.json + benchmarks/estimator_comparison.json + benchmarks/estimator_coverage.json + scaled-report: name: Scaled benchmark report runs-on: ubuntu-latest diff --git a/benchmarks/bench_estimators_sklearn.py b/benchmarks/bench_estimators_sklearn.py index a99c9cc..5d5ea9b 100644 --- a/benchmarks/bench_estimators_sklearn.py +++ b/benchmarks/bench_estimators_sklearn.py @@ -91,7 +91,9 @@ def main() -> int: args = ap.parse_args() registry = json.loads(REGISTRY.read_text()) - runnable = [e for e in registry["entries"] if e["bucket"] == "runnable"] + # Matches generate_estimator_bench.py: simplified rows are timed on both + # sides and shown without a ratio. + runnable = [e for e in registry["entries"] if e["bucket"] in ("runnable", "simplified")] classes = dict(all_estimators()) constructors = _constructors() diff --git a/benchmarks/check_estimator_matrix.py b/benchmarks/check_estimator_matrix.py new file mode 100755 index 0000000..6cd1a31 --- /dev/null +++ b/benchmarks/check_estimator_matrix.py @@ -0,0 +1,66 @@ +#!/usr/bin/env python3 +"""Gate the wide estimator matrix: every ranked row has to favour Flow. + +The matrix was measured by hand for a long time, which is why a published +number sat at 145 of 166 while the working tree was at 155. A gate in CI is +what keeps the claim and the code in step: a row that turns slower than +scikit-learn fails the run, and a row that stops reporting fails it too, +because a chunk that dies mid-run is indistinguishable from a chunk with +nothing to say. + +Rows the registry marks simplified carry no ratio and are not ranked. Rows +under the clock's floor carry no ratio either. Both are reported and neither +fails the run. +""" +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[1] + + +def main() -> int: + ap = argparse.ArgumentParser() + ap.add_argument("comparison", nargs="?", type=Path, + default=ROOT / "benchmarks" / "estimator_comparison.json") + ap.add_argument("--min-compared", type=int, default=160, + help="fail if fewer rows than this produced a ratio") + ap.add_argument("--tolerance", type=float, default=1.0, + help="the ratio a row has to reach, sklearn_ms / flow_ms") + args = ap.parse_args() + + payload = json.loads(args.comparison.read_text()) + rows = payload["rows"] + ranked = [r for r in rows if r.get("status") == "ok" and r.get("speedup") is not None] + losers = sorted((r for r in ranked if r["speedup"] < args.tolerance), + key=lambda r: r["speedup"]) + missing = [r for r in rows if r.get("status") in ("flow_missing", "sklearn_missing")] + + print(f"ranked {len(ranked)} rows, {len(ranked) - len(losers)} at or above " + f"{args.tolerance:.2f}x") + for r in losers: + print(f" SLOWER {r['speedup']:6.3f}x {r['flow_estimator']:32s} " + f"flow={r['flow_ms']:9.3f} sklearn={r['sklearn_ms']:9.3f}") + for r in missing: + print(f" MISSING {r['flow_estimator']:32s} {r['status']}: {r.get('reason', '')}") + + failed = False + if losers: + print(f"{len(losers)} rows are slower than scikit-learn") + failed = True + if missing: + print(f"{len(missing)} rows reported no timing") + failed = True + if len(ranked) < args.min_compared: + print(f"only {len(ranked)} rows produced a ratio, expected at least {args.min_compared}") + failed = True + if failed: + return 1 + print("every ranked row favours Flow") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/benchmarks/generate_estimator_bench.py b/benchmarks/generate_estimator_bench.py index 52b601d..3c02d20 100644 --- a/benchmarks/generate_estimator_bench.py +++ b/benchmarks/generate_estimator_bench.py @@ -267,24 +267,35 @@ def main() -> int: ap = argparse.ArgumentParser() ap.add_argument("--per-file", type=int, default=20) ap.add_argument("--outdir", type=Path, default=OUTDIR) - ap.add_argument("--only", help="generate a single estimator, for bisecting a compile failure") + ap.add_argument("--only", help="comma separated estimator names, for bisecting a compile " + "failure or re-timing one group") + ap.add_argument("--prefix", default="bench_estimators", + help="output file prefix, so a subset written elsewhere cannot collide with " + "the committed files or with another process in the shared build directory") args = ap.parse_args() registry = json.loads(REGISTRY.read_text()) - runnable = [e for e in registry["entries"] if e["bucket"] == "runnable"] + # Simplified implementations are timed too, because compare_estimators.py + # shows their times while withholding a ratio. Leaving them out of the + # race would make the page quietly drop four rows. + timed = ("runnable", "simplified") + runnable = [e for e in registry["entries"] if e["bucket"] in timed] if args.only: - runnable = [e for e in runnable if e["flow_estimator"] == args.only] - if not runnable: - raise SystemExit(f"{args.only} is not a runnable registry entry") + wanted = [n.strip() for n in args.only.split(",") if n.strip()] + runnable = [e for e in runnable if e["flow_estimator"] in wanted] + found = {e["flow_estimator"] for e in runnable} + missing = [n for n in wanted if n not in found] + if missing: + raise SystemExit(f"not timed registry entries: {', '.join(missing)}") args.outdir.mkdir(parents=True, exist_ok=True) - for stale in args.outdir.glob("bench_estimators_*.flow"): + for stale in args.outdir.glob(f"{args.prefix}_*.flow"): stale.unlink() written = [] for i in range(0, len(runnable), args.per_file): part = runnable[i : i + args.per_file] - path = args.outdir / f"bench_estimators_{i // args.per_file:02d}.flow" + path = args.outdir / f"{args.prefix}_{i // args.per_file:02d}.flow" path.write_text(chunk_file(i // args.per_file, part)) written.append((path.name, len(part))) diff --git a/benchmarks/run_estimator_bench.py b/benchmarks/run_estimator_bench.py new file mode 100755 index 0000000..ab9e43a --- /dev/null +++ b/benchmarks/run_estimator_bench.py @@ -0,0 +1,121 @@ +#!/usr/bin/env python3 +"""Compile and run the generated estimator benchmarks, and collect their lines. + +The wide matrix used to be driven by a shell loop typed out by hand, which is +how a chunk that died mid-run once passed for a chunk that had no rows. This +does the same work with the failures visible: every chunk's exit status is +checked, and a chunk that dies takes its own row count down with it in the +summary rather than disappearing. + +Each chunk prints one `ESTIMATOR|name|fit_ms|pred_ms|reps|ok` line per +estimator and flushes it, so a crash costs only the rows after it. +compare_estimators.py keeps the fastest observation per estimator, so several +rounds of the same chunk are appended to one file and joined there. + +The first round compiles through `flow run`. Later rounds re-exec the binary +that left behind, because compiling the library at -O3 costs far more than +running the timings. +""" +from __future__ import annotations + +import argparse +import os +import platform +import re +import shutil +import subprocess +import sys +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[1] +GENERATED = ROOT / "benchmarks" / "generated" +LINE = re.compile(r"^ESTIMATOR\|([a-z_0-9]+)\|") + +MAC_LDFLAGS = "-framework Accelerate lib/scikit/flow_time.c lib/scikit/flow_parallel.c" +LINUX_LDFLAGS = "-lm -lopenblas lib/scikit/flow_time.c lib/scikit/flow_parallel.c" + + +def default_ldflags() -> str: + return MAC_LDFLAGS if platform.system() == "Darwin" else LINUX_LDFLAGS + + +def flow_binary() -> str: + explicit = os.environ.get("FLOW_BIN") + if explicit: + return explicit + found = shutil.which("flow") + if not found: + raise SystemExit("no flow binary on PATH and FLOW_BIN is unset") + return found + + +def built_binary(flow_bin: str, stem: str) -> Path: + # flow run writes its C and its executable into one build directory next to + # the compiler, whatever directory the source came from. + return Path(flow_bin).resolve().parent / "build" / stem + + +def run(cmd: list[str], env: dict[str, str], cwd: Path, timeout: int) -> tuple[int, str]: + try: + proc = subprocess.run( + cmd, cwd=cwd, env=env, capture_output=True, text=True, timeout=timeout + ) + except subprocess.TimeoutExpired as exc: + return 124, (exc.stdout or "") if isinstance(exc.stdout, str) else "" + return proc.returncode, proc.stdout + proc.stderr + + +def main() -> int: + ap = argparse.ArgumentParser() + ap.add_argument("--rounds", type=int, default=3) + ap.add_argument("--out", type=Path, default=ROOT / "benchmarks" / "estimator_flow_raw.txt") + ap.add_argument("--outdir", type=Path, default=GENERATED) + ap.add_argument("--opt", default="3") + ap.add_argument("--timeout", type=int, default=1800) + ap.add_argument("--only", help="one chunk stem, for bisecting a failure") + args = ap.parse_args() + + chunks = sorted(args.outdir.glob("bench_estimators_*.flow")) + if args.only: + chunks = [c for c in chunks if c.stem == args.only] + if not chunks: + raise SystemExit(f"no generated benchmarks in {args.outdir}") + + flow_bin = flow_binary() + env = dict(os.environ) + env["FLOW_HOST"] = env.get("FLOW_HOST", "python") + env["FLOW_OPT_LEVEL"] = args.opt + env["FLOW_LDFLAGS"] = env.get("FLOW_LDFLAGS") or default_ldflags() + + collected: list[str] = [] + seen: set[str] = set() + failures: list[str] = [] + + for round_index in range(args.rounds): + for chunk in chunks: + binary = built_binary(flow_bin, chunk.stem) + if round_index == 0 or not binary.exists(): + cmd = [flow_bin, "run", str(chunk)] + else: + cmd = [str(binary)] + code, output = run(cmd, env, ROOT, args.timeout) + rows = [l for l in output.splitlines() if LINE.match(l.strip())] + collected.extend(rows) + for line in rows: + seen.add(line.split("|")[1]) + status = "ok" if code == 0 else f"exit={code}" + if code != 0: + failures.append(f"round {round_index} {chunk.stem} {status}") + print(f"round {round_index} {chunk.stem}: {len(rows)} rows {status}", flush=True) + + args.out.write_text("\n".join(collected) + "\n") + print(f"{len(seen)} estimators, {len(collected)} lines -> {args.out}") + if failures: + print("failures:") + for f in failures: + print(f" {f}") + return 1 if failures else 0 + + +if __name__ == "__main__": + sys.exit(main()) From c24460465440c422a7ae8abb84bbb297d3e4bc18 Mon Sep 17 00:00:00 2001 From: godofecht Date: Sun, 27 Sep 2026 22:10:51 +0100 Subject: [PATCH 06/12] Race thirteen more estimators, and fix the fit that crashed CI The wide matrix raced 172 of 203 exported estimators. Twenty-five sat in different_shape, which meant only that their fit does not begin with a feature matrix and the generic path builds one shape of call. Thirteen of them race scikit-learn perfectly well once the call is written out, so the registry now carries that call and they are ranked like any other row: the seven kernel approximation and projection transformers, the two dummy estimators, the two label encoders, the isotonic fit and the model selector. The twelve that are left take a Pipeline, a ColumnTransformer, a FeatureUnion, an estimator array, a document array or a dict array, and each keeps its written reason. factor_analysis_fit allocated its component array with malloc, tested each row against null, and wrote through whatever the allocator had left there when the test failed. On macOS that memory happened to be zero, so the fit worked and the row has been timed for weeks. On the CI runner it was not, and the process died partway through its chunk, taking the nine estimators after it down with it. Every row is allocated up front now. The harness also adds a sink. Two rows came back at exactly 0.000000 ms because clang at -O3 is free to delete a transform whose result is freed without being read. One value from every result is now folded into a running total that main prints on a line the parser ignores. --- benchmarks/bench_estimators_sklearn.py | 37 +- benchmarks/compare_estimators.py | 2 +- benchmarks/estimator_coverage.json | 241 +++++++-- benchmarks/estimator_coverage.py | 109 ++++- benchmarks/generate_estimator_bench.py | 96 +++- benchmarks/generated/bench_estimators_00.flow | 86 +++- benchmarks/generated/bench_estimators_01.flow | 192 +++++--- benchmarks/generated/bench_estimators_02.flow | 281 ++++++----- benchmarks/generated/bench_estimators_03.flow | 360 ++++++++------ benchmarks/generated/bench_estimators_04.flow | 373 +++++++------- benchmarks/generated/bench_estimators_05.flow | 400 ++++++++------- benchmarks/generated/bench_estimators_06.flow | 456 ++++++++++-------- benchmarks/generated/bench_estimators_07.flow | 456 ++++++++++-------- benchmarks/generated/bench_estimators_08.flow | 447 ++++++++++++----- benchmarks/generated/bench_estimators_09.flow | 224 +++++++++ lib/scikit/decomposition.flow | 9 +- 16 files changed, 2499 insertions(+), 1270 deletions(-) create mode 100644 benchmarks/generated/bench_estimators_09.flow diff --git a/benchmarks/bench_estimators_sklearn.py b/benchmarks/bench_estimators_sklearn.py index 5d5ea9b..f276f67 100644 --- a/benchmarks/bench_estimators_sklearn.py +++ b/benchmarks/bench_estimators_sklearn.py @@ -57,6 +57,8 @@ def _constructors() -> dict: "SparseCoder": lambda c: c(dictionary=np.eye(4)), # nu=0.5 is infeasible for this class balance. "NuSVC": lambda c: c(nu=0.1), + # A selector needs something to read importances from. + "SelectFromModel": lambda c: c(tree_c(), threshold=-np.inf, max_features=2), } @@ -91,9 +93,10 @@ def main() -> int: args = ap.parse_args() registry = json.loads(REGISTRY.read_text()) - # Matches generate_estimator_bench.py: simplified rows are timed on both - # sides and shown without a ratio. - runnable = [e for e in registry["entries"] if e["bucket"] in ("runnable", "simplified")] + # Matches generate_estimator_bench.py: shaped rows are raced and ranked, + # simplified rows are timed on both sides and shown without a ratio. + runnable = [e for e in registry["entries"] + if e["bucket"] in ("runnable", "shaped", "simplified")] classes = dict(all_estimators()) constructors = _constructors() @@ -117,19 +120,39 @@ def main() -> int: rows.append({"flow_estimator": entry["flow_estimator"], "sklearn_estimator": name, "status": "unavailable", "reason": "not in sklearn.utils.all_estimators()"}) continue - kind = dataset_kind(entry) + shape = entry.get("shape") + kind = shape["dataset"] if shape else dataset_kind(entry) X, y = data[kind] + ctor_kwargs = {} + if shape and shape.get("sklearn_ctor"): + # The string is the constructor's own keyword list, kept next to the + # Flow call it matches so the two cannot drift apart. + ctor_kwargs = eval(f"dict({shape['sklearn_ctor']})") # noqa: S307 + # A row whose scikit-learn side fits a target vector or a single ordered + # variable rather than a design. The Flow side of these is written out + # in the registry for the same reason. + fit_input = (shape or {}).get("sklearn_input", "X") + if fit_input == "y": + first, second = y, None + elif fit_input == "x1d": + first, second = X[:, 0], y + else: + first, second = X, y try: with warnings.catch_warnings(): warnings.simplefilter("ignore") build = constructors.get(name) - model = build(cls) if build else cls() - fit_ms = timed((lambda: model.fit(X)) if y is None else (lambda: model.fit(X, y)), args.repeats) + model = build(cls) if build else cls(**ctor_kwargs) + fit_ms = timed( + (lambda: model.fit(first)) if second is None + else (lambda: model.fit(first, second)), + args.repeats, + ) pred_ms = 0.0 for method in ("predict", "transform"): if hasattr(model, method): try: - pred_ms = timed(lambda m=method: getattr(model, m)(X), args.repeats) + pred_ms = timed(lambda m=method: getattr(model, m)(first), args.repeats) except Exception: pred_ms = 0.0 break diff --git a/benchmarks/compare_estimators.py b/benchmarks/compare_estimators.py index 3973b57..2fff435 100644 --- a/benchmarks/compare_estimators.py +++ b/benchmarks/compare_estimators.py @@ -55,7 +55,7 @@ def main() -> int: rows = [] for entry in registry["entries"]: name = entry["flow_estimator"] - if entry["bucket"] != "runnable": + if entry["bucket"] not in ("runnable", "shaped"): row = {"flow_estimator": name, "status": entry["bucket"], "reason": entry.get("reason", "")} if entry.get("sklearn_estimator"): row["sklearn_estimator"] = entry["sklearn_estimator"] diff --git a/benchmarks/estimator_coverage.json b/benchmarks/estimator_coverage.json index cc1fbea..f9eccb6 100644 --- a/benchmarks/estimator_coverage.json +++ b/benchmarks/estimator_coverage.json @@ -3,7 +3,8 @@ "counts": { "estimators": 203, "runnable": 168, - "different_shape": 25, + "shaped": 13, + "different_shape": 12, "simplified": 4, "flow_only": 6 }, @@ -196,9 +197,20 @@ "n_features", "sample_steps" ], - "bucket": "different_shape", - "reason": "takes i32 first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "AdditiveChi2Sampler" + "bucket": "shaped", + "sklearn_estimator": "AdditiveChi2Sampler", + "shape": { + "dataset": "classification", + "flow_fit": [ + "n_c", + "f_c", + "2" + ], + "flow_work": [ + "X_c" + ], + "sklearn_input": "X" + } }, { "flow_estimator": "affinity_propagation", @@ -1719,9 +1731,23 @@ "constant_label", "seed" ], - "bucket": "different_shape", - "reason": "takes ptr first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "DummyClassifier" + "bucket": "shaped", + "sklearn_estimator": "DummyClassifier", + "shape": { + "dataset": "classification", + "flow_fit": [ + "y_c", + "n_c", + "3", + "0", + "0.0", + "42" + ], + "flow_work": [ + "n_c" + ], + "sklearn_input": "X" + } }, { "flow_estimator": "dummy_regressor", @@ -1780,9 +1806,21 @@ "strategy", "constant_value" ], - "bucket": "different_shape", - "reason": "takes ptr first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "DummyRegressor" + "bucket": "shaped", + "sklearn_estimator": "DummyRegressor", + "shape": { + "dataset": "regression", + "flow_fit": [ + "y_r", + "n_r", + "0", + "0.0" + ], + "flow_work": [ + "n_r" + ], + "sklearn_input": "X" + } }, { "flow_estimator": "elastic_net_cv", @@ -2897,9 +2935,21 @@ "n_components", "seed" ], - "bucket": "different_shape", - "reason": "takes i32 first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "GaussianRandomProjection" + "bucket": "shaped", + "sklearn_estimator": "GaussianRandomProjection", + "shape": { + "dataset": "classification", + "flow_fit": [ + "f_c", + "2", + "42" + ], + "flow_work": [ + "X_c" + ], + "sklearn_input": "X", + "sklearn_ctor": "n_components=2, random_state=42" + } }, { "flow_estimator": "gradient_boosting_classifier", @@ -3604,9 +3654,22 @@ "n", "increasing" ], - "bucket": "different_shape", - "reason": "takes ptr first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "IsotonicRegression" + "bucket": "shaped", + "sklearn_estimator": "IsotonicRegression", + "shape": { + "dataset": "regression", + "flow_fit": [ + "x1d_r", + "y_r", + "n_r", + "true" + ], + "flow_work": [ + "x1d_r", + "n_r" + ], + "sklearn_input": "x1d" + } }, { "flow_estimator": "iterative_imputer", @@ -4456,9 +4519,22 @@ "neg_label", "pos_label" ], - "bucket": "different_shape", - "reason": "takes ptr first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "LabelBinarizer" + "bucket": "shaped", + "sklearn_estimator": "LabelBinarizer", + "shape": { + "dataset": "classification", + "flow_fit": [ + "y_c", + "n_c", + "0.0", + "1.0" + ], + "flow_work": [ + "y_c", + "n_c" + ], + "sklearn_input": "y" + } }, { "flow_estimator": "label_encoder", @@ -4511,9 +4587,20 @@ "y", "n" ], - "bucket": "different_shape", - "reason": "takes ptr first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "LabelEncoder" + "bucket": "shaped", + "sklearn_estimator": "LabelEncoder", + "shape": { + "dataset": "classification", + "flow_fit": [ + "y_c", + "n_c" + ], + "flow_work": [ + "y_c", + "n_c" + ], + "sklearn_input": "y" + } }, { "flow_estimator": "label_propagation", @@ -8771,9 +8858,21 @@ "n_components", "degree" ], - "bucket": "different_shape", - "reason": "takes i32 first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "PolynomialCountSketch" + "bucket": "shaped", + "sklearn_estimator": "PolynomialCountSketch", + "shape": { + "dataset": "classification", + "flow_fit": [ + "f_c", + "2", + "2" + ], + "flow_work": [ + "X_c" + ], + "sklearn_input": "X", + "sklearn_ctor": "n_components=2, degree=2, random_state=42" + } }, { "flow_estimator": "polynomial_features", @@ -8832,9 +8931,21 @@ "interaction_only", "include_bias" ], - "bucket": "different_shape", - "reason": "takes i32 first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "PolynomialFeatures" + "bucket": "shaped", + "sklearn_estimator": "PolynomialFeatures", + "shape": { + "dataset": "classification", + "flow_fit": [ + "f_c", + "2", + "false", + "true" + ], + "flow_work": [ + "X_c" + ], + "sklearn_input": "X" + } }, { "flow_estimator": "power_transformer", @@ -9577,9 +9688,22 @@ "n_components", "seed" ], - "bucket": "different_shape", - "reason": "takes i32 first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "RBFSampler" + "bucket": "shaped", + "sklearn_estimator": "RBFSampler", + "shape": { + "dataset": "classification", + "flow_fit": [ + "f_c", + "0.1", + "2", + "42" + ], + "flow_work": [ + "X_c" + ], + "sklearn_input": "X", + "sklearn_ctor": "gamma=0.1, n_components=2, random_state=42" + } }, { "flow_estimator": "regressor_chain", @@ -10238,9 +10362,20 @@ "n_features", "threshold" ], - "bucket": "different_shape", - "reason": "takes ptr first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "SelectFromModel" + "bucket": "shaped", + "sklearn_estimator": "SelectFromModel", + "shape": { + "dataset": "classification", + "flow_fit": [ + "w_f", + "f_c", + "0.5" + ], + "flow_work": [ + "X_c" + ], + "sklearn_input": "X" + } }, { "flow_estimator": "select_fwe", @@ -10921,9 +11056,22 @@ "n_components", "seed" ], - "bucket": "different_shape", - "reason": "takes i32 first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "SkewedChi2Sampler" + "bucket": "shaped", + "sklearn_estimator": "SkewedChi2Sampler", + "shape": { + "dataset": "classification", + "flow_fit": [ + "f_c", + "1.0", + "2", + "42" + ], + "flow_work": [ + "X_c" + ], + "sklearn_input": "X", + "sklearn_ctor": "skewedness=1.0, n_components=2, random_state=42" + } }, { "flow_estimator": "sparse_coder", @@ -11102,9 +11250,22 @@ "density", "seed" ], - "bucket": "different_shape", - "reason": "takes i32 first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "SparseRandomProjection" + "bucket": "shaped", + "sklearn_estimator": "SparseRandomProjection", + "shape": { + "dataset": "classification", + "flow_fit": [ + "f_c", + "2", + "0.3", + "42" + ], + "flow_work": [ + "X_c" + ], + "sklearn_input": "X", + "sklearn_ctor": "n_components=2, density=0.3, random_state=42" + } }, { "flow_estimator": "spectral_biclustering", diff --git a/benchmarks/estimator_coverage.py b/benchmarks/estimator_coverage.py index ef436c5..d969bee 100644 --- a/benchmarks/estimator_coverage.py +++ b/benchmarks/estimator_coverage.py @@ -7,9 +7,11 @@ from: every exported `*_fit`, its companion predict/transform, the arguments to call it with, and the scikit-learn class to race it against. -An estimator lands in one of three buckets, and every one carries a reason: +An estimator lands in one of these buckets, and every one carries a reason: runnable arguments resolved and a scikit-learn counterpart exists + shaped raced through a written out call, because its fit does not begin + with a feature matrix flow_only Flow implements it and scikit-learn has no equivalent blocked something about the signature is not resolved yet @@ -175,6 +177,103 @@ "n_categories": ("array", "[4, 4, 4, 4]"), } +# Estimators whose fit does not begin with a feature matrix, and which race +# scikit-learn perfectly well once the call is written out. Thirteen rows sat +# in different_shape only because the generic path builds one call shape. +# +# `flow_fit` and `flow_work` are argument expressions in terms of the variables +# the generated harness declares (X_c, y_c, n_c, f_c and the regression pair). +# `sklearn_input` says what the scikit-learn side fits and transforms: the +# feature matrix, the target vector, or the first column of the matrix as a +# one-dimensional x. `sklearn_ctor` overrides the constructor where the default +# would measure a different size of problem, as it would for a random +# projection whose n_components is chosen by Johnson-Lindenstrauss. +SHAPED: dict[str, dict] = { + "additive_chi2_sampler": { + "dataset": "classification", + "flow_fit": ["n_c", "f_c", "2"], + "flow_work": ["X_c"], + "sklearn_input": "X", + }, + "gaussian_random_projection": { + "dataset": "classification", + "flow_fit": ["f_c", "2", "42"], + "flow_work": ["X_c"], + "sklearn_input": "X", + "sklearn_ctor": "n_components=2, random_state=42", + }, + "polynomial_count_sketch": { + "dataset": "classification", + "flow_fit": ["f_c", "2", "2"], + "flow_work": ["X_c"], + "sklearn_input": "X", + "sklearn_ctor": "n_components=2, degree=2, random_state=42", + }, + "polynomial_features": { + "dataset": "classification", + "flow_fit": ["f_c", "2", "false", "true"], + "flow_work": ["X_c"], + "sklearn_input": "X", + }, + "rbf_sampler": { + "dataset": "classification", + "flow_fit": ["f_c", "0.1", "2", "42"], + "flow_work": ["X_c"], + "sklearn_input": "X", + "sklearn_ctor": "gamma=0.1, n_components=2, random_state=42", + }, + "skewed_chi2_sampler": { + "dataset": "classification", + "flow_fit": ["f_c", "1.0", "2", "42"], + "flow_work": ["X_c"], + "sklearn_input": "X", + "sklearn_ctor": "skewedness=1.0, n_components=2, random_state=42", + }, + "sparse_random_projection": { + "dataset": "classification", + "flow_fit": ["f_c", "2", "0.3", "42"], + "flow_work": ["X_c"], + "sklearn_input": "X", + "sklearn_ctor": "n_components=2, density=0.3, random_state=42", + }, + "select_from_model": { + "dataset": "classification", + "flow_fit": ["w_f", "f_c", "0.5"], + "flow_work": ["X_c"], + "sklearn_input": "X", + }, + "dummy_classifier": { + "dataset": "classification", + "flow_fit": ["y_c", "n_c", "3", "0", "0.0", "42"], + "flow_work": ["n_c"], + "sklearn_input": "X", + }, + "dummy_regressor": { + "dataset": "regression", + "flow_fit": ["y_r", "n_r", "0", "0.0"], + "flow_work": ["n_r"], + "sklearn_input": "X", + }, + "label_encoder": { + "dataset": "classification", + "flow_fit": ["y_c", "n_c"], + "flow_work": ["y_c", "n_c"], + "sklearn_input": "y", + }, + "label_binarizer": { + "dataset": "classification", + "flow_fit": ["y_c", "n_c", "0.0", "1.0"], + "flow_work": ["y_c", "n_c"], + "sklearn_input": "y", + }, + "isotonic": { + "dataset": "regression", + "flow_fit": ["x1d_r", "y_r", "n_r", "true"], + "flow_work": ["x1d_r", "n_r"], + "sklearn_input": "x1d", + }, +} + # Parameters whose value depends on the dataset rather than on a constant. DATASET_ARGS = {"n_classes", "n_samples", "n_features"} @@ -302,6 +401,9 @@ def classify(base: str, spec: dict, known: set[str], exports: dict[str, dict]) - unresolved.append(f"{p['name']}: {p['type']}") head = spec["params"][0]["type"] if spec["params"] else "absent" sk = sklearn_name(base, known) + if base in SHAPED and sk is not None: + entry.update(bucket="shaped", sklearn_estimator=sk, shape=SHAPED[base]) + return entry if head != "Matrix": entry.update( bucket="different_shape", @@ -326,7 +428,8 @@ def argument_tables() -> dict: def main() -> int: ap = argparse.ArgumentParser() ap.add_argument("--out", type=Path, default=OUT) - ap.add_argument("--show", choices=["blocked", "flow_only", "runnable", "simplified", "different_shape"], help="list one bucket and exit") + ap.add_argument("--show", choices=["blocked", "flow_only", "runnable", "shaped", "simplified", + "different_shape"], help="list one bucket and exit") args = ap.parse_args() exports = parse_exports() @@ -359,7 +462,7 @@ def main() -> int: return 0 print(f"{len(entries)} exported Flow estimators against a {len(known)}-estimator scikit-learn surface") - for bucket in ("runnable", "simplified", "different_shape", "flow_only", "blocked"): + for bucket in ("runnable", "shaped", "simplified", "different_shape", "flow_only", "blocked"): print(f" {bucket:10s} {counts.get(bucket, 0)}") return 0 diff --git a/benchmarks/generate_estimator_bench.py b/benchmarks/generate_estimator_bench.py index 3c02d20..ffb1651 100644 --- a/benchmarks/generate_estimator_bench.py +++ b/benchmarks/generate_estimator_bench.py @@ -119,6 +119,20 @@ def fit_arguments(entry: dict, kind: str) -> tuple[list[str], list[str]]: return pre, args +# One value out of every result is added to a running sink, which main prints +# on a line the parser ignores. Without it, clang at -O3 is free to delete a +# transform whose result is freed without being read, and two rows came back at +# exactly 0.000000 ms because it did. +def sink_line(var: str, returns: str, indent: str = " ") -> list[str]: + if returns == "Matrix": + return [ + f"{indent}if {var}.rows > 0 {{", + f"{indent} if {var}.cols > 0 {{ sink = sink + {var}.data[0] }}", + f"{indent}}}", + ] + return [f"{indent}sink = sink + {var}[0]"] + + def block(entry: dict) -> str: """One estimator: probe, choose a repeat count, then time fit and predict. @@ -175,6 +189,7 @@ def release(var: str, kind_: str) -> str: lines.append(" t2 = flow_now_ns()") lines.append(" for rep2 in 0 to reps {") lines.append(f" let o_{name}: {comp['returns']} = {comp['name']}(fitted_{name}, X_{suffix})") + lines += sink_line(f"o_{name}", comp["returns"]) lines.append(" " + release(f"o_{name}", comp["returns"])) lines.append(" }") lines.append(" t3 = flow_now_ns()") @@ -194,8 +209,69 @@ def release(var: str, kind_: str) -> str: return "\n".join(lines) + "\n" +def shaped_block(entry: dict) -> str: + """One estimator whose fit does not begin with a feature matrix. + + The registry carries the call for these, because the generic path above + builds one shape of call and these take a target vector, a feature count or + a pair of scalars instead. Everything else about the timing is the same, so + a shaped row is ranked like any other. + """ + name = entry["flow_estimator"] + shape = entry["shape"] + kind = shape["dataset"] + ret = entry["fit"]["returns"] + call = f"{entry['fit']['name']}({', '.join(shape['flow_fit'])})" + free = entry["companions"].get("free") + free_ok = free is not None and len(free["parameters"]) == 1 + comp = entry["companions"].get("transform") or entry["companions"].get("predict") + work = shape.get("flow_work") + comp_ok = comp is not None and work is not None and comp["returns"] in ("Matrix", "ptr") + + lines = [f" # ---- {name} ({kind}, written out) ----"] + lines.append(" t0 = flow_now_ns()") + lines.append(f" let probe_{name}: {ret} = {call}") + lines.append(" t1 = flow_now_ns()") + if free_ok: + lines.append(f" {free['name']}(probe_{name})") + lines.append(" reps = 1") + lines.append(" if (t1 - t0) < 200000 { reps = 200 }") + lines.append(" elif (t1 - t0) < 2000000 { reps = 20 }") + lines.append(" t0 = flow_now_ns()") + lines.append(" for rep in 0 to reps {") + lines.append(f" let m_{name}: {ret} = {call}") + if free_ok: + lines.append(f" {free['name']}(m_{name})") + lines.append(" }") + lines.append(" t1 = flow_now_ns()") + + if comp_ok: + release = "matrix_free" if comp["returns"] == "Matrix" else "array_free_f32" + lines.append(f" let fitted_{name}: {ret} = {call}") + lines.append(" t2 = flow_now_ns()") + lines.append(" for rep2 in 0 to reps {") + lines.append(f" let o_{name}: {comp['returns']} = " + f"{comp['name']}(fitted_{name}, {', '.join(work)})") + lines += sink_line(f"o_{name}", comp["returns"]) + lines.append(f" {release}(o_{name})") + lines.append(" }") + lines.append(" t3 = flow_now_ns()") + if free_ok: + lines.append(f" {free['name']}(fitted_{name})") + pred_expr = "ms_between(t2, t3) / (reps as f32)" + else: + pred_expr = "0.0" + + lines.append( + f' printf("ESTIMATOR|{name}|%.9f|%.9f|%d|ok\\n", ' + f"ms_between(t0, t1) / (reps as f32), {pred_expr}, reps)" + ) + lines.append(" fflush(null)") + return "\n".join(lines) + "\n" + + def chunk_file(index: int, entries: list[dict]) -> str: - body = "\n".join(block(e) for e in entries) + body = "\n".join(shaped_block(e) if e["bucket"] == "shaped" else block(e) for e in entries) return f'''{HEADER} function main() -> i32 {{ let iris: Dataset = load_iris() @@ -243,11 +319,22 @@ def chunk_file(index: int, entries: list[dict]) -> str: Y_rows[i] = row }} + # A one-dimensional x for the isotonic row, which regresses against a + # single ordered variable rather than a design. + let x1d_r: ptr = array_new_f32(n_r) + for i in 0 to n_r {{ x1d_r[i] = matrix_at(X_r, i, 0) }} + + # Per-feature importances for the selector row, which takes the weights a + # fitted model would hand it rather than a design. + let w_f: ptr = array_new_f32(f_c) + for i in 0 to f_c {{ w_f[i] = 1.0 / ((i + 1) as f32) }} + let mut t0: i64 = 0 let mut t1: i64 = 0 let mut t2: i64 = 0 let mut t3: i64 = 0 let mut reps: i32 = 1 + let mut sink: f32 = 0.0 {body} for i in 0 to n_c {{ array_free_f32(Y_label_rows[i]) }} @@ -258,6 +345,11 @@ def chunk_file(index: int, entries: list[dict]) -> str: matrix_free(Y_multi) free(yi_c as ptr) free(yi_r as ptr) + array_free_f32(x1d_r) + array_free_f32(w_f) + # The sink is printed so the work above cannot be optimized away. The + # parser matches ESTIMATOR lines only, so this one is ignored. + printf("SINK|%.9f\\n", sink) return 0 }} ''' @@ -278,7 +370,7 @@ def main() -> int: # Simplified implementations are timed too, because compare_estimators.py # shows their times while withholding a ratio. Leaving them out of the # race would make the page quietly drop four rows. - timed = ("runnable", "simplified") + timed = ("runnable", "shaped", "simplified") runnable = [e for e in registry["entries"] if e["bucket"] in timed] if args.only: wanted = [n.strip() for n in args.only.split(",") if n.strip()] diff --git a/benchmarks/generated/bench_estimators_00.flow b/benchmarks/generated/bench_estimators_00.flow index f67f874..3c09d03 100644 --- a/benchmarks/generated/bench_estimators_00.flow +++ b/benchmarks/generated/bench_estimators_00.flow @@ -65,11 +65,22 @@ function main() -> i32 { Y_rows[i] = row } + # A one-dimensional x for the isotonic row, which regresses against a + # single ordered variable rather than a design. + let x1d_r: ptr = array_new_f32(n_r) + for i in 0 to n_r { x1d_r[i] = matrix_at(X_r, i, 0) } + + # Per-feature importances for the selector row, which takes the weights a + # fitted model would hand it rather than a design. + let w_f: ptr = array_new_f32(f_c) + for i in 0 to f_c { w_f[i] = 1.0 / ((i + 1) as f32) } + let mut t0: i64 = 0 let mut t1: i64 = 0 let mut t2: i64 = 0 let mut t3: i64 = 0 let mut reps: i32 = 1 + let mut sink: f32 = 0.0 # ---- adaboost_classifier (classification) ---- t0 = flow_now_ns() @@ -89,6 +100,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_adaboost_classifier: ptr = adaboost_classifier_predict(fitted_adaboost_classifier, X_c) + sink = sink + o_adaboost_classifier[0] array_free_f32(o_adaboost_classifier) } t3 = flow_now_ns() @@ -114,6 +126,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_adaboost_regressor: ptr = adaboost_regressor_predict(fitted_adaboost_regressor, X_r) + sink = sink + o_adaboost_regressor[0] array_free_f32(o_adaboost_regressor) } t3 = flow_now_ns() @@ -121,6 +134,34 @@ function main() -> i32 { printf("ESTIMATOR|adaboost_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) + # ---- additive_chi2_sampler (classification, written out) ---- + t0 = flow_now_ns() + let probe_additive_chi2_sampler: AdditiveChi2Sampler = additive_chi2_sampler_fit(n_c, f_c, 2) + t1 = flow_now_ns() + additive_chi2_sampler_free(probe_additive_chi2_sampler) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_additive_chi2_sampler: AdditiveChi2Sampler = additive_chi2_sampler_fit(n_c, f_c, 2) + additive_chi2_sampler_free(m_additive_chi2_sampler) + } + t1 = flow_now_ns() + let fitted_additive_chi2_sampler: AdditiveChi2Sampler = additive_chi2_sampler_fit(n_c, f_c, 2) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_additive_chi2_sampler: Matrix = additive_chi2_sampler_transform(fitted_additive_chi2_sampler, X_c) + if o_additive_chi2_sampler.rows > 0 { + if o_additive_chi2_sampler.cols > 0 { sink = sink + o_additive_chi2_sampler.data[0] } + } + matrix_free(o_additive_chi2_sampler) + } + t3 = flow_now_ns() + additive_chi2_sampler_free(fitted_additive_chi2_sampler) + printf("ESTIMATOR|additive_chi2_sampler|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- affinity_propagation (unsupervised) ---- t0 = flow_now_ns() let probe_affinity_propagation: AffinityPropagation = affinity_propagation_fit(X_c, 0.5, 100, 15) @@ -173,6 +214,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_ard_regression: ptr = ard_regression_predict(fitted_ard_regression, X_r) + sink = sink + o_ard_regression[0] array_free_f32(o_ard_regression) } t3 = flow_now_ns() @@ -198,6 +240,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_bagging_classifier: ptr = bagging_classifier_predict(fitted_bagging_classifier, X_c) + sink = sink + o_bagging_classifier[0] array_free_f32(o_bagging_classifier) } t3 = flow_now_ns() @@ -223,6 +266,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_bagging_regressor: ptr = bagging_regressor_predict(fitted_bagging_regressor, X_r) + sink = sink + o_bagging_regressor[0] array_free_f32(o_bagging_regressor) } t3 = flow_now_ns() @@ -248,6 +292,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_bayesian_gaussian_mixture: ptr = bayesian_gaussian_mixture_predict(fitted_bayesian_gaussian_mixture, X_c) + sink = sink + o_bayesian_gaussian_mixture[0] array_free_f32(o_bayesian_gaussian_mixture) } t3 = flow_now_ns() @@ -273,6 +318,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_bayesian_ridge: ptr = bayesian_ridge_predict(fitted_bayesian_ridge, X_r) + sink = sink + o_bayesian_ridge[0] array_free_f32(o_bayesian_ridge) } t3 = flow_now_ns() @@ -298,6 +344,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_bernoulli_nb: ptr = bernoulli_nb_predict(fitted_bernoulli_nb, X_c) + sink = sink + o_bernoulli_nb[0] array_free_f32(o_bernoulli_nb) } t3 = flow_now_ns() @@ -323,6 +370,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_bernoulli_rbm: Matrix = bernoulli_rbm_transform(fitted_bernoulli_rbm, X_c) + if o_bernoulli_rbm.rows > 0 { + if o_bernoulli_rbm.cols > 0 { sink = sink + o_bernoulli_rbm.data[0] } + } matrix_free(o_bernoulli_rbm) } t3 = flow_now_ns() @@ -382,6 +432,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_calibrated_classifier_cv: ptr = calibrated_classifier_cv_predict(fitted_calibrated_classifier_cv, X_c) + sink = sink + o_calibrated_classifier_cv[0] array_free_f32(o_calibrated_classifier_cv) } t3 = flow_now_ns() @@ -408,6 +459,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_categorical_nb: ptr = categorical_nb_predict(fitted_categorical_nb, X_c) + sink = sink + o_categorical_nb[0] array_free_f32(o_categorical_nb) } t3 = flow_now_ns() @@ -433,6 +485,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_cca: Matrix = cca_transform(fitted_cca, X_r) + if o_cca.rows > 0 { + if o_cca.cols > 0 { sink = sink + o_cca.data[0] } + } matrix_free(o_cca) } t3 = flow_now_ns() @@ -475,6 +530,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_complement_nb: ptr = complement_nb_predict(fitted_complement_nb, X_c) + sink = sink + o_complement_nb[0] array_free_f32(o_complement_nb) } t3 = flow_now_ns() @@ -499,31 +555,6 @@ function main() -> i32 { printf("ESTIMATOR|dbscan|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) fflush(null) - # ---- decision_tree_classifier (classification) ---- - t0 = flow_now_ns() - let probe_decision_tree_classifier: DecisionTreeClassifier = decision_tree_classifier_fit(X_c, y_c, 3, 5, 0) - t1 = flow_now_ns() - decision_tree_classifier_free(probe_decision_tree_classifier) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_decision_tree_classifier: DecisionTreeClassifier = decision_tree_classifier_fit(X_c, y_c, 3, 5, 0) - decision_tree_classifier_free(m_decision_tree_classifier) - } - t1 = flow_now_ns() - let fitted_decision_tree_classifier: DecisionTreeClassifier = decision_tree_classifier_fit(X_c, y_c, 3, 5, 0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_decision_tree_classifier: ptr = decision_tree_classifier_predict(fitted_decision_tree_classifier, X_c) - array_free_f32(o_decision_tree_classifier) - } - t3 = flow_now_ns() - decision_tree_classifier_free(fitted_decision_tree_classifier) - printf("ESTIMATOR|decision_tree_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) @@ -532,5 +563,10 @@ function main() -> i32 { matrix_free(Y_multi) free(yi_c as ptr) free(yi_r as ptr) + array_free_f32(x1d_r) + array_free_f32(w_f) + # The sink is printed so the work above cannot be optimized away. The + # parser matches ESTIMATOR lines only, so this one is ignored. + printf("SINK|%.9f\n", sink) return 0 } diff --git a/benchmarks/generated/bench_estimators_01.flow b/benchmarks/generated/bench_estimators_01.flow index 4101be8..dcf0255 100644 --- a/benchmarks/generated/bench_estimators_01.flow +++ b/benchmarks/generated/bench_estimators_01.flow @@ -65,11 +65,48 @@ function main() -> i32 { Y_rows[i] = row } + # A one-dimensional x for the isotonic row, which regresses against a + # single ordered variable rather than a design. + let x1d_r: ptr = array_new_f32(n_r) + for i in 0 to n_r { x1d_r[i] = matrix_at(X_r, i, 0) } + + # Per-feature importances for the selector row, which takes the weights a + # fitted model would hand it rather than a design. + let w_f: ptr = array_new_f32(f_c) + for i in 0 to f_c { w_f[i] = 1.0 / ((i + 1) as f32) } + let mut t0: i64 = 0 let mut t1: i64 = 0 let mut t2: i64 = 0 let mut t3: i64 = 0 let mut reps: i32 = 1 + let mut sink: f32 = 0.0 + + # ---- decision_tree_classifier (classification) ---- + t0 = flow_now_ns() + let probe_decision_tree_classifier: DecisionTreeClassifier = decision_tree_classifier_fit(X_c, y_c, 3, 5, 0) + t1 = flow_now_ns() + decision_tree_classifier_free(probe_decision_tree_classifier) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_decision_tree_classifier: DecisionTreeClassifier = decision_tree_classifier_fit(X_c, y_c, 3, 5, 0) + decision_tree_classifier_free(m_decision_tree_classifier) + } + t1 = flow_now_ns() + let fitted_decision_tree_classifier: DecisionTreeClassifier = decision_tree_classifier_fit(X_c, y_c, 3, 5, 0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_decision_tree_classifier: ptr = decision_tree_classifier_predict(fitted_decision_tree_classifier, X_c) + sink = sink + o_decision_tree_classifier[0] + array_free_f32(o_decision_tree_classifier) + } + t3 = flow_now_ns() + decision_tree_classifier_free(fitted_decision_tree_classifier) + printf("ESTIMATOR|decision_tree_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) # ---- decision_tree_regressor (regression) ---- t0 = flow_now_ns() @@ -89,6 +126,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_decision_tree_regressor: ptr = decision_tree_regressor_predict(fitted_decision_tree_regressor, X_r) + sink = sink + o_decision_tree_regressor[0] array_free_f32(o_decision_tree_regressor) } t3 = flow_now_ns() @@ -114,6 +152,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_dictionary_learning: Matrix = dictionary_learning_transform(fitted_dictionary_learning, X_c) + if o_dictionary_learning.rows > 0 { + if o_dictionary_learning.cols > 0 { sink = sink + o_dictionary_learning.data[0] } + } matrix_free(o_dictionary_learning) } t3 = flow_now_ns() @@ -139,6 +180,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_discriminant_lda: ptr = discriminant_lda_predict(fitted_discriminant_lda, X_c) + sink = sink + o_discriminant_lda[0] array_free_f32(o_discriminant_lda) } t3 = flow_now_ns() @@ -146,6 +188,58 @@ function main() -> i32 { printf("ESTIMATOR|discriminant_lda|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) + # ---- dummy_classifier (classification, written out) ---- + t0 = flow_now_ns() + let probe_dummy_classifier: DummyClassifier = dummy_classifier_fit(y_c, n_c, 3, 0, 0.0, 42) + t1 = flow_now_ns() + dummy_classifier_free(probe_dummy_classifier) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_dummy_classifier: DummyClassifier = dummy_classifier_fit(y_c, n_c, 3, 0, 0.0, 42) + dummy_classifier_free(m_dummy_classifier) + } + t1 = flow_now_ns() + let fitted_dummy_classifier: DummyClassifier = dummy_classifier_fit(y_c, n_c, 3, 0, 0.0, 42) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_dummy_classifier: ptr = dummy_classifier_predict(fitted_dummy_classifier, n_c) + sink = sink + o_dummy_classifier[0] + array_free_f32(o_dummy_classifier) + } + t3 = flow_now_ns() + dummy_classifier_free(fitted_dummy_classifier) + printf("ESTIMATOR|dummy_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- dummy_regressor (regression, written out) ---- + t0 = flow_now_ns() + let probe_dummy_regressor: DummyRegressor = dummy_regressor_fit(y_r, n_r, 0, 0.0) + t1 = flow_now_ns() + dummy_regressor_free(probe_dummy_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_dummy_regressor: DummyRegressor = dummy_regressor_fit(y_r, n_r, 0, 0.0) + dummy_regressor_free(m_dummy_regressor) + } + t1 = flow_now_ns() + let fitted_dummy_regressor: DummyRegressor = dummy_regressor_fit(y_r, n_r, 0, 0.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_dummy_regressor: ptr = dummy_regressor_predict(fitted_dummy_regressor, n_r) + sink = sink + o_dummy_regressor[0] + array_free_f32(o_dummy_regressor) + } + t3 = flow_now_ns() + dummy_regressor_free(fitted_dummy_regressor) + printf("ESTIMATOR|dummy_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- elastic_net_cv (regression) ---- t0 = flow_now_ns() let probe_elastic_net_cv: ElasticNetCV = elastic_net_cv_fit(X_r, y_r, 10) @@ -164,6 +258,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_elastic_net_cv: ptr = elastic_net_cv_predict(fitted_elastic_net_cv, X_r) + sink = sink + o_elastic_net_cv[0] array_free_f32(o_elastic_net_cv) } t3 = flow_now_ns() @@ -189,6 +284,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_elastic_net: ptr = elastic_net_predict(fitted_elastic_net, X_r) + sink = sink + o_elastic_net[0] array_free_f32(o_elastic_net) } t3 = flow_now_ns() @@ -214,6 +310,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_elliptic_envelope: ptr = elliptic_envelope_predict(fitted_elliptic_envelope, X_c) + sink = sink + o_elliptic_envelope[0] array_free_f32(o_elliptic_envelope) } t3 = flow_now_ns() @@ -256,6 +353,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_extra_tree_classifier: ptr = extra_tree_classifier_predict(fitted_extra_tree_classifier, X_c) + sink = sink + o_extra_tree_classifier[0] array_free_f32(o_extra_tree_classifier) } t3 = flow_now_ns() @@ -281,6 +379,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_extra_tree_regressor: ptr = extra_tree_regressor_predict(fitted_extra_tree_regressor, X_r) + sink = sink + o_extra_tree_regressor[0] array_free_f32(o_extra_tree_regressor) } t3 = flow_now_ns() @@ -306,6 +405,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_extra_trees_classifier: ptr = extra_trees_classifier_predict(fitted_extra_trees_classifier, X_c) + sink = sink + o_extra_trees_classifier[0] array_free_f32(o_extra_trees_classifier) } t3 = flow_now_ns() @@ -331,6 +431,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_extra_trees_regressor: ptr = extra_trees_regressor_predict(fitted_extra_trees_regressor, X_r) + sink = sink + o_extra_trees_regressor[0] array_free_f32(o_extra_trees_regressor) } t3 = flow_now_ns() @@ -356,6 +457,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_factor_analysis: Matrix = factor_analysis_transform(fitted_factor_analysis, X_c) + if o_factor_analysis.rows > 0 { + if o_factor_analysis.cols > 0 { sink = sink + o_factor_analysis.data[0] } + } matrix_free(o_factor_analysis) } t3 = flow_now_ns() @@ -381,6 +485,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_fast_ica: Matrix = fast_ica_transform(fitted_fast_ica, X_c) + if o_fast_ica.rows > 0 { + if o_fast_ica.cols > 0 { sink = sink + o_fast_ica.data[0] } + } matrix_free(o_fast_ica) } t3 = flow_now_ns() @@ -406,6 +513,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_feature_agglomeration: Matrix = feature_agglomeration_transform(fitted_feature_agglomeration, X_c) + if o_feature_agglomeration.rows > 0 { + if o_feature_agglomeration.cols > 0 { sink = sink + o_feature_agglomeration.data[0] } + } matrix_free(o_feature_agglomeration) } t3 = flow_now_ns() @@ -431,6 +541,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_gamma_regressor: ptr = gamma_regressor_predict(fitted_gamma_regressor, X_r) + sink = sink + o_gamma_regressor[0] array_free_f32(o_gamma_regressor) } t3 = flow_now_ns() @@ -473,6 +584,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_gaussian_nb: ptr = gaussian_nb_predict(fitted_gaussian_nb, X_c) + sink = sink + o_gaussian_nb[0] array_free_f32(o_gaussian_nb) } t3 = flow_now_ns() @@ -480,81 +592,6 @@ function main() -> i32 { printf("ESTIMATOR|gaussian_nb|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- gaussian_process_classifier (regression) ---- - t0 = flow_now_ns() - let probe_gaussian_process_classifier: GaussianProcessClassifier = gaussian_process_classifier_fit(X_r, y_r, 0.1, 100) - t1 = flow_now_ns() - gaussian_process_classifier_free(probe_gaussian_process_classifier) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_gaussian_process_classifier: GaussianProcessClassifier = gaussian_process_classifier_fit(X_r, y_r, 0.1, 100) - gaussian_process_classifier_free(m_gaussian_process_classifier) - } - t1 = flow_now_ns() - let fitted_gaussian_process_classifier: GaussianProcessClassifier = gaussian_process_classifier_fit(X_r, y_r, 0.1, 100) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_gaussian_process_classifier: ptr = gaussian_process_classifier_predict(fitted_gaussian_process_classifier, X_r) - array_free_f32(o_gaussian_process_classifier) - } - t3 = flow_now_ns() - gaussian_process_classifier_free(fitted_gaussian_process_classifier) - printf("ESTIMATOR|gaussian_process_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- gaussian_process_regressor (regression) ---- - t0 = flow_now_ns() - let probe_gaussian_process_regressor: GaussianProcessRegressor = gaussian_process_regressor_fit(X_r, y_r, 1.0, 1.0, 0) - t1 = flow_now_ns() - gaussian_process_regressor_free(probe_gaussian_process_regressor) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_gaussian_process_regressor: GaussianProcessRegressor = gaussian_process_regressor_fit(X_r, y_r, 1.0, 1.0, 0) - gaussian_process_regressor_free(m_gaussian_process_regressor) - } - t1 = flow_now_ns() - let fitted_gaussian_process_regressor: GaussianProcessRegressor = gaussian_process_regressor_fit(X_r, y_r, 1.0, 1.0, 0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_gaussian_process_regressor: ptr = gaussian_process_regressor_predict(fitted_gaussian_process_regressor, X_r) - array_free_f32(o_gaussian_process_regressor) - } - t3 = flow_now_ns() - gaussian_process_regressor_free(fitted_gaussian_process_regressor) - printf("ESTIMATOR|gaussian_process_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- gradient_boosting_classifier (classification) ---- - t0 = flow_now_ns() - let probe_gradient_boosting_classifier: GradientBoostingClassifier = gradient_boosting_classifier_fit(X_c, y_c, 3, 10, 0.1, 5, 42) - t1 = flow_now_ns() - gradient_boosting_classifier_free(probe_gradient_boosting_classifier) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_gradient_boosting_classifier: GradientBoostingClassifier = gradient_boosting_classifier_fit(X_c, y_c, 3, 10, 0.1, 5, 42) - gradient_boosting_classifier_free(m_gradient_boosting_classifier) - } - t1 = flow_now_ns() - let fitted_gradient_boosting_classifier: GradientBoostingClassifier = gradient_boosting_classifier_fit(X_c, y_c, 3, 10, 0.1, 5, 42) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_gradient_boosting_classifier: ptr = gradient_boosting_classifier_predict(fitted_gradient_boosting_classifier, X_c) - array_free_f32(o_gradient_boosting_classifier) - } - t3 = flow_now_ns() - gradient_boosting_classifier_free(fitted_gradient_boosting_classifier) - printf("ESTIMATOR|gradient_boosting_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) @@ -563,5 +600,10 @@ function main() -> i32 { matrix_free(Y_multi) free(yi_c as ptr) free(yi_r as ptr) + array_free_f32(x1d_r) + array_free_f32(w_f) + # The sink is printed so the work above cannot be optimized away. The + # parser matches ESTIMATOR lines only, so this one is ignored. + printf("SINK|%.9f\n", sink) return 0 } diff --git a/benchmarks/generated/bench_estimators_02.flow b/benchmarks/generated/bench_estimators_02.flow index 79de920..5d45b94 100644 --- a/benchmarks/generated/bench_estimators_02.flow +++ b/benchmarks/generated/bench_estimators_02.flow @@ -65,11 +65,128 @@ function main() -> i32 { Y_rows[i] = row } + # A one-dimensional x for the isotonic row, which regresses against a + # single ordered variable rather than a design. + let x1d_r: ptr = array_new_f32(n_r) + for i in 0 to n_r { x1d_r[i] = matrix_at(X_r, i, 0) } + + # Per-feature importances for the selector row, which takes the weights a + # fitted model would hand it rather than a design. + let w_f: ptr = array_new_f32(f_c) + for i in 0 to f_c { w_f[i] = 1.0 / ((i + 1) as f32) } + let mut t0: i64 = 0 let mut t1: i64 = 0 let mut t2: i64 = 0 let mut t3: i64 = 0 let mut reps: i32 = 1 + let mut sink: f32 = 0.0 + + # ---- gaussian_process_classifier (regression) ---- + t0 = flow_now_ns() + let probe_gaussian_process_classifier: GaussianProcessClassifier = gaussian_process_classifier_fit(X_r, y_r, 0.1, 100) + t1 = flow_now_ns() + gaussian_process_classifier_free(probe_gaussian_process_classifier) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_gaussian_process_classifier: GaussianProcessClassifier = gaussian_process_classifier_fit(X_r, y_r, 0.1, 100) + gaussian_process_classifier_free(m_gaussian_process_classifier) + } + t1 = flow_now_ns() + let fitted_gaussian_process_classifier: GaussianProcessClassifier = gaussian_process_classifier_fit(X_r, y_r, 0.1, 100) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_gaussian_process_classifier: ptr = gaussian_process_classifier_predict(fitted_gaussian_process_classifier, X_r) + sink = sink + o_gaussian_process_classifier[0] + array_free_f32(o_gaussian_process_classifier) + } + t3 = flow_now_ns() + gaussian_process_classifier_free(fitted_gaussian_process_classifier) + printf("ESTIMATOR|gaussian_process_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- gaussian_process_regressor (regression) ---- + t0 = flow_now_ns() + let probe_gaussian_process_regressor: GaussianProcessRegressor = gaussian_process_regressor_fit(X_r, y_r, 1.0, 1.0, 0) + t1 = flow_now_ns() + gaussian_process_regressor_free(probe_gaussian_process_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_gaussian_process_regressor: GaussianProcessRegressor = gaussian_process_regressor_fit(X_r, y_r, 1.0, 1.0, 0) + gaussian_process_regressor_free(m_gaussian_process_regressor) + } + t1 = flow_now_ns() + let fitted_gaussian_process_regressor: GaussianProcessRegressor = gaussian_process_regressor_fit(X_r, y_r, 1.0, 1.0, 0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_gaussian_process_regressor: ptr = gaussian_process_regressor_predict(fitted_gaussian_process_regressor, X_r) + sink = sink + o_gaussian_process_regressor[0] + array_free_f32(o_gaussian_process_regressor) + } + t3 = flow_now_ns() + gaussian_process_regressor_free(fitted_gaussian_process_regressor) + printf("ESTIMATOR|gaussian_process_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- gaussian_random_projection (classification, written out) ---- + t0 = flow_now_ns() + let probe_gaussian_random_projection: GaussianRandomProjection = gaussian_random_projection_fit(f_c, 2, 42) + t1 = flow_now_ns() + gaussian_random_projection_free(probe_gaussian_random_projection) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_gaussian_random_projection: GaussianRandomProjection = gaussian_random_projection_fit(f_c, 2, 42) + gaussian_random_projection_free(m_gaussian_random_projection) + } + t1 = flow_now_ns() + let fitted_gaussian_random_projection: GaussianRandomProjection = gaussian_random_projection_fit(f_c, 2, 42) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_gaussian_random_projection: Matrix = gaussian_random_projection_transform(fitted_gaussian_random_projection, X_c) + if o_gaussian_random_projection.rows > 0 { + if o_gaussian_random_projection.cols > 0 { sink = sink + o_gaussian_random_projection.data[0] } + } + matrix_free(o_gaussian_random_projection) + } + t3 = flow_now_ns() + gaussian_random_projection_free(fitted_gaussian_random_projection) + printf("ESTIMATOR|gaussian_random_projection|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- gradient_boosting_classifier (classification) ---- + t0 = flow_now_ns() + let probe_gradient_boosting_classifier: GradientBoostingClassifier = gradient_boosting_classifier_fit(X_c, y_c, 3, 10, 0.1, 5, 42) + t1 = flow_now_ns() + gradient_boosting_classifier_free(probe_gradient_boosting_classifier) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_gradient_boosting_classifier: GradientBoostingClassifier = gradient_boosting_classifier_fit(X_c, y_c, 3, 10, 0.1, 5, 42) + gradient_boosting_classifier_free(m_gradient_boosting_classifier) + } + t1 = flow_now_ns() + let fitted_gradient_boosting_classifier: GradientBoostingClassifier = gradient_boosting_classifier_fit(X_c, y_c, 3, 10, 0.1, 5, 42) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_gradient_boosting_classifier: ptr = gradient_boosting_classifier_predict(fitted_gradient_boosting_classifier, X_c) + sink = sink + o_gradient_boosting_classifier[0] + array_free_f32(o_gradient_boosting_classifier) + } + t3 = flow_now_ns() + gradient_boosting_classifier_free(fitted_gradient_boosting_classifier) + printf("ESTIMATOR|gradient_boosting_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) # ---- gradient_boosting_regressor (regression) ---- t0 = flow_now_ns() @@ -89,6 +206,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_gradient_boosting_regressor: ptr = gradient_boosting_regressor_predict(fitted_gradient_boosting_regressor, X_r) + sink = sink + o_gradient_boosting_regressor[0] array_free_f32(o_gradient_boosting_regressor) } t3 = flow_now_ns() @@ -148,6 +266,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_hist_gradient_boosting_classifier: ptr = hist_gradient_boosting_classifier_predict(fitted_hist_gradient_boosting_classifier, X_c) + sink = sink + o_hist_gradient_boosting_classifier[0] array_free_f32(o_hist_gradient_boosting_classifier) } t3 = flow_now_ns() @@ -173,6 +292,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_hist_gradient_boosting_regressor: ptr = hist_gradient_boosting_regressor_predict(fitted_hist_gradient_boosting_regressor, X_r) + sink = sink + o_hist_gradient_boosting_regressor[0] array_free_f32(o_hist_gradient_boosting_regressor) } t3 = flow_now_ns() @@ -198,6 +318,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_huber_regressor: ptr = huber_regressor_predict(fitted_huber_regressor, X_r) + sink = sink + o_huber_regressor[0] array_free_f32(o_huber_regressor) } t3 = flow_now_ns() @@ -239,6 +360,32 @@ function main() -> i32 { printf("ESTIMATOR|isomap|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) fflush(null) + # ---- isotonic (regression, written out) ---- + t0 = flow_now_ns() + let probe_isotonic: IsotonicRegression = isotonic_fit(x1d_r, y_r, n_r, true) + t1 = flow_now_ns() + isotonic_free(probe_isotonic) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_isotonic: IsotonicRegression = isotonic_fit(x1d_r, y_r, n_r, true) + isotonic_free(m_isotonic) + } + t1 = flow_now_ns() + let fitted_isotonic: IsotonicRegression = isotonic_fit(x1d_r, y_r, n_r, true) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_isotonic: ptr = isotonic_transform(fitted_isotonic, x1d_r, n_r) + sink = sink + o_isotonic[0] + array_free_f32(o_isotonic) + } + t3 = flow_now_ns() + isotonic_free(fitted_isotonic) + printf("ESTIMATOR|isotonic|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- iterative_imputer (unsupervised) ---- t0 = flow_now_ns() let probe_iterative_imputer: IterativeImputer = iterative_imputer_fit(X_c, 100, 0.0001, 42) @@ -257,6 +404,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_iterative_imputer: Matrix = iterative_imputer_transform(fitted_iterative_imputer, X_c) + if o_iterative_imputer.rows > 0 { + if o_iterative_imputer.cols > 0 { sink = sink + o_iterative_imputer.data[0] } + } matrix_free(o_iterative_imputer) } t3 = flow_now_ns() @@ -282,6 +432,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_kbins_discretizer: Matrix = kbins_discretizer_transform(fitted_kbins_discretizer, X_c) + if o_kbins_discretizer.rows > 0 { + if o_kbins_discretizer.cols > 0 { sink = sink + o_kbins_discretizer.data[0] } + } matrix_free(o_kbins_discretizer) } t3 = flow_now_ns() @@ -324,6 +477,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_kernel_pca: Matrix = kernel_pca_transform(fitted_kernel_pca, X_c) + if o_kernel_pca.rows > 0 { + if o_kernel_pca.cols > 0 { sink = sink + o_kernel_pca.data[0] } + } matrix_free(o_kernel_pca) } t3 = flow_now_ns() @@ -349,6 +505,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_kernel_ridge: ptr = kernel_ridge_predict(fitted_kernel_ridge, X_r) + sink = sink + o_kernel_ridge[0] array_free_f32(o_kernel_ridge) } t3 = flow_now_ns() @@ -374,6 +531,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_kernel_svc: ptr = kernel_svc_predict(fitted_kernel_svc, X_c) + sink = sink + o_kernel_svc[0] array_free_f32(o_kernel_svc) } t3 = flow_now_ns() @@ -399,6 +557,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_kernel_svc_multi: ptr = kernel_svc_multi_predict(fitted_kernel_svc_multi, X_c) + sink = sink + o_kernel_svc_multi[0] array_free_f32(o_kernel_svc_multi) } t3 = flow_now_ns() @@ -406,123 +565,6 @@ function main() -> i32 { printf("ESTIMATOR|kernel_svc_multi|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- kmeans (unsupervised) ---- - t0 = flow_now_ns() - let probe_kmeans: KMeans = kmeans_fit(X_c, 3, 100, 0.0001, 42) - t1 = flow_now_ns() - kmeans_free(probe_kmeans) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_kmeans: KMeans = kmeans_fit(X_c, 3, 100, 0.0001, 42) - kmeans_free(m_kmeans) - } - t1 = flow_now_ns() - printf("ESTIMATOR|kmeans|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) - fflush(null) - - # ---- kneighbors_transformer (unsupervised) ---- - t0 = flow_now_ns() - let probe_kneighbors_transformer: KNeighborsTransformer = kneighbors_transformer_fit(X_c, 5, 0) - t1 = flow_now_ns() - kneighbors_transformer_free(probe_kneighbors_transformer) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_kneighbors_transformer: KNeighborsTransformer = kneighbors_transformer_fit(X_c, 5, 0) - kneighbors_transformer_free(m_kneighbors_transformer) - } - t1 = flow_now_ns() - let fitted_kneighbors_transformer: KNeighborsTransformer = kneighbors_transformer_fit(X_c, 5, 0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_kneighbors_transformer: Matrix = kneighbors_transformer_transform(fitted_kneighbors_transformer, X_c) - matrix_free(o_kneighbors_transformer) - } - t3 = flow_now_ns() - kneighbors_transformer_free(fitted_kneighbors_transformer) - printf("ESTIMATOR|kneighbors_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- knn_classifier (classification) ---- - t0 = flow_now_ns() - let probe_knn_classifier: KNNClassifier = knn_classifier_fit(X_c, y_c, 3, 3) - t1 = flow_now_ns() - knn_classifier_free(probe_knn_classifier) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_knn_classifier: KNNClassifier = knn_classifier_fit(X_c, y_c, 3, 3) - knn_classifier_free(m_knn_classifier) - } - t1 = flow_now_ns() - let fitted_knn_classifier: KNNClassifier = knn_classifier_fit(X_c, y_c, 3, 3) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_knn_classifier: Matrix = knn_classifier_predict(fitted_knn_classifier, X_c) - matrix_free(o_knn_classifier) - } - t3 = flow_now_ns() - knn_classifier_free(fitted_knn_classifier) - printf("ESTIMATOR|knn_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- knn_imputer (unsupervised) ---- - t0 = flow_now_ns() - let probe_knn_imputer: KNNImputer = knn_imputer_fit(X_c, 5, 0) - t1 = flow_now_ns() - knn_imputer_free(probe_knn_imputer) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_knn_imputer: KNNImputer = knn_imputer_fit(X_c, 5, 0) - knn_imputer_free(m_knn_imputer) - } - t1 = flow_now_ns() - let fitted_knn_imputer: KNNImputer = knn_imputer_fit(X_c, 5, 0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_knn_imputer: Matrix = knn_imputer_transform(fitted_knn_imputer, X_c) - matrix_free(o_knn_imputer) - } - t3 = flow_now_ns() - knn_imputer_free(fitted_knn_imputer) - printf("ESTIMATOR|knn_imputer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- knn_regressor (regression) ---- - t0 = flow_now_ns() - let probe_knn_regressor: KNNRegressor = knn_regressor_fit(X_r, y_r, 3) - t1 = flow_now_ns() - knn_regressor_free(probe_knn_regressor) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_knn_regressor: KNNRegressor = knn_regressor_fit(X_r, y_r, 3) - knn_regressor_free(m_knn_regressor) - } - t1 = flow_now_ns() - let fitted_knn_regressor: KNNRegressor = knn_regressor_fit(X_r, y_r, 3) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_knn_regressor: ptr = knn_regressor_predict(fitted_knn_regressor, X_r) - array_free_f32(o_knn_regressor) - } - t3 = flow_now_ns() - knn_regressor_free(fitted_knn_regressor) - printf("ESTIMATOR|knn_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) @@ -531,5 +573,10 @@ function main() -> i32 { matrix_free(Y_multi) free(yi_c as ptr) free(yi_r as ptr) + array_free_f32(x1d_r) + array_free_f32(w_f) + # The sink is printed so the work above cannot be optimized away. The + # parser matches ESTIMATOR lines only, so this one is ignored. + printf("SINK|%.9f\n", sink) return 0 } diff --git a/benchmarks/generated/bench_estimators_03.flow b/benchmarks/generated/bench_estimators_03.flow index 542372f..3f797ba 100644 --- a/benchmarks/generated/bench_estimators_03.flow +++ b/benchmarks/generated/bench_estimators_03.flow @@ -65,11 +65,203 @@ function main() -> i32 { Y_rows[i] = row } + # A one-dimensional x for the isotonic row, which regresses against a + # single ordered variable rather than a design. + let x1d_r: ptr = array_new_f32(n_r) + for i in 0 to n_r { x1d_r[i] = matrix_at(X_r, i, 0) } + + # Per-feature importances for the selector row, which takes the weights a + # fitted model would hand it rather than a design. + let w_f: ptr = array_new_f32(f_c) + for i in 0 to f_c { w_f[i] = 1.0 / ((i + 1) as f32) } + let mut t0: i64 = 0 let mut t1: i64 = 0 let mut t2: i64 = 0 let mut t3: i64 = 0 let mut reps: i32 = 1 + let mut sink: f32 = 0.0 + + # ---- kmeans (unsupervised) ---- + t0 = flow_now_ns() + let probe_kmeans: KMeans = kmeans_fit(X_c, 3, 100, 0.0001, 42) + t1 = flow_now_ns() + kmeans_free(probe_kmeans) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_kmeans: KMeans = kmeans_fit(X_c, 3, 100, 0.0001, 42) + kmeans_free(m_kmeans) + } + t1 = flow_now_ns() + printf("ESTIMATOR|kmeans|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + + # ---- kneighbors_transformer (unsupervised) ---- + t0 = flow_now_ns() + let probe_kneighbors_transformer: KNeighborsTransformer = kneighbors_transformer_fit(X_c, 5, 0) + t1 = flow_now_ns() + kneighbors_transformer_free(probe_kneighbors_transformer) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_kneighbors_transformer: KNeighborsTransformer = kneighbors_transformer_fit(X_c, 5, 0) + kneighbors_transformer_free(m_kneighbors_transformer) + } + t1 = flow_now_ns() + let fitted_kneighbors_transformer: KNeighborsTransformer = kneighbors_transformer_fit(X_c, 5, 0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_kneighbors_transformer: Matrix = kneighbors_transformer_transform(fitted_kneighbors_transformer, X_c) + if o_kneighbors_transformer.rows > 0 { + if o_kneighbors_transformer.cols > 0 { sink = sink + o_kneighbors_transformer.data[0] } + } + matrix_free(o_kneighbors_transformer) + } + t3 = flow_now_ns() + kneighbors_transformer_free(fitted_kneighbors_transformer) + printf("ESTIMATOR|kneighbors_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- knn_classifier (classification) ---- + t0 = flow_now_ns() + let probe_knn_classifier: KNNClassifier = knn_classifier_fit(X_c, y_c, 3, 3) + t1 = flow_now_ns() + knn_classifier_free(probe_knn_classifier) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_knn_classifier: KNNClassifier = knn_classifier_fit(X_c, y_c, 3, 3) + knn_classifier_free(m_knn_classifier) + } + t1 = flow_now_ns() + let fitted_knn_classifier: KNNClassifier = knn_classifier_fit(X_c, y_c, 3, 3) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_knn_classifier: Matrix = knn_classifier_predict(fitted_knn_classifier, X_c) + if o_knn_classifier.rows > 0 { + if o_knn_classifier.cols > 0 { sink = sink + o_knn_classifier.data[0] } + } + matrix_free(o_knn_classifier) + } + t3 = flow_now_ns() + knn_classifier_free(fitted_knn_classifier) + printf("ESTIMATOR|knn_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- knn_imputer (unsupervised) ---- + t0 = flow_now_ns() + let probe_knn_imputer: KNNImputer = knn_imputer_fit(X_c, 5, 0) + t1 = flow_now_ns() + knn_imputer_free(probe_knn_imputer) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_knn_imputer: KNNImputer = knn_imputer_fit(X_c, 5, 0) + knn_imputer_free(m_knn_imputer) + } + t1 = flow_now_ns() + let fitted_knn_imputer: KNNImputer = knn_imputer_fit(X_c, 5, 0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_knn_imputer: Matrix = knn_imputer_transform(fitted_knn_imputer, X_c) + if o_knn_imputer.rows > 0 { + if o_knn_imputer.cols > 0 { sink = sink + o_knn_imputer.data[0] } + } + matrix_free(o_knn_imputer) + } + t3 = flow_now_ns() + knn_imputer_free(fitted_knn_imputer) + printf("ESTIMATOR|knn_imputer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- knn_regressor (regression) ---- + t0 = flow_now_ns() + let probe_knn_regressor: KNNRegressor = knn_regressor_fit(X_r, y_r, 3) + t1 = flow_now_ns() + knn_regressor_free(probe_knn_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_knn_regressor: KNNRegressor = knn_regressor_fit(X_r, y_r, 3) + knn_regressor_free(m_knn_regressor) + } + t1 = flow_now_ns() + let fitted_knn_regressor: KNNRegressor = knn_regressor_fit(X_r, y_r, 3) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_knn_regressor: ptr = knn_regressor_predict(fitted_knn_regressor, X_r) + sink = sink + o_knn_regressor[0] + array_free_f32(o_knn_regressor) + } + t3 = flow_now_ns() + knn_regressor_free(fitted_knn_regressor) + printf("ESTIMATOR|knn_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- label_binarizer (classification, written out) ---- + t0 = flow_now_ns() + let probe_label_binarizer: LabelBinarizer = label_binarizer_fit(y_c, n_c, 0.0, 1.0) + t1 = flow_now_ns() + label_binarizer_free(probe_label_binarizer) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_label_binarizer: LabelBinarizer = label_binarizer_fit(y_c, n_c, 0.0, 1.0) + label_binarizer_free(m_label_binarizer) + } + t1 = flow_now_ns() + let fitted_label_binarizer: LabelBinarizer = label_binarizer_fit(y_c, n_c, 0.0, 1.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_label_binarizer: Matrix = label_binarizer_transform(fitted_label_binarizer, y_c, n_c) + if o_label_binarizer.rows > 0 { + if o_label_binarizer.cols > 0 { sink = sink + o_label_binarizer.data[0] } + } + matrix_free(o_label_binarizer) + } + t3 = flow_now_ns() + label_binarizer_free(fitted_label_binarizer) + printf("ESTIMATOR|label_binarizer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- label_encoder (classification, written out) ---- + t0 = flow_now_ns() + let probe_label_encoder: LabelEncoder = label_encoder_fit(y_c, n_c) + t1 = flow_now_ns() + label_encoder_free(probe_label_encoder) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_label_encoder: LabelEncoder = label_encoder_fit(y_c, n_c) + label_encoder_free(m_label_encoder) + } + t1 = flow_now_ns() + let fitted_label_encoder: LabelEncoder = label_encoder_fit(y_c, n_c) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_label_encoder: ptr = label_encoder_transform(fitted_label_encoder, y_c, n_c) + sink = sink + o_label_encoder[0] + array_free_f32(o_label_encoder) + } + t3 = flow_now_ns() + label_encoder_free(fitted_label_encoder) + printf("ESTIMATOR|label_encoder|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) # ---- label_propagation (classification) ---- t0 = flow_now_ns() @@ -123,6 +315,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_lars_cv: ptr = lars_cv_predict(fitted_lars_cv, X_r) + sink = sink + o_lars_cv[0] array_free_f32(o_lars_cv) } t3 = flow_now_ns() @@ -148,6 +341,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_lars: ptr = lars_predict(fitted_lars, X_r) + sink = sink + o_lars[0] array_free_f32(o_lars) } t3 = flow_now_ns() @@ -173,6 +367,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_lasso_cv: ptr = lasso_cv_predict(fitted_lasso_cv, X_r) + sink = sink + o_lasso_cv[0] array_free_f32(o_lasso_cv) } t3 = flow_now_ns() @@ -198,6 +393,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_lasso: ptr = lasso_predict(fitted_lasso, X_r) + sink = sink + o_lasso[0] array_free_f32(o_lasso) } t3 = flow_now_ns() @@ -223,6 +419,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_lasso_lars_cv: ptr = lasso_lars_cv_predict(fitted_lasso_lars_cv, X_r) + sink = sink + o_lasso_lars_cv[0] array_free_f32(o_lasso_lars_cv) } t3 = flow_now_ns() @@ -248,6 +445,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_lasso_lars: ptr = lasso_lars_predict(fitted_lasso_lars, X_r) + sink = sink + o_lasso_lars[0] array_free_f32(o_lasso_lars) } t3 = flow_now_ns() @@ -273,6 +471,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_lasso_lars_ic: ptr = lasso_lars_ic_predict(fitted_lasso_lars_ic, X_r) + sink = sink + o_lasso_lars_ic[0] array_free_f32(o_lasso_lars_ic) } t3 = flow_now_ns() @@ -298,6 +497,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_lda: Matrix = lda_transform(fitted_lda, X_c) + if o_lda.rows > 0 { + if o_lda.cols > 0 { sink = sink + o_lda.data[0] } + } matrix_free(o_lda) } t3 = flow_now_ns() @@ -338,6 +540,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_linear_regression: ptr = linear_regression_predict(fitted_linear_regression, X_r) + sink = sink + o_linear_regression[0] array_free_f32(o_linear_regression) } t3 = flow_now_ns() @@ -363,6 +566,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_linear_svc: ptr = linear_svc_predict(fitted_linear_svc, X_c) + sink = sink + o_linear_svc[0] array_free_f32(o_linear_svc) } t3 = flow_now_ns() @@ -370,157 +574,6 @@ function main() -> i32 { printf("ESTIMATOR|linear_svc|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- linear_svc_multi (classification) ---- - t0 = flow_now_ns() - let probe_linear_svc_multi: LinearSVCMulti = linear_svc_multi_fit(X_c, y_c, 3, 1.0, 50) - t1 = flow_now_ns() - linear_svc_multi_free(probe_linear_svc_multi) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_linear_svc_multi: LinearSVCMulti = linear_svc_multi_fit(X_c, y_c, 3, 1.0, 50) - linear_svc_multi_free(m_linear_svc_multi) - } - t1 = flow_now_ns() - let fitted_linear_svc_multi: LinearSVCMulti = linear_svc_multi_fit(X_c, y_c, 3, 1.0, 50) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_linear_svc_multi: ptr = linear_svc_multi_predict(fitted_linear_svc_multi, X_c) - array_free_f32(o_linear_svc_multi) - } - t3 = flow_now_ns() - linear_svc_multi_free(fitted_linear_svc_multi) - printf("ESTIMATOR|linear_svc_multi|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- linear_svr (regression) ---- - t0 = flow_now_ns() - let probe_linear_svr: LinearSVR = linear_svr_fit(X_r, y_r, 1.0, 0.1, 50, 0.01) - t1 = flow_now_ns() - linear_svr_free(probe_linear_svr) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_linear_svr: LinearSVR = linear_svr_fit(X_r, y_r, 1.0, 0.1, 50, 0.01) - linear_svr_free(m_linear_svr) - } - t1 = flow_now_ns() - let fitted_linear_svr: LinearSVR = linear_svr_fit(X_r, y_r, 1.0, 0.1, 50, 0.01) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_linear_svr: ptr = linear_svr_predict(fitted_linear_svr, X_r) - array_free_f32(o_linear_svr) - } - t3 = flow_now_ns() - linear_svr_free(fitted_linear_svr) - printf("ESTIMATOR|linear_svr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- lle (unsupervised) ---- - t0 = flow_now_ns() - let probe_lle: LLE = lle_fit(X_c, 2, 5) - t1 = flow_now_ns() - lle_free(probe_lle) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_lle: LLE = lle_fit(X_c, 2, 5) - lle_free(m_lle) - } - t1 = flow_now_ns() - printf("ESTIMATOR|lle|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) - fflush(null) - - # ---- local_outlier_factor (unsupervised) ---- - t0 = flow_now_ns() - let probe_local_outlier_factor: LocalOutlierFactor = local_outlier_factor_fit(X_c, 5) - t1 = flow_now_ns() - local_outlier_factor_free(probe_local_outlier_factor) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_local_outlier_factor: LocalOutlierFactor = local_outlier_factor_fit(X_c, 5) - local_outlier_factor_free(m_local_outlier_factor) - } - t1 = flow_now_ns() - printf("ESTIMATOR|local_outlier_factor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) - fflush(null) - - # ---- logistic_regression_cv (classification) ---- - t0 = flow_now_ns() - let probe_logistic_regression_cv: LogisticRegressionCV = logistic_regression_cv_fit(X_c, y_c, 3, 5) - t1 = flow_now_ns() - logistic_regression_cv_free(probe_logistic_regression_cv) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_logistic_regression_cv: LogisticRegressionCV = logistic_regression_cv_fit(X_c, y_c, 3, 5) - logistic_regression_cv_free(m_logistic_regression_cv) - } - t1 = flow_now_ns() - let fitted_logistic_regression_cv: LogisticRegressionCV = logistic_regression_cv_fit(X_c, y_c, 3, 5) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_logistic_regression_cv: ptr = logistic_regression_cv_predict(fitted_logistic_regression_cv, X_c) - array_free_f32(o_logistic_regression_cv) - } - t3 = flow_now_ns() - logistic_regression_cv_free(fitted_logistic_regression_cv) - printf("ESTIMATOR|logistic_regression_cv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- logistic_regression (classification) ---- - t0 = flow_now_ns() - let probe_logistic_regression: LogisticRegression = logistic_regression_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) - t1 = flow_now_ns() - logistic_regression_free(probe_logistic_regression) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_logistic_regression: LogisticRegression = logistic_regression_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) - logistic_regression_free(m_logistic_regression) - } - t1 = flow_now_ns() - printf("ESTIMATOR|logistic_regression|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) - fflush(null) - - # ---- maxabs_scaler (unsupervised) ---- - t0 = flow_now_ns() - let probe_maxabs_scaler: MaxAbsScaler = maxabs_scaler_fit(X_c) - t1 = flow_now_ns() - maxabs_scaler_free(probe_maxabs_scaler) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_maxabs_scaler: MaxAbsScaler = maxabs_scaler_fit(X_c) - maxabs_scaler_free(m_maxabs_scaler) - } - t1 = flow_now_ns() - let fitted_maxabs_scaler: MaxAbsScaler = maxabs_scaler_fit(X_c) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_maxabs_scaler: Matrix = maxabs_scaler_transform(fitted_maxabs_scaler, X_c) - matrix_free(o_maxabs_scaler) - } - t3 = flow_now_ns() - maxabs_scaler_free(fitted_maxabs_scaler) - printf("ESTIMATOR|maxabs_scaler|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) @@ -529,5 +582,10 @@ function main() -> i32 { matrix_free(Y_multi) free(yi_c as ptr) free(yi_r as ptr) + array_free_f32(x1d_r) + array_free_f32(w_f) + # The sink is printed so the work above cannot be optimized away. The + # parser matches ESTIMATOR lines only, so this one is ignored. + printf("SINK|%.9f\n", sink) return 0 } diff --git a/benchmarks/generated/bench_estimators_04.flow b/benchmarks/generated/bench_estimators_04.flow index 132eb6e..1da5bfe 100644 --- a/benchmarks/generated/bench_estimators_04.flow +++ b/benchmarks/generated/bench_estimators_04.flow @@ -65,11 +65,179 @@ function main() -> i32 { Y_rows[i] = row } + # A one-dimensional x for the isotonic row, which regresses against a + # single ordered variable rather than a design. + let x1d_r: ptr = array_new_f32(n_r) + for i in 0 to n_r { x1d_r[i] = matrix_at(X_r, i, 0) } + + # Per-feature importances for the selector row, which takes the weights a + # fitted model would hand it rather than a design. + let w_f: ptr = array_new_f32(f_c) + for i in 0 to f_c { w_f[i] = 1.0 / ((i + 1) as f32) } + let mut t0: i64 = 0 let mut t1: i64 = 0 let mut t2: i64 = 0 let mut t3: i64 = 0 let mut reps: i32 = 1 + let mut sink: f32 = 0.0 + + # ---- linear_svc_multi (classification) ---- + t0 = flow_now_ns() + let probe_linear_svc_multi: LinearSVCMulti = linear_svc_multi_fit(X_c, y_c, 3, 1.0, 50) + t1 = flow_now_ns() + linear_svc_multi_free(probe_linear_svc_multi) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_linear_svc_multi: LinearSVCMulti = linear_svc_multi_fit(X_c, y_c, 3, 1.0, 50) + linear_svc_multi_free(m_linear_svc_multi) + } + t1 = flow_now_ns() + let fitted_linear_svc_multi: LinearSVCMulti = linear_svc_multi_fit(X_c, y_c, 3, 1.0, 50) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_linear_svc_multi: ptr = linear_svc_multi_predict(fitted_linear_svc_multi, X_c) + sink = sink + o_linear_svc_multi[0] + array_free_f32(o_linear_svc_multi) + } + t3 = flow_now_ns() + linear_svc_multi_free(fitted_linear_svc_multi) + printf("ESTIMATOR|linear_svc_multi|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- linear_svr (regression) ---- + t0 = flow_now_ns() + let probe_linear_svr: LinearSVR = linear_svr_fit(X_r, y_r, 1.0, 0.1, 50, 0.01) + t1 = flow_now_ns() + linear_svr_free(probe_linear_svr) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_linear_svr: LinearSVR = linear_svr_fit(X_r, y_r, 1.0, 0.1, 50, 0.01) + linear_svr_free(m_linear_svr) + } + t1 = flow_now_ns() + let fitted_linear_svr: LinearSVR = linear_svr_fit(X_r, y_r, 1.0, 0.1, 50, 0.01) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_linear_svr: ptr = linear_svr_predict(fitted_linear_svr, X_r) + sink = sink + o_linear_svr[0] + array_free_f32(o_linear_svr) + } + t3 = flow_now_ns() + linear_svr_free(fitted_linear_svr) + printf("ESTIMATOR|linear_svr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- lle (unsupervised) ---- + t0 = flow_now_ns() + let probe_lle: LLE = lle_fit(X_c, 2, 5) + t1 = flow_now_ns() + lle_free(probe_lle) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_lle: LLE = lle_fit(X_c, 2, 5) + lle_free(m_lle) + } + t1 = flow_now_ns() + printf("ESTIMATOR|lle|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + + # ---- local_outlier_factor (unsupervised) ---- + t0 = flow_now_ns() + let probe_local_outlier_factor: LocalOutlierFactor = local_outlier_factor_fit(X_c, 5) + t1 = flow_now_ns() + local_outlier_factor_free(probe_local_outlier_factor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_local_outlier_factor: LocalOutlierFactor = local_outlier_factor_fit(X_c, 5) + local_outlier_factor_free(m_local_outlier_factor) + } + t1 = flow_now_ns() + printf("ESTIMATOR|local_outlier_factor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + + # ---- logistic_regression_cv (classification) ---- + t0 = flow_now_ns() + let probe_logistic_regression_cv: LogisticRegressionCV = logistic_regression_cv_fit(X_c, y_c, 3, 5) + t1 = flow_now_ns() + logistic_regression_cv_free(probe_logistic_regression_cv) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_logistic_regression_cv: LogisticRegressionCV = logistic_regression_cv_fit(X_c, y_c, 3, 5) + logistic_regression_cv_free(m_logistic_regression_cv) + } + t1 = flow_now_ns() + let fitted_logistic_regression_cv: LogisticRegressionCV = logistic_regression_cv_fit(X_c, y_c, 3, 5) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_logistic_regression_cv: ptr = logistic_regression_cv_predict(fitted_logistic_regression_cv, X_c) + sink = sink + o_logistic_regression_cv[0] + array_free_f32(o_logistic_regression_cv) + } + t3 = flow_now_ns() + logistic_regression_cv_free(fitted_logistic_regression_cv) + printf("ESTIMATOR|logistic_regression_cv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- logistic_regression (classification) ---- + t0 = flow_now_ns() + let probe_logistic_regression: LogisticRegression = logistic_regression_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) + t1 = flow_now_ns() + logistic_regression_free(probe_logistic_regression) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_logistic_regression: LogisticRegression = logistic_regression_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) + logistic_regression_free(m_logistic_regression) + } + t1 = flow_now_ns() + printf("ESTIMATOR|logistic_regression|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + + # ---- maxabs_scaler (unsupervised) ---- + t0 = flow_now_ns() + let probe_maxabs_scaler: MaxAbsScaler = maxabs_scaler_fit(X_c) + t1 = flow_now_ns() + maxabs_scaler_free(probe_maxabs_scaler) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_maxabs_scaler: MaxAbsScaler = maxabs_scaler_fit(X_c) + maxabs_scaler_free(m_maxabs_scaler) + } + t1 = flow_now_ns() + let fitted_maxabs_scaler: MaxAbsScaler = maxabs_scaler_fit(X_c) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_maxabs_scaler: Matrix = maxabs_scaler_transform(fitted_maxabs_scaler, X_c) + if o_maxabs_scaler.rows > 0 { + if o_maxabs_scaler.cols > 0 { sink = sink + o_maxabs_scaler.data[0] } + } + matrix_free(o_maxabs_scaler) + } + t3 = flow_now_ns() + maxabs_scaler_free(fitted_maxabs_scaler) + printf("ESTIMATOR|maxabs_scaler|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) # ---- mds (unsupervised) ---- t0 = flow_now_ns() @@ -140,6 +308,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_minibatch_dictionary_learning: Matrix = minibatch_dictionary_learning_transform(fitted_minibatch_dictionary_learning, X_c) + if o_minibatch_dictionary_learning.rows > 0 { + if o_minibatch_dictionary_learning.cols > 0 { sink = sink + o_minibatch_dictionary_learning.data[0] } + } matrix_free(o_minibatch_dictionary_learning) } t3 = flow_now_ns() @@ -182,6 +353,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_minibatch_nmf: Matrix = minibatch_nmf_transform(fitted_minibatch_nmf, X_c) + if o_minibatch_nmf.rows > 0 { + if o_minibatch_nmf.cols > 0 { sink = sink + o_minibatch_nmf.data[0] } + } matrix_free(o_minibatch_nmf) } t3 = flow_now_ns() @@ -207,6 +381,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_minibatch_sparse_pca: Matrix = minibatch_sparse_pca_transform(fitted_minibatch_sparse_pca, X_c) + if o_minibatch_sparse_pca.rows > 0 { + if o_minibatch_sparse_pca.cols > 0 { sink = sink + o_minibatch_sparse_pca.data[0] } + } matrix_free(o_minibatch_sparse_pca) } t3 = flow_now_ns() @@ -232,6 +409,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_minmax_scaler: Matrix = minmax_scaler_transform(fitted_minmax_scaler, X_c) + if o_minmax_scaler.rows > 0 { + if o_minmax_scaler.cols > 0 { sink = sink + o_minmax_scaler.data[0] } + } matrix_free(o_minmax_scaler) } t3 = flow_now_ns() @@ -257,6 +437,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_missing_indicator: Matrix = missing_indicator_transform(fitted_missing_indicator, X_c) + if o_missing_indicator.rows > 0 { + if o_missing_indicator.cols > 0 { sink = sink + o_missing_indicator.data[0] } + } matrix_free(o_missing_indicator) } t3 = flow_now_ns() @@ -283,6 +466,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_mlp_classifier: Matrix = mlp_classifier_predict(fitted_mlp_classifier, X_c) + if o_mlp_classifier.rows > 0 { + if o_mlp_classifier.cols > 0 { sink = sink + o_mlp_classifier.data[0] } + } matrix_free(o_mlp_classifier) } t3 = flow_now_ns() @@ -309,6 +495,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_mlp_regressor: ptr = mlp_regressor_predict(fitted_mlp_regressor, X_r) + sink = sink + o_mlp_regressor[0] array_free_f32(o_mlp_regressor) } t3 = flow_now_ns() @@ -334,6 +521,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_multi_output_classifier: Matrix = multi_output_classifier_predict(fitted_multi_output_classifier, X_c) + if o_multi_output_classifier.rows > 0 { + if o_multi_output_classifier.cols > 0 { sink = sink + o_multi_output_classifier.data[0] } + } matrix_free(o_multi_output_classifier) } t3 = flow_now_ns() @@ -359,6 +549,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_multi_output_regressor: Matrix = multi_output_regressor_predict(fitted_multi_output_regressor, X_r) + if o_multi_output_regressor.rows > 0 { + if o_multi_output_regressor.cols > 0 { sink = sink + o_multi_output_regressor.data[0] } + } matrix_free(o_multi_output_regressor) } t3 = flow_now_ns() @@ -366,181 +559,6 @@ function main() -> i32 { printf("ESTIMATOR|multi_output_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- multiclass_logistic (classification) ---- - t0 = flow_now_ns() - let probe_multiclass_logistic: MultiClassLogisticRegression = multiclass_logistic_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) - t1 = flow_now_ns() - multiclass_logistic_free(probe_multiclass_logistic) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_multiclass_logistic: MultiClassLogisticRegression = multiclass_logistic_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) - multiclass_logistic_free(m_multiclass_logistic) - } - t1 = flow_now_ns() - let fitted_multiclass_logistic: MultiClassLogisticRegression = multiclass_logistic_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_multiclass_logistic: Matrix = multiclass_logistic_predict(fitted_multiclass_logistic, X_c) - matrix_free(o_multiclass_logistic) - } - t3 = flow_now_ns() - multiclass_logistic_free(fitted_multiclass_logistic) - printf("ESTIMATOR|multiclass_logistic|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- multinomial_nb (classification) ---- - t0 = flow_now_ns() - let probe_multinomial_nb: MultinomialNB = multinomial_nb_fit(X_c, y_c, 3, 1.0) - t1 = flow_now_ns() - multinomial_nb_free(probe_multinomial_nb) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_multinomial_nb: MultinomialNB = multinomial_nb_fit(X_c, y_c, 3, 1.0) - multinomial_nb_free(m_multinomial_nb) - } - t1 = flow_now_ns() - let fitted_multinomial_nb: MultinomialNB = multinomial_nb_fit(X_c, y_c, 3, 1.0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_multinomial_nb: ptr = multinomial_nb_predict(fitted_multinomial_nb, X_c) - array_free_f32(o_multinomial_nb) - } - t3 = flow_now_ns() - multinomial_nb_free(fitted_multinomial_nb) - printf("ESTIMATOR|multinomial_nb|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- multitask_elastic_net_cv (multioutput) ---- - t0 = flow_now_ns() - let probe_multitask_elastic_net_cv: MultiTaskElasticNetCV = multitask_elastic_net_cv_fit(X_r, Y_multi, 10) - t1 = flow_now_ns() - multitask_elastic_net_cv_free(probe_multitask_elastic_net_cv) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_multitask_elastic_net_cv: MultiTaskElasticNetCV = multitask_elastic_net_cv_fit(X_r, Y_multi, 10) - multitask_elastic_net_cv_free(m_multitask_elastic_net_cv) - } - t1 = flow_now_ns() - let fitted_multitask_elastic_net_cv: MultiTaskElasticNetCV = multitask_elastic_net_cv_fit(X_r, Y_multi, 10) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_multitask_elastic_net_cv: Matrix = multitask_elastic_net_cv_predict(fitted_multitask_elastic_net_cv, X_r) - matrix_free(o_multitask_elastic_net_cv) - } - t3 = flow_now_ns() - multitask_elastic_net_cv_free(fitted_multitask_elastic_net_cv) - printf("ESTIMATOR|multitask_elastic_net_cv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- multitask_elastic_net (multioutput) ---- - t0 = flow_now_ns() - let probe_multitask_elastic_net: MultiTaskElasticNet = multitask_elastic_net_fit(X_r, Y_multi, 1.0, 0.5, 50, 0.01) - t1 = flow_now_ns() - multitask_elastic_net_free(probe_multitask_elastic_net) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_multitask_elastic_net: MultiTaskElasticNet = multitask_elastic_net_fit(X_r, Y_multi, 1.0, 0.5, 50, 0.01) - multitask_elastic_net_free(m_multitask_elastic_net) - } - t1 = flow_now_ns() - let fitted_multitask_elastic_net: MultiTaskElasticNet = multitask_elastic_net_fit(X_r, Y_multi, 1.0, 0.5, 50, 0.01) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_multitask_elastic_net: Matrix = multitask_elastic_net_predict(fitted_multitask_elastic_net, X_r) - matrix_free(o_multitask_elastic_net) - } - t3 = flow_now_ns() - multitask_elastic_net_free(fitted_multitask_elastic_net) - printf("ESTIMATOR|multitask_elastic_net|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- multitask_lasso_cv (multioutput) ---- - t0 = flow_now_ns() - let probe_multitask_lasso_cv: MultiTaskLassoCV = multitask_lasso_cv_fit(X_r, Y_multi, 10) - t1 = flow_now_ns() - multitask_lasso_cv_free(probe_multitask_lasso_cv) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_multitask_lasso_cv: MultiTaskLassoCV = multitask_lasso_cv_fit(X_r, Y_multi, 10) - multitask_lasso_cv_free(m_multitask_lasso_cv) - } - t1 = flow_now_ns() - let fitted_multitask_lasso_cv: MultiTaskLassoCV = multitask_lasso_cv_fit(X_r, Y_multi, 10) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_multitask_lasso_cv: Matrix = multitask_lasso_cv_predict(fitted_multitask_lasso_cv, X_r) - matrix_free(o_multitask_lasso_cv) - } - t3 = flow_now_ns() - multitask_lasso_cv_free(fitted_multitask_lasso_cv) - printf("ESTIMATOR|multitask_lasso_cv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- multitask_lasso (multioutput) ---- - t0 = flow_now_ns() - let probe_multitask_lasso: MultiTaskLasso = multitask_lasso_fit(X_r, Y_multi, 1.0, 50, 0.01) - t1 = flow_now_ns() - multitask_lasso_free(probe_multitask_lasso) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_multitask_lasso: MultiTaskLasso = multitask_lasso_fit(X_r, Y_multi, 1.0, 50, 0.01) - multitask_lasso_free(m_multitask_lasso) - } - t1 = flow_now_ns() - let fitted_multitask_lasso: MultiTaskLasso = multitask_lasso_fit(X_r, Y_multi, 1.0, 50, 0.01) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_multitask_lasso: Matrix = multitask_lasso_predict(fitted_multitask_lasso, X_r) - matrix_free(o_multitask_lasso) - } - t3 = flow_now_ns() - multitask_lasso_free(fitted_multitask_lasso) - printf("ESTIMATOR|multitask_lasso|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- nca (regression) ---- - t0 = flow_now_ns() - let probe_nca: NeighborhoodComponentsAnalysis = nca_fit(X_r, y_r, 2, 50, 0.01) - t1 = flow_now_ns() - nca_free(probe_nca) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_nca: NeighborhoodComponentsAnalysis = nca_fit(X_r, y_r, 2, 50, 0.01) - nca_free(m_nca) - } - t1 = flow_now_ns() - let fitted_nca: NeighborhoodComponentsAnalysis = nca_fit(X_r, y_r, 2, 50, 0.01) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_nca: Matrix = nca_transform(fitted_nca, X_r) - matrix_free(o_nca) - } - t3 = flow_now_ns() - nca_free(fitted_nca) - printf("ESTIMATOR|nca|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) @@ -549,5 +567,10 @@ function main() -> i32 { matrix_free(Y_multi) free(yi_c as ptr) free(yi_r as ptr) + array_free_f32(x1d_r) + array_free_f32(w_f) + # The sink is printed so the work above cannot be optimized away. The + # parser matches ESTIMATOR lines only, so this one is ignored. + printf("SINK|%.9f\n", sink) return 0 } diff --git a/benchmarks/generated/bench_estimators_05.flow b/benchmarks/generated/bench_estimators_05.flow index 37b3de4..0c54ec5 100644 --- a/benchmarks/generated/bench_estimators_05.flow +++ b/benchmarks/generated/bench_estimators_05.flow @@ -65,11 +65,216 @@ function main() -> i32 { Y_rows[i] = row } + # A one-dimensional x for the isotonic row, which regresses against a + # single ordered variable rather than a design. + let x1d_r: ptr = array_new_f32(n_r) + for i in 0 to n_r { x1d_r[i] = matrix_at(X_r, i, 0) } + + # Per-feature importances for the selector row, which takes the weights a + # fitted model would hand it rather than a design. + let w_f: ptr = array_new_f32(f_c) + for i in 0 to f_c { w_f[i] = 1.0 / ((i + 1) as f32) } + let mut t0: i64 = 0 let mut t1: i64 = 0 let mut t2: i64 = 0 let mut t3: i64 = 0 let mut reps: i32 = 1 + let mut sink: f32 = 0.0 + + # ---- multiclass_logistic (classification) ---- + t0 = flow_now_ns() + let probe_multiclass_logistic: MultiClassLogisticRegression = multiclass_logistic_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) + t1 = flow_now_ns() + multiclass_logistic_free(probe_multiclass_logistic) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_multiclass_logistic: MultiClassLogisticRegression = multiclass_logistic_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) + multiclass_logistic_free(m_multiclass_logistic) + } + t1 = flow_now_ns() + let fitted_multiclass_logistic: MultiClassLogisticRegression = multiclass_logistic_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_multiclass_logistic: Matrix = multiclass_logistic_predict(fitted_multiclass_logistic, X_c) + if o_multiclass_logistic.rows > 0 { + if o_multiclass_logistic.cols > 0 { sink = sink + o_multiclass_logistic.data[0] } + } + matrix_free(o_multiclass_logistic) + } + t3 = flow_now_ns() + multiclass_logistic_free(fitted_multiclass_logistic) + printf("ESTIMATOR|multiclass_logistic|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- multinomial_nb (classification) ---- + t0 = flow_now_ns() + let probe_multinomial_nb: MultinomialNB = multinomial_nb_fit(X_c, y_c, 3, 1.0) + t1 = flow_now_ns() + multinomial_nb_free(probe_multinomial_nb) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_multinomial_nb: MultinomialNB = multinomial_nb_fit(X_c, y_c, 3, 1.0) + multinomial_nb_free(m_multinomial_nb) + } + t1 = flow_now_ns() + let fitted_multinomial_nb: MultinomialNB = multinomial_nb_fit(X_c, y_c, 3, 1.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_multinomial_nb: ptr = multinomial_nb_predict(fitted_multinomial_nb, X_c) + sink = sink + o_multinomial_nb[0] + array_free_f32(o_multinomial_nb) + } + t3 = flow_now_ns() + multinomial_nb_free(fitted_multinomial_nb) + printf("ESTIMATOR|multinomial_nb|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- multitask_elastic_net_cv (multioutput) ---- + t0 = flow_now_ns() + let probe_multitask_elastic_net_cv: MultiTaskElasticNetCV = multitask_elastic_net_cv_fit(X_r, Y_multi, 10) + t1 = flow_now_ns() + multitask_elastic_net_cv_free(probe_multitask_elastic_net_cv) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_multitask_elastic_net_cv: MultiTaskElasticNetCV = multitask_elastic_net_cv_fit(X_r, Y_multi, 10) + multitask_elastic_net_cv_free(m_multitask_elastic_net_cv) + } + t1 = flow_now_ns() + let fitted_multitask_elastic_net_cv: MultiTaskElasticNetCV = multitask_elastic_net_cv_fit(X_r, Y_multi, 10) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_multitask_elastic_net_cv: Matrix = multitask_elastic_net_cv_predict(fitted_multitask_elastic_net_cv, X_r) + if o_multitask_elastic_net_cv.rows > 0 { + if o_multitask_elastic_net_cv.cols > 0 { sink = sink + o_multitask_elastic_net_cv.data[0] } + } + matrix_free(o_multitask_elastic_net_cv) + } + t3 = flow_now_ns() + multitask_elastic_net_cv_free(fitted_multitask_elastic_net_cv) + printf("ESTIMATOR|multitask_elastic_net_cv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- multitask_elastic_net (multioutput) ---- + t0 = flow_now_ns() + let probe_multitask_elastic_net: MultiTaskElasticNet = multitask_elastic_net_fit(X_r, Y_multi, 1.0, 0.5, 50, 0.01) + t1 = flow_now_ns() + multitask_elastic_net_free(probe_multitask_elastic_net) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_multitask_elastic_net: MultiTaskElasticNet = multitask_elastic_net_fit(X_r, Y_multi, 1.0, 0.5, 50, 0.01) + multitask_elastic_net_free(m_multitask_elastic_net) + } + t1 = flow_now_ns() + let fitted_multitask_elastic_net: MultiTaskElasticNet = multitask_elastic_net_fit(X_r, Y_multi, 1.0, 0.5, 50, 0.01) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_multitask_elastic_net: Matrix = multitask_elastic_net_predict(fitted_multitask_elastic_net, X_r) + if o_multitask_elastic_net.rows > 0 { + if o_multitask_elastic_net.cols > 0 { sink = sink + o_multitask_elastic_net.data[0] } + } + matrix_free(o_multitask_elastic_net) + } + t3 = flow_now_ns() + multitask_elastic_net_free(fitted_multitask_elastic_net) + printf("ESTIMATOR|multitask_elastic_net|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- multitask_lasso_cv (multioutput) ---- + t0 = flow_now_ns() + let probe_multitask_lasso_cv: MultiTaskLassoCV = multitask_lasso_cv_fit(X_r, Y_multi, 10) + t1 = flow_now_ns() + multitask_lasso_cv_free(probe_multitask_lasso_cv) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_multitask_lasso_cv: MultiTaskLassoCV = multitask_lasso_cv_fit(X_r, Y_multi, 10) + multitask_lasso_cv_free(m_multitask_lasso_cv) + } + t1 = flow_now_ns() + let fitted_multitask_lasso_cv: MultiTaskLassoCV = multitask_lasso_cv_fit(X_r, Y_multi, 10) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_multitask_lasso_cv: Matrix = multitask_lasso_cv_predict(fitted_multitask_lasso_cv, X_r) + if o_multitask_lasso_cv.rows > 0 { + if o_multitask_lasso_cv.cols > 0 { sink = sink + o_multitask_lasso_cv.data[0] } + } + matrix_free(o_multitask_lasso_cv) + } + t3 = flow_now_ns() + multitask_lasso_cv_free(fitted_multitask_lasso_cv) + printf("ESTIMATOR|multitask_lasso_cv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- multitask_lasso (multioutput) ---- + t0 = flow_now_ns() + let probe_multitask_lasso: MultiTaskLasso = multitask_lasso_fit(X_r, Y_multi, 1.0, 50, 0.01) + t1 = flow_now_ns() + multitask_lasso_free(probe_multitask_lasso) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_multitask_lasso: MultiTaskLasso = multitask_lasso_fit(X_r, Y_multi, 1.0, 50, 0.01) + multitask_lasso_free(m_multitask_lasso) + } + t1 = flow_now_ns() + let fitted_multitask_lasso: MultiTaskLasso = multitask_lasso_fit(X_r, Y_multi, 1.0, 50, 0.01) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_multitask_lasso: Matrix = multitask_lasso_predict(fitted_multitask_lasso, X_r) + if o_multitask_lasso.rows > 0 { + if o_multitask_lasso.cols > 0 { sink = sink + o_multitask_lasso.data[0] } + } + matrix_free(o_multitask_lasso) + } + t3 = flow_now_ns() + multitask_lasso_free(fitted_multitask_lasso) + printf("ESTIMATOR|multitask_lasso|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- nca (regression) ---- + t0 = flow_now_ns() + let probe_nca: NeighborhoodComponentsAnalysis = nca_fit(X_r, y_r, 2, 50, 0.01) + t1 = flow_now_ns() + nca_free(probe_nca) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_nca: NeighborhoodComponentsAnalysis = nca_fit(X_r, y_r, 2, 50, 0.01) + nca_free(m_nca) + } + t1 = flow_now_ns() + let fitted_nca: NeighborhoodComponentsAnalysis = nca_fit(X_r, y_r, 2, 50, 0.01) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_nca: Matrix = nca_transform(fitted_nca, X_r) + if o_nca.rows > 0 { + if o_nca.cols > 0 { sink = sink + o_nca.data[0] } + } + matrix_free(o_nca) + } + t3 = flow_now_ns() + nca_free(fitted_nca) + printf("ESTIMATOR|nca|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) # ---- nearest_centroid (classification) ---- t0 = flow_now_ns() @@ -89,6 +294,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_nearest_centroid: ptr = nearest_centroid_predict(fitted_nearest_centroid, X_c) + sink = sink + o_nearest_centroid[0] array_free_f32(o_nearest_centroid) } t3 = flow_now_ns() @@ -131,6 +337,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_nmf: Matrix = nmf_transform(fitted_nmf, X_c) + if o_nmf.rows > 0 { + if o_nmf.cols > 0 { sink = sink + o_nmf.data[0] } + } matrix_free(o_nmf) } t3 = flow_now_ns() @@ -156,6 +365,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_nu_svc: ptr = nu_svc_predict(fitted_nu_svc, X_r) + sink = sink + o_nu_svc[0] array_free_f32(o_nu_svc) } t3 = flow_now_ns() @@ -181,6 +391,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_nu_svr: ptr = nu_svr_predict(fitted_nu_svr, X_r) + sink = sink + o_nu_svr[0] array_free_f32(o_nu_svr) } t3 = flow_now_ns() @@ -206,6 +417,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_nystroem: Matrix = nystroem_transform(fitted_nystroem, X_c) + if o_nystroem.rows > 0 { + if o_nystroem.cols > 0 { sink = sink + o_nystroem.data[0] } + } matrix_free(o_nystroem) } t3 = flow_now_ns() @@ -246,6 +460,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_omp_cv: ptr = omp_cv_predict(fitted_omp_cv, X_r) + sink = sink + o_omp_cv[0] array_free_f32(o_omp_cv) } t3 = flow_now_ns() @@ -288,6 +503,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_one_vs_one: ptr = one_vs_one_predict(fitted_one_vs_one, X_c) + sink = sink + o_one_vs_one[0] array_free_f32(o_one_vs_one) } t3 = flow_now_ns() @@ -313,6 +529,7 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_one_vs_rest: ptr = one_vs_rest_predict(fitted_one_vs_rest, X_c) + sink = sink + o_one_vs_rest[0] array_free_f32(o_one_vs_rest) } t3 = flow_now_ns() @@ -338,6 +555,9 @@ function main() -> i32 { t2 = flow_now_ns() for rep2 in 0 to reps { let o_onehot_encoder: Matrix = onehot_encoder_transform(fitted_onehot_encoder, X_c) + if o_onehot_encoder.rows > 0 { + if o_onehot_encoder.cols > 0 { sink = sink + o_onehot_encoder.data[0] } + } matrix_free(o_onehot_encoder) } t3 = flow_now_ns() @@ -362,181 +582,6 @@ function main() -> i32 { printf("ESTIMATOR|optics|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) fflush(null) - # ---- ordinal_encoder (unsupervised) ---- - t0 = flow_now_ns() - let probe_ordinal_encoder: OrdinalEncoder = ordinal_encoder_fit(X_c, 0) - t1 = flow_now_ns() - ordinal_encoder_free(probe_ordinal_encoder) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_ordinal_encoder: OrdinalEncoder = ordinal_encoder_fit(X_c, 0) - ordinal_encoder_free(m_ordinal_encoder) - } - t1 = flow_now_ns() - let fitted_ordinal_encoder: OrdinalEncoder = ordinal_encoder_fit(X_c, 0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_ordinal_encoder: Matrix = ordinal_encoder_transform(fitted_ordinal_encoder, X_c) - matrix_free(o_ordinal_encoder) - } - t3 = flow_now_ns() - ordinal_encoder_free(fitted_ordinal_encoder) - printf("ESTIMATOR|ordinal_encoder|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- orthogonal_matching_pursuit (regression) ---- - t0 = flow_now_ns() - let probe_orthogonal_matching_pursuit: OrthogonalMatchingPursuit = orthogonal_matching_pursuit_fit(X_r, y_r, 3) - t1 = flow_now_ns() - orthogonal_matching_pursuit_free(probe_orthogonal_matching_pursuit) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_orthogonal_matching_pursuit: OrthogonalMatchingPursuit = orthogonal_matching_pursuit_fit(X_r, y_r, 3) - orthogonal_matching_pursuit_free(m_orthogonal_matching_pursuit) - } - t1 = flow_now_ns() - let fitted_orthogonal_matching_pursuit: OrthogonalMatchingPursuit = orthogonal_matching_pursuit_fit(X_r, y_r, 3) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_orthogonal_matching_pursuit: ptr = orthogonal_matching_pursuit_predict(fitted_orthogonal_matching_pursuit, X_r) - array_free_f32(o_orthogonal_matching_pursuit) - } - t3 = flow_now_ns() - orthogonal_matching_pursuit_free(fitted_orthogonal_matching_pursuit) - printf("ESTIMATOR|orthogonal_matching_pursuit|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- output_code (classification) ---- - t0 = flow_now_ns() - let probe_output_code: OutputCodeClassifier = output_code_fit(X_c, y_c, 3, 4, 50, 0.01, penalty_none()) - t1 = flow_now_ns() - output_code_free(probe_output_code) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_output_code: OutputCodeClassifier = output_code_fit(X_c, y_c, 3, 4, 50, 0.01, penalty_none()) - output_code_free(m_output_code) - } - t1 = flow_now_ns() - let fitted_output_code: OutputCodeClassifier = output_code_fit(X_c, y_c, 3, 4, 50, 0.01, penalty_none()) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_output_code: ptr = output_code_predict(fitted_output_code, X_c) - array_free_f32(o_output_code) - } - t3 = flow_now_ns() - output_code_free(fitted_output_code) - printf("ESTIMATOR|output_code|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- passive_aggressive_classifier (regression) ---- - t0 = flow_now_ns() - let probe_passive_aggressive_classifier: PassiveAggressiveClassifier = passive_aggressive_classifier_fit(X_r, y_r, 50, 1.0) - t1 = flow_now_ns() - passive_aggressive_classifier_free(probe_passive_aggressive_classifier) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_passive_aggressive_classifier: PassiveAggressiveClassifier = passive_aggressive_classifier_fit(X_r, y_r, 50, 1.0) - passive_aggressive_classifier_free(m_passive_aggressive_classifier) - } - t1 = flow_now_ns() - let fitted_passive_aggressive_classifier: PassiveAggressiveClassifier = passive_aggressive_classifier_fit(X_r, y_r, 50, 1.0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_passive_aggressive_classifier: ptr = passive_aggressive_classifier_predict(fitted_passive_aggressive_classifier, X_r) - array_free_f32(o_passive_aggressive_classifier) - } - t3 = flow_now_ns() - passive_aggressive_classifier_free(fitted_passive_aggressive_classifier) - printf("ESTIMATOR|passive_aggressive_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- passive_aggressive_regressor (regression) ---- - t0 = flow_now_ns() - let probe_passive_aggressive_regressor: PassiveAggressiveRegressor = passive_aggressive_regressor_fit(X_r, y_r, 1.0, 100) - t1 = flow_now_ns() - passive_aggressive_regressor_free(probe_passive_aggressive_regressor) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_passive_aggressive_regressor: PassiveAggressiveRegressor = passive_aggressive_regressor_fit(X_r, y_r, 1.0, 100) - passive_aggressive_regressor_free(m_passive_aggressive_regressor) - } - t1 = flow_now_ns() - let fitted_passive_aggressive_regressor: PassiveAggressiveRegressor = passive_aggressive_regressor_fit(X_r, y_r, 1.0, 100) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_passive_aggressive_regressor: ptr = passive_aggressive_regressor_predict(fitted_passive_aggressive_regressor, X_r) - array_free_f32(o_passive_aggressive_regressor) - } - t3 = flow_now_ns() - passive_aggressive_regressor_free(fitted_passive_aggressive_regressor) - printf("ESTIMATOR|passive_aggressive_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- pca (unsupervised) ---- - t0 = flow_now_ns() - let probe_pca: PCA = pca_fit(X_c, 2) - t1 = flow_now_ns() - pca_free(probe_pca) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_pca: PCA = pca_fit(X_c, 2) - pca_free(m_pca) - } - t1 = flow_now_ns() - let fitted_pca: PCA = pca_fit(X_c, 2) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_pca: Matrix = pca_transform(fitted_pca, X_c) - matrix_free(o_pca) - } - t3 = flow_now_ns() - pca_free(fitted_pca) - printf("ESTIMATOR|pca|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- perceptron (regression) ---- - t0 = flow_now_ns() - let probe_perceptron: Perceptron = perceptron_fit(X_r, y_r, 50, 0.01, 42) - t1 = flow_now_ns() - perceptron_free(probe_perceptron) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_perceptron: Perceptron = perceptron_fit(X_r, y_r, 50, 0.01, 42) - perceptron_free(m_perceptron) - } - t1 = flow_now_ns() - let fitted_perceptron: Perceptron = perceptron_fit(X_r, y_r, 50, 0.01, 42) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_perceptron: ptr = perceptron_predict(fitted_perceptron, X_r) - array_free_f32(o_perceptron) - } - t3 = flow_now_ns() - perceptron_free(fitted_perceptron) - printf("ESTIMATOR|perceptron|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) @@ -545,5 +590,10 @@ function main() -> i32 { matrix_free(Y_multi) free(yi_c as ptr) free(yi_r as ptr) + array_free_f32(x1d_r) + array_free_f32(w_f) + # The sink is printed so the work above cannot be optimized away. The + # parser matches ESTIMATOR lines only, so this one is ignored. + printf("SINK|%.9f\n", sink) return 0 } diff --git a/benchmarks/generated/bench_estimators_06.flow b/benchmarks/generated/bench_estimators_06.flow index 043cadb..c443010 100644 --- a/benchmarks/generated/bench_estimators_06.flow +++ b/benchmarks/generated/bench_estimators_06.flow @@ -65,494 +65,561 @@ function main() -> i32 { Y_rows[i] = row } + # A one-dimensional x for the isotonic row, which regresses against a + # single ordered variable rather than a design. + let x1d_r: ptr = array_new_f32(n_r) + for i in 0 to n_r { x1d_r[i] = matrix_at(X_r, i, 0) } + + # Per-feature importances for the selector row, which takes the weights a + # fitted model would hand it rather than a design. + let w_f: ptr = array_new_f32(f_c) + for i in 0 to f_c { w_f[i] = 1.0 / ((i + 1) as f32) } + let mut t0: i64 = 0 let mut t1: i64 = 0 let mut t2: i64 = 0 let mut t3: i64 = 0 let mut reps: i32 = 1 + let mut sink: f32 = 0.0 - # ---- pls_canonical (multioutput) ---- + # ---- ordinal_encoder (unsupervised) ---- t0 = flow_now_ns() - let probe_pls_canonical: PLSCanonical = pls_canonical_fit(X_r, Y_multi, 2) + let probe_ordinal_encoder: OrdinalEncoder = ordinal_encoder_fit(X_c, 0) t1 = flow_now_ns() - pls_canonical_free(probe_pls_canonical) + ordinal_encoder_free(probe_ordinal_encoder) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_pls_canonical: PLSCanonical = pls_canonical_fit(X_r, Y_multi, 2) - pls_canonical_free(m_pls_canonical) + let m_ordinal_encoder: OrdinalEncoder = ordinal_encoder_fit(X_c, 0) + ordinal_encoder_free(m_ordinal_encoder) } t1 = flow_now_ns() - let fitted_pls_canonical: PLSCanonical = pls_canonical_fit(X_r, Y_multi, 2) + let fitted_ordinal_encoder: OrdinalEncoder = ordinal_encoder_fit(X_c, 0) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_pls_canonical: Matrix = pls_canonical_transform(fitted_pls_canonical, X_r) - matrix_free(o_pls_canonical) + let o_ordinal_encoder: Matrix = ordinal_encoder_transform(fitted_ordinal_encoder, X_c) + if o_ordinal_encoder.rows > 0 { + if o_ordinal_encoder.cols > 0 { sink = sink + o_ordinal_encoder.data[0] } + } + matrix_free(o_ordinal_encoder) } t3 = flow_now_ns() - pls_canonical_free(fitted_pls_canonical) - printf("ESTIMATOR|pls_canonical|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + ordinal_encoder_free(fitted_ordinal_encoder) + printf("ESTIMATOR|ordinal_encoder|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- pls (multioutput) ---- + # ---- orthogonal_matching_pursuit (regression) ---- t0 = flow_now_ns() - let probe_pls: PLSRegression = pls_fit(X_r, Y_multi, 2, 100, 0.0001) + let probe_orthogonal_matching_pursuit: OrthogonalMatchingPursuit = orthogonal_matching_pursuit_fit(X_r, y_r, 3) t1 = flow_now_ns() - pls_free(probe_pls) + orthogonal_matching_pursuit_free(probe_orthogonal_matching_pursuit) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_pls: PLSRegression = pls_fit(X_r, Y_multi, 2, 100, 0.0001) - pls_free(m_pls) + let m_orthogonal_matching_pursuit: OrthogonalMatchingPursuit = orthogonal_matching_pursuit_fit(X_r, y_r, 3) + orthogonal_matching_pursuit_free(m_orthogonal_matching_pursuit) } t1 = flow_now_ns() - let fitted_pls: PLSRegression = pls_fit(X_r, Y_multi, 2, 100, 0.0001) + let fitted_orthogonal_matching_pursuit: OrthogonalMatchingPursuit = orthogonal_matching_pursuit_fit(X_r, y_r, 3) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_pls: Matrix = pls_predict(fitted_pls, X_r) - matrix_free(o_pls) + let o_orthogonal_matching_pursuit: ptr = orthogonal_matching_pursuit_predict(fitted_orthogonal_matching_pursuit, X_r) + sink = sink + o_orthogonal_matching_pursuit[0] + array_free_f32(o_orthogonal_matching_pursuit) } t3 = flow_now_ns() - pls_free(fitted_pls) - printf("ESTIMATOR|pls|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + orthogonal_matching_pursuit_free(fitted_orthogonal_matching_pursuit) + printf("ESTIMATOR|orthogonal_matching_pursuit|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- pls_svd (multioutput) ---- + # ---- output_code (classification) ---- t0 = flow_now_ns() - let probe_pls_svd: PLSSVD = pls_svd_fit(X_r, Y_multi, 2) + let probe_output_code: OutputCodeClassifier = output_code_fit(X_c, y_c, 3, 4, 50, 0.01, penalty_none()) t1 = flow_now_ns() - pls_svd_free(probe_pls_svd) + output_code_free(probe_output_code) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_pls_svd: PLSSVD = pls_svd_fit(X_r, Y_multi, 2) - pls_svd_free(m_pls_svd) + let m_output_code: OutputCodeClassifier = output_code_fit(X_c, y_c, 3, 4, 50, 0.01, penalty_none()) + output_code_free(m_output_code) } t1 = flow_now_ns() - let fitted_pls_svd: PLSSVD = pls_svd_fit(X_r, Y_multi, 2) + let fitted_output_code: OutputCodeClassifier = output_code_fit(X_c, y_c, 3, 4, 50, 0.01, penalty_none()) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_pls_svd: Matrix = pls_svd_transform(fitted_pls_svd, X_r) - matrix_free(o_pls_svd) + let o_output_code: ptr = output_code_predict(fitted_output_code, X_c) + sink = sink + o_output_code[0] + array_free_f32(o_output_code) } t3 = flow_now_ns() - pls_svd_free(fitted_pls_svd) - printf("ESTIMATOR|pls_svd|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + output_code_free(fitted_output_code) + printf("ESTIMATOR|output_code|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- poisson_regressor (regression) ---- + # ---- passive_aggressive_classifier (regression) ---- t0 = flow_now_ns() - let probe_poisson_regressor: PoissonRegressor = poisson_regressor_fit(X_r, y_r, 1.0, 100, 0.01) + let probe_passive_aggressive_classifier: PassiveAggressiveClassifier = passive_aggressive_classifier_fit(X_r, y_r, 50, 1.0) t1 = flow_now_ns() - poisson_regressor_free(probe_poisson_regressor) + passive_aggressive_classifier_free(probe_passive_aggressive_classifier) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_poisson_regressor: PoissonRegressor = poisson_regressor_fit(X_r, y_r, 1.0, 100, 0.01) - poisson_regressor_free(m_poisson_regressor) + let m_passive_aggressive_classifier: PassiveAggressiveClassifier = passive_aggressive_classifier_fit(X_r, y_r, 50, 1.0) + passive_aggressive_classifier_free(m_passive_aggressive_classifier) } t1 = flow_now_ns() - let fitted_poisson_regressor: PoissonRegressor = poisson_regressor_fit(X_r, y_r, 1.0, 100, 0.01) + let fitted_passive_aggressive_classifier: PassiveAggressiveClassifier = passive_aggressive_classifier_fit(X_r, y_r, 50, 1.0) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_poisson_regressor: ptr = poisson_regressor_predict(fitted_poisson_regressor, X_r) - array_free_f32(o_poisson_regressor) + let o_passive_aggressive_classifier: ptr = passive_aggressive_classifier_predict(fitted_passive_aggressive_classifier, X_r) + sink = sink + o_passive_aggressive_classifier[0] + array_free_f32(o_passive_aggressive_classifier) } t3 = flow_now_ns() - poisson_regressor_free(fitted_poisson_regressor) - printf("ESTIMATOR|poisson_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + passive_aggressive_classifier_free(fitted_passive_aggressive_classifier) + printf("ESTIMATOR|passive_aggressive_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- power_transformer (unsupervised) ---- + # ---- passive_aggressive_regressor (regression) ---- t0 = flow_now_ns() - let probe_power_transformer: PowerTransformer = power_transformer_fit(X_c, 0) + let probe_passive_aggressive_regressor: PassiveAggressiveRegressor = passive_aggressive_regressor_fit(X_r, y_r, 1.0, 100) t1 = flow_now_ns() - power_transformer_free(probe_power_transformer) + passive_aggressive_regressor_free(probe_passive_aggressive_regressor) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_power_transformer: PowerTransformer = power_transformer_fit(X_c, 0) - power_transformer_free(m_power_transformer) + let m_passive_aggressive_regressor: PassiveAggressiveRegressor = passive_aggressive_regressor_fit(X_r, y_r, 1.0, 100) + passive_aggressive_regressor_free(m_passive_aggressive_regressor) } t1 = flow_now_ns() - let fitted_power_transformer: PowerTransformer = power_transformer_fit(X_c, 0) + let fitted_passive_aggressive_regressor: PassiveAggressiveRegressor = passive_aggressive_regressor_fit(X_r, y_r, 1.0, 100) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_power_transformer: Matrix = power_transformer_transform(fitted_power_transformer, X_c) - matrix_free(o_power_transformer) + let o_passive_aggressive_regressor: ptr = passive_aggressive_regressor_predict(fitted_passive_aggressive_regressor, X_r) + sink = sink + o_passive_aggressive_regressor[0] + array_free_f32(o_passive_aggressive_regressor) } t3 = flow_now_ns() - power_transformer_free(fitted_power_transformer) - printf("ESTIMATOR|power_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + passive_aggressive_regressor_free(fitted_passive_aggressive_regressor) + printf("ESTIMATOR|passive_aggressive_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- qda (classification) ---- + # ---- pca (unsupervised) ---- t0 = flow_now_ns() - let probe_qda: QuadraticDiscriminantAnalysis = qda_fit(X_c, y_c, 3) + let probe_pca: PCA = pca_fit(X_c, 2) t1 = flow_now_ns() - qda_free(probe_qda) + pca_free(probe_pca) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_qda: QuadraticDiscriminantAnalysis = qda_fit(X_c, y_c, 3) - qda_free(m_qda) + let m_pca: PCA = pca_fit(X_c, 2) + pca_free(m_pca) } t1 = flow_now_ns() - let fitted_qda: QuadraticDiscriminantAnalysis = qda_fit(X_c, y_c, 3) + let fitted_pca: PCA = pca_fit(X_c, 2) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_qda: ptr = qda_predict(fitted_qda, X_c) - array_free_f32(o_qda) + let o_pca: Matrix = pca_transform(fitted_pca, X_c) + if o_pca.rows > 0 { + if o_pca.cols > 0 { sink = sink + o_pca.data[0] } + } + matrix_free(o_pca) } t3 = flow_now_ns() - qda_free(fitted_qda) - printf("ESTIMATOR|qda|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + pca_free(fitted_pca) + printf("ESTIMATOR|pca|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- quantile_regressor (regression) ---- + # ---- perceptron (regression) ---- t0 = flow_now_ns() - let probe_quantile_regressor: QuantileRegressor = quantile_regressor_fit(X_r, y_r, 0.5, 1.0, 50, 0.01) + let probe_perceptron: Perceptron = perceptron_fit(X_r, y_r, 50, 0.01, 42) t1 = flow_now_ns() - quantile_regressor_free(probe_quantile_regressor) + perceptron_free(probe_perceptron) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_quantile_regressor: QuantileRegressor = quantile_regressor_fit(X_r, y_r, 0.5, 1.0, 50, 0.01) - quantile_regressor_free(m_quantile_regressor) + let m_perceptron: Perceptron = perceptron_fit(X_r, y_r, 50, 0.01, 42) + perceptron_free(m_perceptron) } t1 = flow_now_ns() - let fitted_quantile_regressor: QuantileRegressor = quantile_regressor_fit(X_r, y_r, 0.5, 1.0, 50, 0.01) + let fitted_perceptron: Perceptron = perceptron_fit(X_r, y_r, 50, 0.01, 42) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_quantile_regressor: ptr = quantile_regressor_predict(fitted_quantile_regressor, X_r) - array_free_f32(o_quantile_regressor) + let o_perceptron: ptr = perceptron_predict(fitted_perceptron, X_r) + sink = sink + o_perceptron[0] + array_free_f32(o_perceptron) } t3 = flow_now_ns() - quantile_regressor_free(fitted_quantile_regressor) - printf("ESTIMATOR|quantile_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + perceptron_free(fitted_perceptron) + printf("ESTIMATOR|perceptron|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- quantile_transformer (unsupervised) ---- + # ---- pls_canonical (multioutput) ---- t0 = flow_now_ns() - let probe_quantile_transformer: QuantileTransformer = quantile_transformer_fit(X_c, 10, 0) + let probe_pls_canonical: PLSCanonical = pls_canonical_fit(X_r, Y_multi, 2) t1 = flow_now_ns() - quantile_transformer_free(probe_quantile_transformer) + pls_canonical_free(probe_pls_canonical) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_quantile_transformer: QuantileTransformer = quantile_transformer_fit(X_c, 10, 0) - quantile_transformer_free(m_quantile_transformer) + let m_pls_canonical: PLSCanonical = pls_canonical_fit(X_r, Y_multi, 2) + pls_canonical_free(m_pls_canonical) } t1 = flow_now_ns() - let fitted_quantile_transformer: QuantileTransformer = quantile_transformer_fit(X_c, 10, 0) + let fitted_pls_canonical: PLSCanonical = pls_canonical_fit(X_r, Y_multi, 2) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_quantile_transformer: Matrix = quantile_transformer_transform(fitted_quantile_transformer, X_c) - matrix_free(o_quantile_transformer) + let o_pls_canonical: Matrix = pls_canonical_transform(fitted_pls_canonical, X_r) + if o_pls_canonical.rows > 0 { + if o_pls_canonical.cols > 0 { sink = sink + o_pls_canonical.data[0] } + } + matrix_free(o_pls_canonical) } t3 = flow_now_ns() - quantile_transformer_free(fitted_quantile_transformer) - printf("ESTIMATOR|quantile_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + pls_canonical_free(fitted_pls_canonical) + printf("ESTIMATOR|pls_canonical|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- radius_neighbors_classifier (classification) ---- + # ---- pls (multioutput) ---- t0 = flow_now_ns() - let probe_radius_neighbors_classifier: RadiusNeighborsClassifier = radius_neighbors_classifier_fit(X_c, y_c, 1.0, 3, 0.0) + let probe_pls: PLSRegression = pls_fit(X_r, Y_multi, 2, 100, 0.0001) t1 = flow_now_ns() - radius_neighbors_classifier_free(probe_radius_neighbors_classifier) + pls_free(probe_pls) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_radius_neighbors_classifier: RadiusNeighborsClassifier = radius_neighbors_classifier_fit(X_c, y_c, 1.0, 3, 0.0) - radius_neighbors_classifier_free(m_radius_neighbors_classifier) + let m_pls: PLSRegression = pls_fit(X_r, Y_multi, 2, 100, 0.0001) + pls_free(m_pls) } t1 = flow_now_ns() - let fitted_radius_neighbors_classifier: RadiusNeighborsClassifier = radius_neighbors_classifier_fit(X_c, y_c, 1.0, 3, 0.0) + let fitted_pls: PLSRegression = pls_fit(X_r, Y_multi, 2, 100, 0.0001) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_radius_neighbors_classifier: ptr = radius_neighbors_classifier_predict(fitted_radius_neighbors_classifier, X_c) - array_free_f32(o_radius_neighbors_classifier) + let o_pls: Matrix = pls_predict(fitted_pls, X_r) + if o_pls.rows > 0 { + if o_pls.cols > 0 { sink = sink + o_pls.data[0] } + } + matrix_free(o_pls) } t3 = flow_now_ns() - radius_neighbors_classifier_free(fitted_radius_neighbors_classifier) - printf("ESTIMATOR|radius_neighbors_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + pls_free(fitted_pls) + printf("ESTIMATOR|pls|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- radius_neighbors_regressor (regression) ---- + # ---- pls_svd (multioutput) ---- t0 = flow_now_ns() - let probe_radius_neighbors_regressor: RadiusNeighborsRegressor = radius_neighbors_regressor_fit(X_r, y_r, 1.0) + let probe_pls_svd: PLSSVD = pls_svd_fit(X_r, Y_multi, 2) t1 = flow_now_ns() - radius_neighbors_regressor_free(probe_radius_neighbors_regressor) + pls_svd_free(probe_pls_svd) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_radius_neighbors_regressor: RadiusNeighborsRegressor = radius_neighbors_regressor_fit(X_r, y_r, 1.0) - radius_neighbors_regressor_free(m_radius_neighbors_regressor) + let m_pls_svd: PLSSVD = pls_svd_fit(X_r, Y_multi, 2) + pls_svd_free(m_pls_svd) } t1 = flow_now_ns() - let fitted_radius_neighbors_regressor: RadiusNeighborsRegressor = radius_neighbors_regressor_fit(X_r, y_r, 1.0) + let fitted_pls_svd: PLSSVD = pls_svd_fit(X_r, Y_multi, 2) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_radius_neighbors_regressor: ptr = radius_neighbors_regressor_predict(fitted_radius_neighbors_regressor, X_r) - array_free_f32(o_radius_neighbors_regressor) + let o_pls_svd: Matrix = pls_svd_transform(fitted_pls_svd, X_r) + if o_pls_svd.rows > 0 { + if o_pls_svd.cols > 0 { sink = sink + o_pls_svd.data[0] } + } + matrix_free(o_pls_svd) } t3 = flow_now_ns() - radius_neighbors_regressor_free(fitted_radius_neighbors_regressor) - printf("ESTIMATOR|radius_neighbors_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + pls_svd_free(fitted_pls_svd) + printf("ESTIMATOR|pls_svd|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- radius_neighbors_transformer (unsupervised) ---- + # ---- poisson_regressor (regression) ---- t0 = flow_now_ns() - let probe_radius_neighbors_transformer: RadiusNeighborsTransformer = radius_neighbors_transformer_fit(X_c, 1.0, 0) + let probe_poisson_regressor: PoissonRegressor = poisson_regressor_fit(X_r, y_r, 1.0, 100, 0.01) t1 = flow_now_ns() - radius_neighbors_transformer_free(probe_radius_neighbors_transformer) + poisson_regressor_free(probe_poisson_regressor) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_radius_neighbors_transformer: RadiusNeighborsTransformer = radius_neighbors_transformer_fit(X_c, 1.0, 0) - radius_neighbors_transformer_free(m_radius_neighbors_transformer) + let m_poisson_regressor: PoissonRegressor = poisson_regressor_fit(X_r, y_r, 1.0, 100, 0.01) + poisson_regressor_free(m_poisson_regressor) } t1 = flow_now_ns() - let fitted_radius_neighbors_transformer: RadiusNeighborsTransformer = radius_neighbors_transformer_fit(X_c, 1.0, 0) + let fitted_poisson_regressor: PoissonRegressor = poisson_regressor_fit(X_r, y_r, 1.0, 100, 0.01) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_radius_neighbors_transformer: Matrix = radius_neighbors_transformer_transform(fitted_radius_neighbors_transformer, X_c) - matrix_free(o_radius_neighbors_transformer) + let o_poisson_regressor: ptr = poisson_regressor_predict(fitted_poisson_regressor, X_r) + sink = sink + o_poisson_regressor[0] + array_free_f32(o_poisson_regressor) } t3 = flow_now_ns() - radius_neighbors_transformer_free(fitted_radius_neighbors_transformer) - printf("ESTIMATOR|radius_neighbors_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + poisson_regressor_free(fitted_poisson_regressor) + printf("ESTIMATOR|poisson_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- random_forest_classifier (classification) ---- + # ---- polynomial_count_sketch (classification, written out) ---- t0 = flow_now_ns() - let probe_random_forest_classifier: RandomForestClassifier = random_forest_classifier_fit(X_c, y_c, 3, 10, 5, 42) + let probe_polynomial_count_sketch: PolynomialCountSketch = polynomial_count_sketch_fit(f_c, 2, 2) t1 = flow_now_ns() - random_forest_classifier_free(probe_random_forest_classifier) + polynomial_count_sketch_free(probe_polynomial_count_sketch) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_random_forest_classifier: RandomForestClassifier = random_forest_classifier_fit(X_c, y_c, 3, 10, 5, 42) - random_forest_classifier_free(m_random_forest_classifier) + let m_polynomial_count_sketch: PolynomialCountSketch = polynomial_count_sketch_fit(f_c, 2, 2) + polynomial_count_sketch_free(m_polynomial_count_sketch) } t1 = flow_now_ns() - let fitted_random_forest_classifier: RandomForestClassifier = random_forest_classifier_fit(X_c, y_c, 3, 10, 5, 42) + let fitted_polynomial_count_sketch: PolynomialCountSketch = polynomial_count_sketch_fit(f_c, 2, 2) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_random_forest_classifier: ptr = random_forest_classifier_predict(fitted_random_forest_classifier, X_c) - array_free_f32(o_random_forest_classifier) + let o_polynomial_count_sketch: Matrix = polynomial_count_sketch_transform(fitted_polynomial_count_sketch, X_c) + if o_polynomial_count_sketch.rows > 0 { + if o_polynomial_count_sketch.cols > 0 { sink = sink + o_polynomial_count_sketch.data[0] } + } + matrix_free(o_polynomial_count_sketch) } t3 = flow_now_ns() - random_forest_classifier_free(fitted_random_forest_classifier) - printf("ESTIMATOR|random_forest_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + polynomial_count_sketch_free(fitted_polynomial_count_sketch) + printf("ESTIMATOR|polynomial_count_sketch|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- random_forest_regressor (regression) ---- + # ---- polynomial_features (classification, written out) ---- t0 = flow_now_ns() - let probe_random_forest_regressor: RandomForestRegressor = random_forest_regressor_fit(X_r, y_r, 10, 5, 42) + let probe_polynomial_features: PolynomialFeatures = polynomial_features_fit(f_c, 2, false, true) t1 = flow_now_ns() - random_forest_regressor_free(probe_random_forest_regressor) + polynomial_features_free(probe_polynomial_features) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_random_forest_regressor: RandomForestRegressor = random_forest_regressor_fit(X_r, y_r, 10, 5, 42) - random_forest_regressor_free(m_random_forest_regressor) + let m_polynomial_features: PolynomialFeatures = polynomial_features_fit(f_c, 2, false, true) + polynomial_features_free(m_polynomial_features) } t1 = flow_now_ns() - let fitted_random_forest_regressor: RandomForestRegressor = random_forest_regressor_fit(X_r, y_r, 10, 5, 42) + let fitted_polynomial_features: PolynomialFeatures = polynomial_features_fit(f_c, 2, false, true) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_random_forest_regressor: ptr = random_forest_regressor_predict(fitted_random_forest_regressor, X_r) - array_free_f32(o_random_forest_regressor) + let o_polynomial_features: Matrix = polynomial_features_transform(fitted_polynomial_features, X_c) + if o_polynomial_features.rows > 0 { + if o_polynomial_features.cols > 0 { sink = sink + o_polynomial_features.data[0] } + } + matrix_free(o_polynomial_features) } t3 = flow_now_ns() - random_forest_regressor_free(fitted_random_forest_regressor) - printf("ESTIMATOR|random_forest_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + polynomial_features_free(fitted_polynomial_features) + printf("ESTIMATOR|polynomial_features|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- random_trees_embedding (unsupervised) ---- + # ---- power_transformer (unsupervised) ---- t0 = flow_now_ns() - let probe_random_trees_embedding: RandomTreesEmbedding = random_trees_embedding_fit(X_c, 10, 5, 42) + let probe_power_transformer: PowerTransformer = power_transformer_fit(X_c, 0) t1 = flow_now_ns() - random_trees_embedding_free(probe_random_trees_embedding) + power_transformer_free(probe_power_transformer) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_random_trees_embedding: RandomTreesEmbedding = random_trees_embedding_fit(X_c, 10, 5, 42) - random_trees_embedding_free(m_random_trees_embedding) + let m_power_transformer: PowerTransformer = power_transformer_fit(X_c, 0) + power_transformer_free(m_power_transformer) } t1 = flow_now_ns() - let fitted_random_trees_embedding: RandomTreesEmbedding = random_trees_embedding_fit(X_c, 10, 5, 42) + let fitted_power_transformer: PowerTransformer = power_transformer_fit(X_c, 0) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_random_trees_embedding: Matrix = random_trees_embedding_transform(fitted_random_trees_embedding, X_c) - matrix_free(o_random_trees_embedding) + let o_power_transformer: Matrix = power_transformer_transform(fitted_power_transformer, X_c) + if o_power_transformer.rows > 0 { + if o_power_transformer.cols > 0 { sink = sink + o_power_transformer.data[0] } + } + matrix_free(o_power_transformer) } t3 = flow_now_ns() - random_trees_embedding_free(fitted_random_trees_embedding) - printf("ESTIMATOR|random_trees_embedding|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + power_transformer_free(fitted_power_transformer) + printf("ESTIMATOR|power_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- ransac_regressor (regression) ---- + # ---- qda (classification) ---- t0 = flow_now_ns() - let probe_ransac_regressor: RANSACRegressor = ransac_regressor_fit(X_r, y_r, 5, 10, 1.0, 42) + let probe_qda: QuadraticDiscriminantAnalysis = qda_fit(X_c, y_c, 3) t1 = flow_now_ns() - ransac_regressor_free(probe_ransac_regressor) + qda_free(probe_qda) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_ransac_regressor: RANSACRegressor = ransac_regressor_fit(X_r, y_r, 5, 10, 1.0, 42) - ransac_regressor_free(m_ransac_regressor) + let m_qda: QuadraticDiscriminantAnalysis = qda_fit(X_c, y_c, 3) + qda_free(m_qda) } t1 = flow_now_ns() - let fitted_ransac_regressor: RANSACRegressor = ransac_regressor_fit(X_r, y_r, 5, 10, 1.0, 42) + let fitted_qda: QuadraticDiscriminantAnalysis = qda_fit(X_c, y_c, 3) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_ransac_regressor: ptr = ransac_regressor_predict(fitted_ransac_regressor, X_r) - array_free_f32(o_ransac_regressor) + let o_qda: ptr = qda_predict(fitted_qda, X_c) + sink = sink + o_qda[0] + array_free_f32(o_qda) } t3 = flow_now_ns() - ransac_regressor_free(fitted_ransac_regressor) - printf("ESTIMATOR|ransac_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + qda_free(fitted_qda) + printf("ESTIMATOR|qda|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- regressor_chain (multioutput_class) ---- + # ---- quantile_regressor (regression) ---- t0 = flow_now_ns() - let probe_regressor_chain: RegressorChain = regressor_chain_fit(X_c, Y_label_rows, n_c, f_c, 2) + let probe_quantile_regressor: QuantileRegressor = quantile_regressor_fit(X_r, y_r, 0.5, 1.0, 50, 0.01) t1 = flow_now_ns() - regressor_chain_free(probe_regressor_chain) + quantile_regressor_free(probe_quantile_regressor) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_regressor_chain: RegressorChain = regressor_chain_fit(X_c, Y_label_rows, n_c, f_c, 2) - regressor_chain_free(m_regressor_chain) + let m_quantile_regressor: QuantileRegressor = quantile_regressor_fit(X_r, y_r, 0.5, 1.0, 50, 0.01) + quantile_regressor_free(m_quantile_regressor) } t1 = flow_now_ns() - printf("ESTIMATOR|regressor_chain|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + let fitted_quantile_regressor: QuantileRegressor = quantile_regressor_fit(X_r, y_r, 0.5, 1.0, 50, 0.01) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_quantile_regressor: ptr = quantile_regressor_predict(fitted_quantile_regressor, X_r) + sink = sink + o_quantile_regressor[0] + array_free_f32(o_quantile_regressor) + } + t3 = flow_now_ns() + quantile_regressor_free(fitted_quantile_regressor) + printf("ESTIMATOR|quantile_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- rfe (regression) ---- + # ---- quantile_transformer (unsupervised) ---- t0 = flow_now_ns() - let probe_rfe: RFE = rfe_fit(X_r, y_r, 2, null, n_r, f_r) + let probe_quantile_transformer: QuantileTransformer = quantile_transformer_fit(X_c, 10, 0) t1 = flow_now_ns() - rfe_free(probe_rfe) + quantile_transformer_free(probe_quantile_transformer) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_rfe: RFE = rfe_fit(X_r, y_r, 2, null, n_r, f_r) - rfe_free(m_rfe) + let m_quantile_transformer: QuantileTransformer = quantile_transformer_fit(X_c, 10, 0) + quantile_transformer_free(m_quantile_transformer) } t1 = flow_now_ns() - let fitted_rfe: RFE = rfe_fit(X_r, y_r, 2, null, n_r, f_r) + let fitted_quantile_transformer: QuantileTransformer = quantile_transformer_fit(X_c, 10, 0) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_rfe: Matrix = rfe_transform(fitted_rfe, X_r) - matrix_free(o_rfe) + let o_quantile_transformer: Matrix = quantile_transformer_transform(fitted_quantile_transformer, X_c) + if o_quantile_transformer.rows > 0 { + if o_quantile_transformer.cols > 0 { sink = sink + o_quantile_transformer.data[0] } + } + matrix_free(o_quantile_transformer) } t3 = flow_now_ns() - rfe_free(fitted_rfe) - printf("ESTIMATOR|rfe|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + quantile_transformer_free(fitted_quantile_transformer) + printf("ESTIMATOR|quantile_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- rfecv (regression) ---- + # ---- radius_neighbors_classifier (classification) ---- t0 = flow_now_ns() - let probe_rfecv: RFECV = rfecv_fit(X_r, y_r, n_r, f_r, 3, null) + let probe_radius_neighbors_classifier: RadiusNeighborsClassifier = radius_neighbors_classifier_fit(X_c, y_c, 1.0, 3, 0.0) t1 = flow_now_ns() - rfecv_free(probe_rfecv) + radius_neighbors_classifier_free(probe_radius_neighbors_classifier) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_rfecv: RFECV = rfecv_fit(X_r, y_r, n_r, f_r, 3, null) - rfecv_free(m_rfecv) + let m_radius_neighbors_classifier: RadiusNeighborsClassifier = radius_neighbors_classifier_fit(X_c, y_c, 1.0, 3, 0.0) + radius_neighbors_classifier_free(m_radius_neighbors_classifier) } t1 = flow_now_ns() - printf("ESTIMATOR|rfecv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + let fitted_radius_neighbors_classifier: RadiusNeighborsClassifier = radius_neighbors_classifier_fit(X_c, y_c, 1.0, 3, 0.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_radius_neighbors_classifier: ptr = radius_neighbors_classifier_predict(fitted_radius_neighbors_classifier, X_c) + sink = sink + o_radius_neighbors_classifier[0] + array_free_f32(o_radius_neighbors_classifier) + } + t3 = flow_now_ns() + radius_neighbors_classifier_free(fitted_radius_neighbors_classifier) + printf("ESTIMATOR|radius_neighbors_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- ridge_classifier_cv (classification) ---- + # ---- radius_neighbors_regressor (regression) ---- t0 = flow_now_ns() - let probe_ridge_classifier_cv: RidgeClassifierCV = ridge_classifier_cv_fit(X_c, y_c, 3, 10) + let probe_radius_neighbors_regressor: RadiusNeighborsRegressor = radius_neighbors_regressor_fit(X_r, y_r, 1.0) t1 = flow_now_ns() - ridge_classifier_cv_free(probe_ridge_classifier_cv) + radius_neighbors_regressor_free(probe_radius_neighbors_regressor) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_ridge_classifier_cv: RidgeClassifierCV = ridge_classifier_cv_fit(X_c, y_c, 3, 10) - ridge_classifier_cv_free(m_ridge_classifier_cv) + let m_radius_neighbors_regressor: RadiusNeighborsRegressor = radius_neighbors_regressor_fit(X_r, y_r, 1.0) + radius_neighbors_regressor_free(m_radius_neighbors_regressor) } t1 = flow_now_ns() - let fitted_ridge_classifier_cv: RidgeClassifierCV = ridge_classifier_cv_fit(X_c, y_c, 3, 10) + let fitted_radius_neighbors_regressor: RadiusNeighborsRegressor = radius_neighbors_regressor_fit(X_r, y_r, 1.0) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_ridge_classifier_cv: ptr = ridge_classifier_cv_predict(fitted_ridge_classifier_cv, X_c) - array_free_f32(o_ridge_classifier_cv) + let o_radius_neighbors_regressor: ptr = radius_neighbors_regressor_predict(fitted_radius_neighbors_regressor, X_r) + sink = sink + o_radius_neighbors_regressor[0] + array_free_f32(o_radius_neighbors_regressor) } t3 = flow_now_ns() - ridge_classifier_cv_free(fitted_ridge_classifier_cv) - printf("ESTIMATOR|ridge_classifier_cv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + radius_neighbors_regressor_free(fitted_radius_neighbors_regressor) + printf("ESTIMATOR|radius_neighbors_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- ridge_classifier (regression) ---- + # ---- radius_neighbors_transformer (unsupervised) ---- t0 = flow_now_ns() - let probe_ridge_classifier: RidgeClassifier = ridge_classifier_fit(X_r, y_r, 1.0, 50, 0.01) + let probe_radius_neighbors_transformer: RadiusNeighborsTransformer = radius_neighbors_transformer_fit(X_c, 1.0, 0) t1 = flow_now_ns() - ridge_classifier_free(probe_ridge_classifier) + radius_neighbors_transformer_free(probe_radius_neighbors_transformer) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_ridge_classifier: RidgeClassifier = ridge_classifier_fit(X_r, y_r, 1.0, 50, 0.01) - ridge_classifier_free(m_ridge_classifier) + let m_radius_neighbors_transformer: RadiusNeighborsTransformer = radius_neighbors_transformer_fit(X_c, 1.0, 0) + radius_neighbors_transformer_free(m_radius_neighbors_transformer) } t1 = flow_now_ns() - let fitted_ridge_classifier: RidgeClassifier = ridge_classifier_fit(X_r, y_r, 1.0, 50, 0.01) + let fitted_radius_neighbors_transformer: RadiusNeighborsTransformer = radius_neighbors_transformer_fit(X_c, 1.0, 0) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_ridge_classifier: ptr = ridge_classifier_predict(fitted_ridge_classifier, X_r) - array_free_f32(o_ridge_classifier) + let o_radius_neighbors_transformer: Matrix = radius_neighbors_transformer_transform(fitted_radius_neighbors_transformer, X_c) + if o_radius_neighbors_transformer.rows > 0 { + if o_radius_neighbors_transformer.cols > 0 { sink = sink + o_radius_neighbors_transformer.data[0] } + } + matrix_free(o_radius_neighbors_transformer) } t3 = flow_now_ns() - ridge_classifier_free(fitted_ridge_classifier) - printf("ESTIMATOR|ridge_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + radius_neighbors_transformer_free(fitted_radius_neighbors_transformer) + printf("ESTIMATOR|radius_neighbors_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } @@ -563,5 +630,10 @@ function main() -> i32 { matrix_free(Y_multi) free(yi_c as ptr) free(yi_r as ptr) + array_free_f32(x1d_r) + array_free_f32(w_f) + # The sink is printed so the work above cannot be optimized away. The + # parser matches ESTIMATOR lines only, so this one is ignored. + printf("SINK|%.9f\n", sink) return 0 } diff --git a/benchmarks/generated/bench_estimators_07.flow b/benchmarks/generated/bench_estimators_07.flow index 17e83a2..cbba33a 100644 --- a/benchmarks/generated/bench_estimators_07.flow +++ b/benchmarks/generated/bench_estimators_07.flow @@ -65,470 +65,545 @@ function main() -> i32 { Y_rows[i] = row } + # A one-dimensional x for the isotonic row, which regresses against a + # single ordered variable rather than a design. + let x1d_r: ptr = array_new_f32(n_r) + for i in 0 to n_r { x1d_r[i] = matrix_at(X_r, i, 0) } + + # Per-feature importances for the selector row, which takes the weights a + # fitted model would hand it rather than a design. + let w_f: ptr = array_new_f32(f_c) + for i in 0 to f_c { w_f[i] = 1.0 / ((i + 1) as f32) } + let mut t0: i64 = 0 let mut t1: i64 = 0 let mut t2: i64 = 0 let mut t3: i64 = 0 let mut reps: i32 = 1 + let mut sink: f32 = 0.0 - # ---- ridge_cv (regression) ---- + # ---- random_forest_classifier (classification) ---- t0 = flow_now_ns() - let probe_ridge_cv: RidgeCV = ridge_cv_fit(X_r, y_r, 10) + let probe_random_forest_classifier: RandomForestClassifier = random_forest_classifier_fit(X_c, y_c, 3, 10, 5, 42) t1 = flow_now_ns() - ridge_cv_free(probe_ridge_cv) + random_forest_classifier_free(probe_random_forest_classifier) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_ridge_cv: RidgeCV = ridge_cv_fit(X_r, y_r, 10) - ridge_cv_free(m_ridge_cv) + let m_random_forest_classifier: RandomForestClassifier = random_forest_classifier_fit(X_c, y_c, 3, 10, 5, 42) + random_forest_classifier_free(m_random_forest_classifier) } t1 = flow_now_ns() - let fitted_ridge_cv: RidgeCV = ridge_cv_fit(X_r, y_r, 10) + let fitted_random_forest_classifier: RandomForestClassifier = random_forest_classifier_fit(X_c, y_c, 3, 10, 5, 42) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_ridge_cv: ptr = ridge_cv_predict(fitted_ridge_cv, X_r) - array_free_f32(o_ridge_cv) + let o_random_forest_classifier: ptr = random_forest_classifier_predict(fitted_random_forest_classifier, X_c) + sink = sink + o_random_forest_classifier[0] + array_free_f32(o_random_forest_classifier) } t3 = flow_now_ns() - ridge_cv_free(fitted_ridge_cv) - printf("ESTIMATOR|ridge_cv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + random_forest_classifier_free(fitted_random_forest_classifier) + printf("ESTIMATOR|random_forest_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- ridge (regression) ---- + # ---- random_forest_regressor (regression) ---- t0 = flow_now_ns() - let probe_ridge: Ridge = ridge_fit(X_r, y_r, 1.0, 50, 0.01) + let probe_random_forest_regressor: RandomForestRegressor = random_forest_regressor_fit(X_r, y_r, 10, 5, 42) t1 = flow_now_ns() - ridge_free(probe_ridge) + random_forest_regressor_free(probe_random_forest_regressor) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_ridge: Ridge = ridge_fit(X_r, y_r, 1.0, 50, 0.01) - ridge_free(m_ridge) + let m_random_forest_regressor: RandomForestRegressor = random_forest_regressor_fit(X_r, y_r, 10, 5, 42) + random_forest_regressor_free(m_random_forest_regressor) } t1 = flow_now_ns() - let fitted_ridge: Ridge = ridge_fit(X_r, y_r, 1.0, 50, 0.01) + let fitted_random_forest_regressor: RandomForestRegressor = random_forest_regressor_fit(X_r, y_r, 10, 5, 42) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_ridge: ptr = ridge_predict(fitted_ridge, X_r) - array_free_f32(o_ridge) + let o_random_forest_regressor: ptr = random_forest_regressor_predict(fitted_random_forest_regressor, X_r) + sink = sink + o_random_forest_regressor[0] + array_free_f32(o_random_forest_regressor) } t3 = flow_now_ns() - ridge_free(fitted_ridge) - printf("ESTIMATOR|ridge|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + random_forest_regressor_free(fitted_random_forest_regressor) + printf("ESTIMATOR|random_forest_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- robust_scaler (unsupervised) ---- + # ---- random_trees_embedding (unsupervised) ---- t0 = flow_now_ns() - let probe_robust_scaler: RobustScaler = robust_scaler_fit(X_c) + let probe_random_trees_embedding: RandomTreesEmbedding = random_trees_embedding_fit(X_c, 10, 5, 42) t1 = flow_now_ns() - robust_scaler_free(probe_robust_scaler) + random_trees_embedding_free(probe_random_trees_embedding) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_robust_scaler: RobustScaler = robust_scaler_fit(X_c) - robust_scaler_free(m_robust_scaler) + let m_random_trees_embedding: RandomTreesEmbedding = random_trees_embedding_fit(X_c, 10, 5, 42) + random_trees_embedding_free(m_random_trees_embedding) } t1 = flow_now_ns() - let fitted_robust_scaler: RobustScaler = robust_scaler_fit(X_c) + let fitted_random_trees_embedding: RandomTreesEmbedding = random_trees_embedding_fit(X_c, 10, 5, 42) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_robust_scaler: Matrix = robust_scaler_transform(fitted_robust_scaler, X_c) - matrix_free(o_robust_scaler) + let o_random_trees_embedding: Matrix = random_trees_embedding_transform(fitted_random_trees_embedding, X_c) + if o_random_trees_embedding.rows > 0 { + if o_random_trees_embedding.cols > 0 { sink = sink + o_random_trees_embedding.data[0] } + } + matrix_free(o_random_trees_embedding) } t3 = flow_now_ns() - robust_scaler_free(fitted_robust_scaler) - printf("ESTIMATOR|robust_scaler|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + random_trees_embedding_free(fitted_random_trees_embedding) + printf("ESTIMATOR|random_trees_embedding|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- select_fdr (regression) ---- + # ---- ransac_regressor (regression) ---- t0 = flow_now_ns() - let probe_select_fdr: SelectFdr = select_fdr_fit(X_r, y_r, 1.0) + let probe_ransac_regressor: RANSACRegressor = ransac_regressor_fit(X_r, y_r, 5, 10, 1.0, 42) t1 = flow_now_ns() - select_fdr_free(probe_select_fdr) + ransac_regressor_free(probe_ransac_regressor) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_select_fdr: SelectFdr = select_fdr_fit(X_r, y_r, 1.0) - select_fdr_free(m_select_fdr) + let m_ransac_regressor: RANSACRegressor = ransac_regressor_fit(X_r, y_r, 5, 10, 1.0, 42) + ransac_regressor_free(m_ransac_regressor) } t1 = flow_now_ns() - let fitted_select_fdr: SelectFdr = select_fdr_fit(X_r, y_r, 1.0) + let fitted_ransac_regressor: RANSACRegressor = ransac_regressor_fit(X_r, y_r, 5, 10, 1.0, 42) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_select_fdr: Matrix = select_fdr_transform(fitted_select_fdr, X_r) - matrix_free(o_select_fdr) + let o_ransac_regressor: ptr = ransac_regressor_predict(fitted_ransac_regressor, X_r) + sink = sink + o_ransac_regressor[0] + array_free_f32(o_ransac_regressor) } t3 = flow_now_ns() - select_fdr_free(fitted_select_fdr) - printf("ESTIMATOR|select_fdr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + ransac_regressor_free(fitted_ransac_regressor) + printf("ESTIMATOR|ransac_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- select_fpr (regression) ---- + # ---- rbf_sampler (classification, written out) ---- t0 = flow_now_ns() - let probe_select_fpr: SelectFpr = select_fpr_fit(X_r, y_r, 1.0) + let probe_rbf_sampler: RBFSampler = rbf_sampler_fit(f_c, 0.1, 2, 42) t1 = flow_now_ns() - select_fpr_free(probe_select_fpr) + rbf_sampler_free(probe_rbf_sampler) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_select_fpr: SelectFpr = select_fpr_fit(X_r, y_r, 1.0) - select_fpr_free(m_select_fpr) + let m_rbf_sampler: RBFSampler = rbf_sampler_fit(f_c, 0.1, 2, 42) + rbf_sampler_free(m_rbf_sampler) } t1 = flow_now_ns() - let fitted_select_fpr: SelectFpr = select_fpr_fit(X_r, y_r, 1.0) + let fitted_rbf_sampler: RBFSampler = rbf_sampler_fit(f_c, 0.1, 2, 42) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_select_fpr: Matrix = select_fpr_transform(fitted_select_fpr, X_r) - matrix_free(o_select_fpr) + let o_rbf_sampler: Matrix = rbf_sampler_transform(fitted_rbf_sampler, X_c) + if o_rbf_sampler.rows > 0 { + if o_rbf_sampler.cols > 0 { sink = sink + o_rbf_sampler.data[0] } + } + matrix_free(o_rbf_sampler) } t3 = flow_now_ns() - select_fpr_free(fitted_select_fpr) - printf("ESTIMATOR|select_fpr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + rbf_sampler_free(fitted_rbf_sampler) + printf("ESTIMATOR|rbf_sampler|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- select_fwe (regression) ---- + # ---- regressor_chain (multioutput_class) ---- t0 = flow_now_ns() - let probe_select_fwe: SelectFwe = select_fwe_fit(X_r, y_r, 1.0) + let probe_regressor_chain: RegressorChain = regressor_chain_fit(X_c, Y_label_rows, n_c, f_c, 2) t1 = flow_now_ns() - select_fwe_free(probe_select_fwe) + regressor_chain_free(probe_regressor_chain) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_select_fwe: SelectFwe = select_fwe_fit(X_r, y_r, 1.0) - select_fwe_free(m_select_fwe) + let m_regressor_chain: RegressorChain = regressor_chain_fit(X_c, Y_label_rows, n_c, f_c, 2) + regressor_chain_free(m_regressor_chain) } t1 = flow_now_ns() - let fitted_select_fwe: SelectFwe = select_fwe_fit(X_r, y_r, 1.0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_select_fwe: Matrix = select_fwe_transform(fitted_select_fwe, X_r) - matrix_free(o_select_fwe) - } - t3 = flow_now_ns() - select_fwe_free(fitted_select_fwe) - printf("ESTIMATOR|select_fwe|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + printf("ESTIMATOR|regressor_chain|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) fflush(null) - # ---- select_k_best (regression) ---- + # ---- rfe (regression) ---- t0 = flow_now_ns() - let probe_select_k_best: SelectKBest = select_k_best_fit(X_r, y_r, 3, 0) + let probe_rfe: RFE = rfe_fit(X_r, y_r, 2, null, n_r, f_r) t1 = flow_now_ns() - select_k_best_free(probe_select_k_best) + rfe_free(probe_rfe) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_select_k_best: SelectKBest = select_k_best_fit(X_r, y_r, 3, 0) - select_k_best_free(m_select_k_best) + let m_rfe: RFE = rfe_fit(X_r, y_r, 2, null, n_r, f_r) + rfe_free(m_rfe) } t1 = flow_now_ns() - let fitted_select_k_best: SelectKBest = select_k_best_fit(X_r, y_r, 3, 0) + let fitted_rfe: RFE = rfe_fit(X_r, y_r, 2, null, n_r, f_r) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_select_k_best: Matrix = select_k_best_transform(fitted_select_k_best, X_r) - matrix_free(o_select_k_best) + let o_rfe: Matrix = rfe_transform(fitted_rfe, X_r) + if o_rfe.rows > 0 { + if o_rfe.cols > 0 { sink = sink + o_rfe.data[0] } + } + matrix_free(o_rfe) } t3 = flow_now_ns() - select_k_best_free(fitted_select_k_best) - printf("ESTIMATOR|select_k_best|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + rfe_free(fitted_rfe) + printf("ESTIMATOR|rfe|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- select_percentile (regression) ---- + # ---- rfecv (regression) ---- t0 = flow_now_ns() - let probe_select_percentile: SelectPercentile = select_percentile_fit(X_r, y_r, 50) + let probe_rfecv: RFECV = rfecv_fit(X_r, y_r, n_r, f_r, 3, null) t1 = flow_now_ns() - select_percentile_free(probe_select_percentile) + rfecv_free(probe_rfecv) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_select_percentile: SelectPercentile = select_percentile_fit(X_r, y_r, 50) - select_percentile_free(m_select_percentile) + let m_rfecv: RFECV = rfecv_fit(X_r, y_r, n_r, f_r, 3, null) + rfecv_free(m_rfecv) } t1 = flow_now_ns() - let fitted_select_percentile: SelectPercentile = select_percentile_fit(X_r, y_r, 50) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_select_percentile: Matrix = select_percentile_transform(fitted_select_percentile, X_r) - matrix_free(o_select_percentile) - } - t3 = flow_now_ns() - select_percentile_free(fitted_select_percentile) - printf("ESTIMATOR|select_percentile|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + printf("ESTIMATOR|rfecv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) fflush(null) - # ---- sequential_feature_selector (regression) ---- + # ---- ridge_classifier_cv (classification) ---- t0 = flow_now_ns() - let probe_sequential_feature_selector: SequentialFeatureSelector = sequential_feature_selector_fit(X_r, y_r, 2, 0, n_r, f_r) + let probe_ridge_classifier_cv: RidgeClassifierCV = ridge_classifier_cv_fit(X_c, y_c, 3, 10) t1 = flow_now_ns() - sequential_feature_selector_free(probe_sequential_feature_selector) + ridge_classifier_cv_free(probe_ridge_classifier_cv) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_sequential_feature_selector: SequentialFeatureSelector = sequential_feature_selector_fit(X_r, y_r, 2, 0, n_r, f_r) - sequential_feature_selector_free(m_sequential_feature_selector) + let m_ridge_classifier_cv: RidgeClassifierCV = ridge_classifier_cv_fit(X_c, y_c, 3, 10) + ridge_classifier_cv_free(m_ridge_classifier_cv) } t1 = flow_now_ns() - let fitted_sequential_feature_selector: SequentialFeatureSelector = sequential_feature_selector_fit(X_r, y_r, 2, 0, n_r, f_r) + let fitted_ridge_classifier_cv: RidgeClassifierCV = ridge_classifier_cv_fit(X_c, y_c, 3, 10) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_sequential_feature_selector: Matrix = sequential_feature_selector_transform(fitted_sequential_feature_selector, X_r) - matrix_free(o_sequential_feature_selector) + let o_ridge_classifier_cv: ptr = ridge_classifier_cv_predict(fitted_ridge_classifier_cv, X_c) + sink = sink + o_ridge_classifier_cv[0] + array_free_f32(o_ridge_classifier_cv) } t3 = flow_now_ns() - sequential_feature_selector_free(fitted_sequential_feature_selector) - printf("ESTIMATOR|sequential_feature_selector|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + ridge_classifier_cv_free(fitted_ridge_classifier_cv) + printf("ESTIMATOR|ridge_classifier_cv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- sgd_classifier (classification) ---- + # ---- ridge_classifier (regression) ---- t0 = flow_now_ns() - let probe_sgd_classifier: SGDClassifier = sgd_classifier_fit(X_c, y_c, 3, 0, 1.0, 50, 0.01) + let probe_ridge_classifier: RidgeClassifier = ridge_classifier_fit(X_r, y_r, 1.0, 50, 0.01) t1 = flow_now_ns() - sgd_classifier_free(probe_sgd_classifier) + ridge_classifier_free(probe_ridge_classifier) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_sgd_classifier: SGDClassifier = sgd_classifier_fit(X_c, y_c, 3, 0, 1.0, 50, 0.01) - sgd_classifier_free(m_sgd_classifier) + let m_ridge_classifier: RidgeClassifier = ridge_classifier_fit(X_r, y_r, 1.0, 50, 0.01) + ridge_classifier_free(m_ridge_classifier) } t1 = flow_now_ns() - let fitted_sgd_classifier: SGDClassifier = sgd_classifier_fit(X_c, y_c, 3, 0, 1.0, 50, 0.01) + let fitted_ridge_classifier: RidgeClassifier = ridge_classifier_fit(X_r, y_r, 1.0, 50, 0.01) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_sgd_classifier: ptr = sgd_classifier_predict(fitted_sgd_classifier, X_c) - array_free_f32(o_sgd_classifier) + let o_ridge_classifier: ptr = ridge_classifier_predict(fitted_ridge_classifier, X_r) + sink = sink + o_ridge_classifier[0] + array_free_f32(o_ridge_classifier) } t3 = flow_now_ns() - sgd_classifier_free(fitted_sgd_classifier) - printf("ESTIMATOR|sgd_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + ridge_classifier_free(fitted_ridge_classifier) + printf("ESTIMATOR|ridge_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- sgd_one_class_svm (unsupervised) ---- + # ---- ridge_cv (regression) ---- t0 = flow_now_ns() - let probe_sgd_one_class_svm: SGDOneClassSVM = sgd_one_class_svm_fit(X_c, 0.5, 50, 0.01) + let probe_ridge_cv: RidgeCV = ridge_cv_fit(X_r, y_r, 10) t1 = flow_now_ns() - sgd_one_class_svm_free(probe_sgd_one_class_svm) + ridge_cv_free(probe_ridge_cv) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_sgd_one_class_svm: SGDOneClassSVM = sgd_one_class_svm_fit(X_c, 0.5, 50, 0.01) - sgd_one_class_svm_free(m_sgd_one_class_svm) + let m_ridge_cv: RidgeCV = ridge_cv_fit(X_r, y_r, 10) + ridge_cv_free(m_ridge_cv) } t1 = flow_now_ns() - let fitted_sgd_one_class_svm: SGDOneClassSVM = sgd_one_class_svm_fit(X_c, 0.5, 50, 0.01) + let fitted_ridge_cv: RidgeCV = ridge_cv_fit(X_r, y_r, 10) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_sgd_one_class_svm: ptr = sgd_one_class_svm_predict(fitted_sgd_one_class_svm, X_c) - array_free_f32(o_sgd_one_class_svm) + let o_ridge_cv: ptr = ridge_cv_predict(fitted_ridge_cv, X_r) + sink = sink + o_ridge_cv[0] + array_free_f32(o_ridge_cv) } t3 = flow_now_ns() - sgd_one_class_svm_free(fitted_sgd_one_class_svm) - printf("ESTIMATOR|sgd_one_class_svm|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + ridge_cv_free(fitted_ridge_cv) + printf("ESTIMATOR|ridge_cv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- sgd_regressor (regression) ---- + # ---- ridge (regression) ---- t0 = flow_now_ns() - let probe_sgd_regressor: SGDRegressor = sgd_regressor_fit(X_r, y_r, 0, 1.0, 0.1, 50, 0.01) + let probe_ridge: Ridge = ridge_fit(X_r, y_r, 1.0, 50, 0.01) t1 = flow_now_ns() - sgd_regressor_free(probe_sgd_regressor) + ridge_free(probe_ridge) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_sgd_regressor: SGDRegressor = sgd_regressor_fit(X_r, y_r, 0, 1.0, 0.1, 50, 0.01) - sgd_regressor_free(m_sgd_regressor) + let m_ridge: Ridge = ridge_fit(X_r, y_r, 1.0, 50, 0.01) + ridge_free(m_ridge) } t1 = flow_now_ns() - let fitted_sgd_regressor: SGDRegressor = sgd_regressor_fit(X_r, y_r, 0, 1.0, 0.1, 50, 0.01) + let fitted_ridge: Ridge = ridge_fit(X_r, y_r, 1.0, 50, 0.01) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_sgd_regressor: ptr = sgd_regressor_predict(fitted_sgd_regressor, X_r) - array_free_f32(o_sgd_regressor) + let o_ridge: ptr = ridge_predict(fitted_ridge, X_r) + sink = sink + o_ridge[0] + array_free_f32(o_ridge) } t3 = flow_now_ns() - sgd_regressor_free(fitted_sgd_regressor) - printf("ESTIMATOR|sgd_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + ridge_free(fitted_ridge) + printf("ESTIMATOR|ridge|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- shrunk_covariance (unsupervised) ---- + # ---- robust_scaler (unsupervised) ---- t0 = flow_now_ns() - let probe_shrunk_covariance: ShrunkCovariance = shrunk_covariance_fit(X_c, 0.1) + let probe_robust_scaler: RobustScaler = robust_scaler_fit(X_c) t1 = flow_now_ns() - shrunk_covariance_free(probe_shrunk_covariance) + robust_scaler_free(probe_robust_scaler) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_shrunk_covariance: ShrunkCovariance = shrunk_covariance_fit(X_c, 0.1) - shrunk_covariance_free(m_shrunk_covariance) + let m_robust_scaler: RobustScaler = robust_scaler_fit(X_c) + robust_scaler_free(m_robust_scaler) } t1 = flow_now_ns() - printf("ESTIMATOR|shrunk_covariance|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + let fitted_robust_scaler: RobustScaler = robust_scaler_fit(X_c) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_robust_scaler: Matrix = robust_scaler_transform(fitted_robust_scaler, X_c) + if o_robust_scaler.rows > 0 { + if o_robust_scaler.cols > 0 { sink = sink + o_robust_scaler.data[0] } + } + matrix_free(o_robust_scaler) + } + t3 = flow_now_ns() + robust_scaler_free(fitted_robust_scaler) + printf("ESTIMATOR|robust_scaler|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- simple_imputer (unsupervised) ---- + # ---- select_fdr (regression) ---- t0 = flow_now_ns() - let probe_simple_imputer: SimpleImputer = simple_imputer_fit(X_c, 0, 1.0) + let probe_select_fdr: SelectFdr = select_fdr_fit(X_r, y_r, 1.0) t1 = flow_now_ns() - simple_imputer_free(probe_simple_imputer) + select_fdr_free(probe_select_fdr) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_simple_imputer: SimpleImputer = simple_imputer_fit(X_c, 0, 1.0) - simple_imputer_free(m_simple_imputer) + let m_select_fdr: SelectFdr = select_fdr_fit(X_r, y_r, 1.0) + select_fdr_free(m_select_fdr) } t1 = flow_now_ns() - let fitted_simple_imputer: SimpleImputer = simple_imputer_fit(X_c, 0, 1.0) + let fitted_select_fdr: SelectFdr = select_fdr_fit(X_r, y_r, 1.0) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_simple_imputer: Matrix = simple_imputer_transform(fitted_simple_imputer, X_c) - matrix_free(o_simple_imputer) + let o_select_fdr: Matrix = select_fdr_transform(fitted_select_fdr, X_r) + if o_select_fdr.rows > 0 { + if o_select_fdr.cols > 0 { sink = sink + o_select_fdr.data[0] } + } + matrix_free(o_select_fdr) } t3 = flow_now_ns() - simple_imputer_free(fitted_simple_imputer) - printf("ESTIMATOR|simple_imputer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + select_fdr_free(fitted_select_fdr) + printf("ESTIMATOR|select_fdr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- sparse_coder (unsupervised) ---- + # ---- select_fpr (regression) ---- t0 = flow_now_ns() - let probe_sparse_coder: SparseCoder = sparse_coder_fit(X_c, 3) + let probe_select_fpr: SelectFpr = select_fpr_fit(X_r, y_r, 1.0) t1 = flow_now_ns() - sparse_coder_free(probe_sparse_coder) + select_fpr_free(probe_select_fpr) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_sparse_coder: SparseCoder = sparse_coder_fit(X_c, 3) - sparse_coder_free(m_sparse_coder) + let m_select_fpr: SelectFpr = select_fpr_fit(X_r, y_r, 1.0) + select_fpr_free(m_select_fpr) } t1 = flow_now_ns() - let fitted_sparse_coder: SparseCoder = sparse_coder_fit(X_c, 3) + let fitted_select_fpr: SelectFpr = select_fpr_fit(X_r, y_r, 1.0) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_sparse_coder: Matrix = sparse_coder_transform(fitted_sparse_coder, X_c) - matrix_free(o_sparse_coder) + let o_select_fpr: Matrix = select_fpr_transform(fitted_select_fpr, X_r) + if o_select_fpr.rows > 0 { + if o_select_fpr.cols > 0 { sink = sink + o_select_fpr.data[0] } + } + matrix_free(o_select_fpr) } t3 = flow_now_ns() - sparse_coder_free(fitted_sparse_coder) - printf("ESTIMATOR|sparse_coder|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + select_fpr_free(fitted_select_fpr) + printf("ESTIMATOR|select_fpr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- sparse_pca (unsupervised) ---- + # ---- select_from_model (classification, written out) ---- t0 = flow_now_ns() - let probe_sparse_pca: SparsePCA = sparse_pca_fit(X_c, 2, 1.0, 100, 0.0001, 42) + let probe_select_from_model: SelectFromModel = select_from_model_fit(w_f, f_c, 0.5) t1 = flow_now_ns() - sparse_pca_free(probe_sparse_pca) + select_from_model_free(probe_select_from_model) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_sparse_pca: SparsePCA = sparse_pca_fit(X_c, 2, 1.0, 100, 0.0001, 42) - sparse_pca_free(m_sparse_pca) + let m_select_from_model: SelectFromModel = select_from_model_fit(w_f, f_c, 0.5) + select_from_model_free(m_select_from_model) } t1 = flow_now_ns() - let fitted_sparse_pca: SparsePCA = sparse_pca_fit(X_c, 2, 1.0, 100, 0.0001, 42) + let fitted_select_from_model: SelectFromModel = select_from_model_fit(w_f, f_c, 0.5) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_sparse_pca: Matrix = sparse_pca_transform(fitted_sparse_pca, X_c) - matrix_free(o_sparse_pca) + let o_select_from_model: Matrix = select_from_model_transform(fitted_select_from_model, X_c) + if o_select_from_model.rows > 0 { + if o_select_from_model.cols > 0 { sink = sink + o_select_from_model.data[0] } + } + matrix_free(o_select_from_model) } t3 = flow_now_ns() - sparse_pca_free(fitted_sparse_pca) - printf("ESTIMATOR|sparse_pca|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + select_from_model_free(fitted_select_from_model) + printf("ESTIMATOR|select_from_model|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- spectral_biclustering (unsupervised) ---- + # ---- select_fwe (regression) ---- t0 = flow_now_ns() - let probe_spectral_biclustering: SpectralBiclustering = spectral_biclustering_fit(X_c, 2, 2) + let probe_select_fwe: SelectFwe = select_fwe_fit(X_r, y_r, 1.0) t1 = flow_now_ns() - spectral_biclustering_free(probe_spectral_biclustering) + select_fwe_free(probe_select_fwe) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_spectral_biclustering: SpectralBiclustering = spectral_biclustering_fit(X_c, 2, 2) - spectral_biclustering_free(m_spectral_biclustering) + let m_select_fwe: SelectFwe = select_fwe_fit(X_r, y_r, 1.0) + select_fwe_free(m_select_fwe) } t1 = flow_now_ns() - printf("ESTIMATOR|spectral_biclustering|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + let fitted_select_fwe: SelectFwe = select_fwe_fit(X_r, y_r, 1.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_select_fwe: Matrix = select_fwe_transform(fitted_select_fwe, X_r) + if o_select_fwe.rows > 0 { + if o_select_fwe.cols > 0 { sink = sink + o_select_fwe.data[0] } + } + matrix_free(o_select_fwe) + } + t3 = flow_now_ns() + select_fwe_free(fitted_select_fwe) + printf("ESTIMATOR|select_fwe|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- spectral_clustering (unsupervised) ---- + # ---- select_k_best (regression) ---- t0 = flow_now_ns() - let probe_spectral_clustering: SpectralClustering = spectral_clustering_fit(X_c, 3, 0.1, 42) + let probe_select_k_best: SelectKBest = select_k_best_fit(X_r, y_r, 3, 0) t1 = flow_now_ns() - spectral_clustering_free(probe_spectral_clustering) + select_k_best_free(probe_select_k_best) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_spectral_clustering: SpectralClustering = spectral_clustering_fit(X_c, 3, 0.1, 42) - spectral_clustering_free(m_spectral_clustering) + let m_select_k_best: SelectKBest = select_k_best_fit(X_r, y_r, 3, 0) + select_k_best_free(m_select_k_best) } t1 = flow_now_ns() - printf("ESTIMATOR|spectral_clustering|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + let fitted_select_k_best: SelectKBest = select_k_best_fit(X_r, y_r, 3, 0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_select_k_best: Matrix = select_k_best_transform(fitted_select_k_best, X_r) + if o_select_k_best.rows > 0 { + if o_select_k_best.cols > 0 { sink = sink + o_select_k_best.data[0] } + } + matrix_free(o_select_k_best) + } + t3 = flow_now_ns() + select_k_best_free(fitted_select_k_best) + printf("ESTIMATOR|select_k_best|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- spectral_coclustering (unsupervised) ---- + # ---- select_percentile (regression) ---- t0 = flow_now_ns() - let probe_spectral_coclustering: SpectralCoclustering = spectral_coclustering_fit(X_c, 3) + let probe_select_percentile: SelectPercentile = select_percentile_fit(X_r, y_r, 50) t1 = flow_now_ns() - spectral_coclustering_free(probe_spectral_coclustering) + select_percentile_free(probe_select_percentile) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_spectral_coclustering: SpectralCoclustering = spectral_coclustering_fit(X_c, 3) - spectral_coclustering_free(m_spectral_coclustering) + let m_select_percentile: SelectPercentile = select_percentile_fit(X_r, y_r, 50) + select_percentile_free(m_select_percentile) } t1 = flow_now_ns() - printf("ESTIMATOR|spectral_coclustering|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + let fitted_select_percentile: SelectPercentile = select_percentile_fit(X_r, y_r, 50) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_select_percentile: Matrix = select_percentile_transform(fitted_select_percentile, X_r) + if o_select_percentile.rows > 0 { + if o_select_percentile.cols > 0 { sink = sink + o_select_percentile.data[0] } + } + matrix_free(o_select_percentile) + } + t3 = flow_now_ns() + select_percentile_free(fitted_select_percentile) + printf("ESTIMATOR|select_percentile|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- spectral_embedding (unsupervised) ---- + # ---- sequential_feature_selector (regression) ---- t0 = flow_now_ns() - let probe_spectral_embedding: SpectralEmbedding = spectral_embedding_fit(X_c, 2, 0.1) + let probe_sequential_feature_selector: SequentialFeatureSelector = sequential_feature_selector_fit(X_r, y_r, 2, 0, n_r, f_r) t1 = flow_now_ns() - spectral_embedding_free(probe_spectral_embedding) + sequential_feature_selector_free(probe_sequential_feature_selector) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_spectral_embedding: SpectralEmbedding = spectral_embedding_fit(X_c, 2, 0.1) - spectral_embedding_free(m_spectral_embedding) + let m_sequential_feature_selector: SequentialFeatureSelector = sequential_feature_selector_fit(X_r, y_r, 2, 0, n_r, f_r) + sequential_feature_selector_free(m_sequential_feature_selector) } t1 = flow_now_ns() - printf("ESTIMATOR|spectral_embedding|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + let fitted_sequential_feature_selector: SequentialFeatureSelector = sequential_feature_selector_fit(X_r, y_r, 2, 0, n_r, f_r) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_sequential_feature_selector: Matrix = sequential_feature_selector_transform(fitted_sequential_feature_selector, X_r) + if o_sequential_feature_selector.rows > 0 { + if o_sequential_feature_selector.cols > 0 { sink = sink + o_sequential_feature_selector.data[0] } + } + matrix_free(o_sequential_feature_selector) + } + t3 = flow_now_ns() + sequential_feature_selector_free(fitted_sequential_feature_selector) + printf("ESTIMATOR|sequential_feature_selector|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } @@ -539,5 +614,10 @@ function main() -> i32 { matrix_free(Y_multi) free(yi_c as ptr) free(yi_r as ptr) + array_free_f32(x1d_r) + array_free_f32(w_f) + # The sink is printed so the work above cannot be optimized away. The + # parser matches ESTIMATOR lines only, so this one is ignored. + printf("SINK|%.9f\n", sink) return 0 } diff --git a/benchmarks/generated/bench_estimators_08.flow b/benchmarks/generated/bench_estimators_08.flow index a654f22..3311597 100644 --- a/benchmarks/generated/bench_estimators_08.flow +++ b/benchmarks/generated/bench_estimators_08.flow @@ -65,302 +65,512 @@ function main() -> i32 { Y_rows[i] = row } + # A one-dimensional x for the isotonic row, which regresses against a + # single ordered variable rather than a design. + let x1d_r: ptr = array_new_f32(n_r) + for i in 0 to n_r { x1d_r[i] = matrix_at(X_r, i, 0) } + + # Per-feature importances for the selector row, which takes the weights a + # fitted model would hand it rather than a design. + let w_f: ptr = array_new_f32(f_c) + for i in 0 to f_c { w_f[i] = 1.0 / ((i + 1) as f32) } + let mut t0: i64 = 0 let mut t1: i64 = 0 let mut t2: i64 = 0 let mut t3: i64 = 0 let mut reps: i32 = 1 + let mut sink: f32 = 0.0 - # ---- spline_transformer (unsupervised) ---- + # ---- sgd_classifier (classification) ---- t0 = flow_now_ns() - let probe_spline_transformer: SplineTransformer = spline_transformer_fit(X_c, 4, 2) + let probe_sgd_classifier: SGDClassifier = sgd_classifier_fit(X_c, y_c, 3, 0, 1.0, 50, 0.01) t1 = flow_now_ns() - spline_transformer_free(probe_spline_transformer) + sgd_classifier_free(probe_sgd_classifier) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_spline_transformer: SplineTransformer = spline_transformer_fit(X_c, 4, 2) - spline_transformer_free(m_spline_transformer) + let m_sgd_classifier: SGDClassifier = sgd_classifier_fit(X_c, y_c, 3, 0, 1.0, 50, 0.01) + sgd_classifier_free(m_sgd_classifier) } t1 = flow_now_ns() - let fitted_spline_transformer: SplineTransformer = spline_transformer_fit(X_c, 4, 2) + let fitted_sgd_classifier: SGDClassifier = sgd_classifier_fit(X_c, y_c, 3, 0, 1.0, 50, 0.01) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_spline_transformer: Matrix = spline_transformer_transform(fitted_spline_transformer, X_c) - matrix_free(o_spline_transformer) + let o_sgd_classifier: ptr = sgd_classifier_predict(fitted_sgd_classifier, X_c) + sink = sink + o_sgd_classifier[0] + array_free_f32(o_sgd_classifier) } t3 = flow_now_ns() - spline_transformer_free(fitted_spline_transformer) - printf("ESTIMATOR|spline_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + sgd_classifier_free(fitted_sgd_classifier) + printf("ESTIMATOR|sgd_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- stacking_regressor (regression) ---- + # ---- sgd_one_class_svm (unsupervised) ---- t0 = flow_now_ns() - let probe_stacking_regressor: StackingRegressor = stacking_regressor_fit(X_r, y_r, 3, 5, 42) + let probe_sgd_one_class_svm: SGDOneClassSVM = sgd_one_class_svm_fit(X_c, 0.5, 50, 0.01) t1 = flow_now_ns() - stacking_regressor_free(probe_stacking_regressor) + sgd_one_class_svm_free(probe_sgd_one_class_svm) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_stacking_regressor: StackingRegressor = stacking_regressor_fit(X_r, y_r, 3, 5, 42) - stacking_regressor_free(m_stacking_regressor) + let m_sgd_one_class_svm: SGDOneClassSVM = sgd_one_class_svm_fit(X_c, 0.5, 50, 0.01) + sgd_one_class_svm_free(m_sgd_one_class_svm) } t1 = flow_now_ns() - let fitted_stacking_regressor: StackingRegressor = stacking_regressor_fit(X_r, y_r, 3, 5, 42) + let fitted_sgd_one_class_svm: SGDOneClassSVM = sgd_one_class_svm_fit(X_c, 0.5, 50, 0.01) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_stacking_regressor: ptr = stacking_regressor_predict(fitted_stacking_regressor, X_r) - array_free_f32(o_stacking_regressor) + let o_sgd_one_class_svm: ptr = sgd_one_class_svm_predict(fitted_sgd_one_class_svm, X_c) + sink = sink + o_sgd_one_class_svm[0] + array_free_f32(o_sgd_one_class_svm) } t3 = flow_now_ns() - stacking_regressor_free(fitted_stacking_regressor) - printf("ESTIMATOR|stacking_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + sgd_one_class_svm_free(fitted_sgd_one_class_svm) + printf("ESTIMATOR|sgd_one_class_svm|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- standard_scaler (unsupervised) ---- + # ---- sgd_regressor (regression) ---- t0 = flow_now_ns() - let probe_standard_scaler: StandardScaler = standard_scaler_fit(X_c) + let probe_sgd_regressor: SGDRegressor = sgd_regressor_fit(X_r, y_r, 0, 1.0, 0.1, 50, 0.01) t1 = flow_now_ns() - standard_scaler_free(probe_standard_scaler) + sgd_regressor_free(probe_sgd_regressor) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_standard_scaler: StandardScaler = standard_scaler_fit(X_c) - standard_scaler_free(m_standard_scaler) + let m_sgd_regressor: SGDRegressor = sgd_regressor_fit(X_r, y_r, 0, 1.0, 0.1, 50, 0.01) + sgd_regressor_free(m_sgd_regressor) } t1 = flow_now_ns() - let fitted_standard_scaler: StandardScaler = standard_scaler_fit(X_c) + let fitted_sgd_regressor: SGDRegressor = sgd_regressor_fit(X_r, y_r, 0, 1.0, 0.1, 50, 0.01) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_standard_scaler: Matrix = standard_scaler_transform(fitted_standard_scaler, X_c) - matrix_free(o_standard_scaler) + let o_sgd_regressor: ptr = sgd_regressor_predict(fitted_sgd_regressor, X_r) + sink = sink + o_sgd_regressor[0] + array_free_f32(o_sgd_regressor) } t3 = flow_now_ns() - standard_scaler_free(fitted_standard_scaler) - printf("ESTIMATOR|standard_scaler|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + sgd_regressor_free(fitted_sgd_regressor) + printf("ESTIMATOR|sgd_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- svc (classification) ---- + # ---- shrunk_covariance (unsupervised) ---- t0 = flow_now_ns() - let probe_svc: SVC = svc_fit(X_c, y_c, 3, 1.0, 0, 0.1, 2, 0.0) + let probe_shrunk_covariance: ShrunkCovariance = shrunk_covariance_fit(X_c, 0.1) t1 = flow_now_ns() - svc_free(probe_svc) + shrunk_covariance_free(probe_shrunk_covariance) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_svc: SVC = svc_fit(X_c, y_c, 3, 1.0, 0, 0.1, 2, 0.0) - svc_free(m_svc) + let m_shrunk_covariance: ShrunkCovariance = shrunk_covariance_fit(X_c, 0.1) + shrunk_covariance_free(m_shrunk_covariance) } t1 = flow_now_ns() - let fitted_svc: SVC = svc_fit(X_c, y_c, 3, 1.0, 0, 0.1, 2, 0.0) + printf("ESTIMATOR|shrunk_covariance|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + + # ---- simple_imputer (unsupervised) ---- + t0 = flow_now_ns() + let probe_simple_imputer: SimpleImputer = simple_imputer_fit(X_c, 0, 1.0) + t1 = flow_now_ns() + simple_imputer_free(probe_simple_imputer) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_simple_imputer: SimpleImputer = simple_imputer_fit(X_c, 0, 1.0) + simple_imputer_free(m_simple_imputer) + } + t1 = flow_now_ns() + let fitted_simple_imputer: SimpleImputer = simple_imputer_fit(X_c, 0, 1.0) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_svc: ptr = svc_predict(fitted_svc, X_c) - array_free_f32(o_svc) + let o_simple_imputer: Matrix = simple_imputer_transform(fitted_simple_imputer, X_c) + if o_simple_imputer.rows > 0 { + if o_simple_imputer.cols > 0 { sink = sink + o_simple_imputer.data[0] } + } + matrix_free(o_simple_imputer) } t3 = flow_now_ns() - svc_free(fitted_svc) - printf("ESTIMATOR|svc|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + simple_imputer_free(fitted_simple_imputer) + printf("ESTIMATOR|simple_imputer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- svr (regression) ---- + # ---- skewed_chi2_sampler (classification, written out) ---- t0 = flow_now_ns() - let probe_svr: SVR = svr_fit(X_r, y_r, 1.0, 0.1, 0, 0.1, 2, 0.0) + let probe_skewed_chi2_sampler: SkewedChi2Sampler = skewed_chi2_sampler_fit(f_c, 1.0, 2, 42) t1 = flow_now_ns() - svr_free(probe_svr) + skewed_chi2_sampler_free(probe_skewed_chi2_sampler) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_svr: SVR = svr_fit(X_r, y_r, 1.0, 0.1, 0, 0.1, 2, 0.0) - svr_free(m_svr) + let m_skewed_chi2_sampler: SkewedChi2Sampler = skewed_chi2_sampler_fit(f_c, 1.0, 2, 42) + skewed_chi2_sampler_free(m_skewed_chi2_sampler) } t1 = flow_now_ns() - let fitted_svr: SVR = svr_fit(X_r, y_r, 1.0, 0.1, 0, 0.1, 2, 0.0) + let fitted_skewed_chi2_sampler: SkewedChi2Sampler = skewed_chi2_sampler_fit(f_c, 1.0, 2, 42) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_svr: ptr = svr_predict(fitted_svr, X_r) - array_free_f32(o_svr) + let o_skewed_chi2_sampler: Matrix = skewed_chi2_sampler_transform(fitted_skewed_chi2_sampler, X_c) + if o_skewed_chi2_sampler.rows > 0 { + if o_skewed_chi2_sampler.cols > 0 { sink = sink + o_skewed_chi2_sampler.data[0] } + } + matrix_free(o_skewed_chi2_sampler) } t3 = flow_now_ns() - svr_free(fitted_svr) - printf("ESTIMATOR|svr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + skewed_chi2_sampler_free(fitted_skewed_chi2_sampler) + printf("ESTIMATOR|skewed_chi2_sampler|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- target_encoder (regression) ---- + # ---- sparse_coder (unsupervised) ---- t0 = flow_now_ns() - let probe_target_encoder: TargetEncoder = target_encoder_fit(X_r, y_r, n_r, 1.0) + let probe_sparse_coder: SparseCoder = sparse_coder_fit(X_c, 3) t1 = flow_now_ns() - target_encoder_free(probe_target_encoder) + sparse_coder_free(probe_sparse_coder) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_target_encoder: TargetEncoder = target_encoder_fit(X_r, y_r, n_r, 1.0) - target_encoder_free(m_target_encoder) + let m_sparse_coder: SparseCoder = sparse_coder_fit(X_c, 3) + sparse_coder_free(m_sparse_coder) } t1 = flow_now_ns() - let fitted_target_encoder: TargetEncoder = target_encoder_fit(X_r, y_r, n_r, 1.0) + let fitted_sparse_coder: SparseCoder = sparse_coder_fit(X_c, 3) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_target_encoder: Matrix = target_encoder_transform(fitted_target_encoder, X_r) - matrix_free(o_target_encoder) + let o_sparse_coder: Matrix = sparse_coder_transform(fitted_sparse_coder, X_c) + if o_sparse_coder.rows > 0 { + if o_sparse_coder.cols > 0 { sink = sink + o_sparse_coder.data[0] } + } + matrix_free(o_sparse_coder) } t3 = flow_now_ns() - target_encoder_free(fitted_target_encoder) - printf("ESTIMATOR|target_encoder|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + sparse_coder_free(fitted_sparse_coder) + printf("ESTIMATOR|sparse_coder|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- theil_sen_regressor (regression) ---- + # ---- sparse_pca (unsupervised) ---- t0 = flow_now_ns() - let probe_theil_sen_regressor: TheilSenRegressor = theil_sen_regressor_fit(X_r, y_r, 10, 100, 42) + let probe_sparse_pca: SparsePCA = sparse_pca_fit(X_c, 2, 1.0, 100, 0.0001, 42) t1 = flow_now_ns() - theil_sen_regressor_free(probe_theil_sen_regressor) + sparse_pca_free(probe_sparse_pca) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_theil_sen_regressor: TheilSenRegressor = theil_sen_regressor_fit(X_r, y_r, 10, 100, 42) - theil_sen_regressor_free(m_theil_sen_regressor) + let m_sparse_pca: SparsePCA = sparse_pca_fit(X_c, 2, 1.0, 100, 0.0001, 42) + sparse_pca_free(m_sparse_pca) } t1 = flow_now_ns() - let fitted_theil_sen_regressor: TheilSenRegressor = theil_sen_regressor_fit(X_r, y_r, 10, 100, 42) + let fitted_sparse_pca: SparsePCA = sparse_pca_fit(X_c, 2, 1.0, 100, 0.0001, 42) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_theil_sen_regressor: ptr = theil_sen_regressor_predict(fitted_theil_sen_regressor, X_r) - array_free_f32(o_theil_sen_regressor) + let o_sparse_pca: Matrix = sparse_pca_transform(fitted_sparse_pca, X_c) + if o_sparse_pca.rows > 0 { + if o_sparse_pca.cols > 0 { sink = sink + o_sparse_pca.data[0] } + } + matrix_free(o_sparse_pca) } t3 = flow_now_ns() - theil_sen_regressor_free(fitted_theil_sen_regressor) - printf("ESTIMATOR|theil_sen_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + sparse_pca_free(fitted_sparse_pca) + printf("ESTIMATOR|sparse_pca|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- sparse_random_projection (classification, written out) ---- + t0 = flow_now_ns() + let probe_sparse_random_projection: SparseRandomProjection = sparse_random_projection_fit(f_c, 2, 0.3, 42) + t1 = flow_now_ns() + sparse_random_projection_free(probe_sparse_random_projection) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_sparse_random_projection: SparseRandomProjection = sparse_random_projection_fit(f_c, 2, 0.3, 42) + sparse_random_projection_free(m_sparse_random_projection) + } + t1 = flow_now_ns() + let fitted_sparse_random_projection: SparseRandomProjection = sparse_random_projection_fit(f_c, 2, 0.3, 42) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_sparse_random_projection: Matrix = sparse_random_projection_transform(fitted_sparse_random_projection, X_c) + if o_sparse_random_projection.rows > 0 { + if o_sparse_random_projection.cols > 0 { sink = sink + o_sparse_random_projection.data[0] } + } + matrix_free(o_sparse_random_projection) + } + t3 = flow_now_ns() + sparse_random_projection_free(fitted_sparse_random_projection) + printf("ESTIMATOR|sparse_random_projection|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- spectral_biclustering (unsupervised) ---- + t0 = flow_now_ns() + let probe_spectral_biclustering: SpectralBiclustering = spectral_biclustering_fit(X_c, 2, 2) + t1 = flow_now_ns() + spectral_biclustering_free(probe_spectral_biclustering) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_spectral_biclustering: SpectralBiclustering = spectral_biclustering_fit(X_c, 2, 2) + spectral_biclustering_free(m_spectral_biclustering) + } + t1 = flow_now_ns() + printf("ESTIMATOR|spectral_biclustering|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + + # ---- spectral_clustering (unsupervised) ---- + t0 = flow_now_ns() + let probe_spectral_clustering: SpectralClustering = spectral_clustering_fit(X_c, 3, 0.1, 42) + t1 = flow_now_ns() + spectral_clustering_free(probe_spectral_clustering) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_spectral_clustering: SpectralClustering = spectral_clustering_fit(X_c, 3, 0.1, 42) + spectral_clustering_free(m_spectral_clustering) + } + t1 = flow_now_ns() + printf("ESTIMATOR|spectral_clustering|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + + # ---- spectral_coclustering (unsupervised) ---- + t0 = flow_now_ns() + let probe_spectral_coclustering: SpectralCoclustering = spectral_coclustering_fit(X_c, 3) + t1 = flow_now_ns() + spectral_coclustering_free(probe_spectral_coclustering) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_spectral_coclustering: SpectralCoclustering = spectral_coclustering_fit(X_c, 3) + spectral_coclustering_free(m_spectral_coclustering) + } + t1 = flow_now_ns() + printf("ESTIMATOR|spectral_coclustering|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) fflush(null) - # ---- transformed_target_regressor (regression) ---- + # ---- spectral_embedding (unsupervised) ---- t0 = flow_now_ns() - let probe_transformed_target_regressor: TransformedTargetRegressor = transformed_target_regressor_fit(X_r, y_r, 0) + let probe_spectral_embedding: SpectralEmbedding = spectral_embedding_fit(X_c, 2, 0.1) t1 = flow_now_ns() - transformed_target_regressor_free(probe_transformed_target_regressor) + spectral_embedding_free(probe_spectral_embedding) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_transformed_target_regressor: TransformedTargetRegressor = transformed_target_regressor_fit(X_r, y_r, 0) - transformed_target_regressor_free(m_transformed_target_regressor) + let m_spectral_embedding: SpectralEmbedding = spectral_embedding_fit(X_c, 2, 0.1) + spectral_embedding_free(m_spectral_embedding) } t1 = flow_now_ns() - let fitted_transformed_target_regressor: TransformedTargetRegressor = transformed_target_regressor_fit(X_r, y_r, 0) + printf("ESTIMATOR|spectral_embedding|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + + # ---- spline_transformer (unsupervised) ---- + t0 = flow_now_ns() + let probe_spline_transformer: SplineTransformer = spline_transformer_fit(X_c, 4, 2) + t1 = flow_now_ns() + spline_transformer_free(probe_spline_transformer) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_spline_transformer: SplineTransformer = spline_transformer_fit(X_c, 4, 2) + spline_transformer_free(m_spline_transformer) + } + t1 = flow_now_ns() + let fitted_spline_transformer: SplineTransformer = spline_transformer_fit(X_c, 4, 2) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_transformed_target_regressor: ptr = transformed_target_regressor_predict(fitted_transformed_target_regressor, X_r) - array_free_f32(o_transformed_target_regressor) + let o_spline_transformer: Matrix = spline_transformer_transform(fitted_spline_transformer, X_c) + if o_spline_transformer.rows > 0 { + if o_spline_transformer.cols > 0 { sink = sink + o_spline_transformer.data[0] } + } + matrix_free(o_spline_transformer) } t3 = flow_now_ns() - transformed_target_regressor_free(fitted_transformed_target_regressor) - printf("ESTIMATOR|transformed_target_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + spline_transformer_free(fitted_spline_transformer) + printf("ESTIMATOR|spline_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- truncated_svd (unsupervised) ---- + # ---- stacking_regressor (regression) ---- t0 = flow_now_ns() - let probe_truncated_svd: TruncatedSVD = truncated_svd_fit(X_c, 2) + let probe_stacking_regressor: StackingRegressor = stacking_regressor_fit(X_r, y_r, 3, 5, 42) t1 = flow_now_ns() - truncated_svd_free(probe_truncated_svd) + stacking_regressor_free(probe_stacking_regressor) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_truncated_svd: TruncatedSVD = truncated_svd_fit(X_c, 2) - truncated_svd_free(m_truncated_svd) + let m_stacking_regressor: StackingRegressor = stacking_regressor_fit(X_r, y_r, 3, 5, 42) + stacking_regressor_free(m_stacking_regressor) } t1 = flow_now_ns() - let fitted_truncated_svd: TruncatedSVD = truncated_svd_fit(X_c, 2) + let fitted_stacking_regressor: StackingRegressor = stacking_regressor_fit(X_r, y_r, 3, 5, 42) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_truncated_svd: Matrix = truncated_svd_transform(fitted_truncated_svd, X_c) - matrix_free(o_truncated_svd) + let o_stacking_regressor: ptr = stacking_regressor_predict(fitted_stacking_regressor, X_r) + sink = sink + o_stacking_regressor[0] + array_free_f32(o_stacking_regressor) } t3 = flow_now_ns() - truncated_svd_free(fitted_truncated_svd) - printf("ESTIMATOR|truncated_svd|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + stacking_regressor_free(fitted_stacking_regressor) + printf("ESTIMATOR|stacking_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- tsne (unsupervised) ---- + # ---- standard_scaler (unsupervised) ---- t0 = flow_now_ns() - let probe_tsne: TSNE = tsne_fit(X_c, 2, 5.0, 0.1, 50, 42) + let probe_standard_scaler: StandardScaler = standard_scaler_fit(X_c) t1 = flow_now_ns() - tsne_free(probe_tsne) + standard_scaler_free(probe_standard_scaler) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_tsne: TSNE = tsne_fit(X_c, 2, 5.0, 0.1, 50, 42) - tsne_free(m_tsne) + let m_standard_scaler: StandardScaler = standard_scaler_fit(X_c) + standard_scaler_free(m_standard_scaler) } t1 = flow_now_ns() - printf("ESTIMATOR|tsne|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + let fitted_standard_scaler: StandardScaler = standard_scaler_fit(X_c) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_standard_scaler: Matrix = standard_scaler_transform(fitted_standard_scaler, X_c) + if o_standard_scaler.rows > 0 { + if o_standard_scaler.cols > 0 { sink = sink + o_standard_scaler.data[0] } + } + matrix_free(o_standard_scaler) + } + t3 = flow_now_ns() + standard_scaler_free(fitted_standard_scaler) + printf("ESTIMATOR|standard_scaler|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- tweedie_regressor (regression) ---- + # ---- svc (classification) ---- t0 = flow_now_ns() - let probe_tweedie_regressor: TweedieRegressor = tweedie_regressor_fit(X_r, y_r, 1.0, 1.5, 100, 0.01) + let probe_svc: SVC = svc_fit(X_c, y_c, 3, 1.0, 0, 0.1, 2, 0.0) t1 = flow_now_ns() - tweedie_regressor_free(probe_tweedie_regressor) + svc_free(probe_svc) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_tweedie_regressor: TweedieRegressor = tweedie_regressor_fit(X_r, y_r, 1.0, 1.5, 100, 0.01) - tweedie_regressor_free(m_tweedie_regressor) + let m_svc: SVC = svc_fit(X_c, y_c, 3, 1.0, 0, 0.1, 2, 0.0) + svc_free(m_svc) } t1 = flow_now_ns() - let fitted_tweedie_regressor: TweedieRegressor = tweedie_regressor_fit(X_r, y_r, 1.0, 1.5, 100, 0.01) + let fitted_svc: SVC = svc_fit(X_c, y_c, 3, 1.0, 0, 0.1, 2, 0.0) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_tweedie_regressor: ptr = tweedie_regressor_predict(fitted_tweedie_regressor, X_r) - array_free_f32(o_tweedie_regressor) + let o_svc: ptr = svc_predict(fitted_svc, X_c) + sink = sink + o_svc[0] + array_free_f32(o_svc) } t3 = flow_now_ns() - tweedie_regressor_free(fitted_tweedie_regressor) - printf("ESTIMATOR|tweedie_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + svc_free(fitted_svc) + printf("ESTIMATOR|svc|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- variance_threshold (unsupervised) ---- + # ---- svr (regression) ---- t0 = flow_now_ns() - let probe_variance_threshold: VarianceThreshold = variance_threshold_fit(X_c, 0.5) + let probe_svr: SVR = svr_fit(X_r, y_r, 1.0, 0.1, 0, 0.1, 2, 0.0) t1 = flow_now_ns() - variance_threshold_free(probe_variance_threshold) + svr_free(probe_svr) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_variance_threshold: VarianceThreshold = variance_threshold_fit(X_c, 0.5) - variance_threshold_free(m_variance_threshold) + let m_svr: SVR = svr_fit(X_r, y_r, 1.0, 0.1, 0, 0.1, 2, 0.0) + svr_free(m_svr) } t1 = flow_now_ns() - let fitted_variance_threshold: VarianceThreshold = variance_threshold_fit(X_c, 0.5) + let fitted_svr: SVR = svr_fit(X_r, y_r, 1.0, 0.1, 0, 0.1, 2, 0.0) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_variance_threshold: Matrix = variance_threshold_transform(fitted_variance_threshold, X_c) - matrix_free(o_variance_threshold) + let o_svr: ptr = svr_predict(fitted_svr, X_r) + sink = sink + o_svr[0] + array_free_f32(o_svr) } t3 = flow_now_ns() - variance_threshold_free(fitted_variance_threshold) - printf("ESTIMATOR|variance_threshold|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + svr_free(fitted_svr) + printf("ESTIMATOR|svr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- target_encoder (regression) ---- + t0 = flow_now_ns() + let probe_target_encoder: TargetEncoder = target_encoder_fit(X_r, y_r, n_r, 1.0) + t1 = flow_now_ns() + target_encoder_free(probe_target_encoder) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_target_encoder: TargetEncoder = target_encoder_fit(X_r, y_r, n_r, 1.0) + target_encoder_free(m_target_encoder) + } + t1 = flow_now_ns() + let fitted_target_encoder: TargetEncoder = target_encoder_fit(X_r, y_r, n_r, 1.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_target_encoder: Matrix = target_encoder_transform(fitted_target_encoder, X_r) + if o_target_encoder.rows > 0 { + if o_target_encoder.cols > 0 { sink = sink + o_target_encoder.data[0] } + } + matrix_free(o_target_encoder) + } + t3 = flow_now_ns() + target_encoder_free(fitted_target_encoder) + printf("ESTIMATOR|target_encoder|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- theil_sen_regressor (regression) ---- + t0 = flow_now_ns() + let probe_theil_sen_regressor: TheilSenRegressor = theil_sen_regressor_fit(X_r, y_r, 10, 100, 42) + t1 = flow_now_ns() + theil_sen_regressor_free(probe_theil_sen_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_theil_sen_regressor: TheilSenRegressor = theil_sen_regressor_fit(X_r, y_r, 10, 100, 42) + theil_sen_regressor_free(m_theil_sen_regressor) + } + t1 = flow_now_ns() + let fitted_theil_sen_regressor: TheilSenRegressor = theil_sen_regressor_fit(X_r, y_r, 10, 100, 42) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_theil_sen_regressor: ptr = theil_sen_regressor_predict(fitted_theil_sen_regressor, X_r) + sink = sink + o_theil_sen_regressor[0] + array_free_f32(o_theil_sen_regressor) + } + t3 = flow_now_ns() + theil_sen_regressor_free(fitted_theil_sen_regressor) + printf("ESTIMATOR|theil_sen_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } @@ -371,5 +581,10 @@ function main() -> i32 { matrix_free(Y_multi) free(yi_c as ptr) free(yi_r as ptr) + array_free_f32(x1d_r) + array_free_f32(w_f) + # The sink is printed so the work above cannot be optimized away. The + # parser matches ESTIMATOR lines only, so this one is ignored. + printf("SINK|%.9f\n", sink) return 0 } diff --git a/benchmarks/generated/bench_estimators_09.flow b/benchmarks/generated/bench_estimators_09.flow new file mode 100644 index 0000000..e29c98f --- /dev/null +++ b/benchmarks/generated/bench_estimators_09.flow @@ -0,0 +1,224 @@ +# Generated by benchmarks/generate_estimator_bench.py. Do not edit. +# +# One timing block per Flow estimator that the registry marks runnable. +# Regenerate with: +# python benchmarks/estimator_coverage.py +# python benchmarks/generate_estimator_bench.py + +import "lib/scikit/scikit.flow" + +extern { + function printf(fmt: string, ...) -> i32 + function flow_now_ns() -> i64 + function malloc(size: i64) -> ptr + function free(p: ptr) -> void + function fflush(stream: ptr) -> i32 +} + +function ms_between(start: i64, finish: i64) -> f32 { + return ((((finish - start) as f64) / 1000000.0) as f32) +} + +function main() -> i32 { + let iris: Dataset = load_iris() + let diabetes: Dataset = load_diabetes() + let X_c: Matrix = iris.X + let y_c: ptr = iris.y + let n_c: i32 = X_c.rows + let f_c: i32 = X_c.cols + let X_r: Matrix = diabetes.X + let y_r: ptr = diabetes.y + let n_r: i32 = X_r.rows + let f_r: i32 = X_r.cols + + let yi_c: ptr = malloc((n_c as i64) * 4) as ptr + for i in 0 to n_c { yi_c[i] = y_c[i] as i32 } + let yi_r: ptr = malloc((n_r as i64) * 4) as ptr + for i in 0 to n_r { yi_r[i] = y_r[i] as i32 } + + # A two-column target for the cross-decomposition estimators. + let Y_multi: Matrix = matrix_new(n_r, 2) + for i in 0 to n_r { + matrix_set(Y_multi, i, 0, y_r[i]) + matrix_set(Y_multi, i, 1, y_r[i] * 0.5) + } + + # Label-valued targets for the multi-output classifiers. + let Y_labels: Matrix = matrix_new(n_c, 2) + let Y_label_rows: ptr > = malloc((n_c as i64) * 8) as ptr > + for i in 0 to n_c { + let a: f32 = y_c[i] + let b: f32 = ((((y_c[i] as i32) + 1) % 3) as f32) + matrix_set(Y_labels, i, 0, a) + matrix_set(Y_labels, i, 1, b) + let lrow: ptr = array_new_f32(2) + lrow[0] = a + lrow[1] = b + Y_label_rows[i] = lrow + } + + let Y_rows: ptr > = malloc((n_r as i64) * 8) as ptr > + for i in 0 to n_r { + let row: ptr = array_new_f32(2) + row[0] = y_r[i] + row[1] = y_r[i] * 0.5 + Y_rows[i] = row + } + + # A one-dimensional x for the isotonic row, which regresses against a + # single ordered variable rather than a design. + let x1d_r: ptr = array_new_f32(n_r) + for i in 0 to n_r { x1d_r[i] = matrix_at(X_r, i, 0) } + + # Per-feature importances for the selector row, which takes the weights a + # fitted model would hand it rather than a design. + let w_f: ptr = array_new_f32(f_c) + for i in 0 to f_c { w_f[i] = 1.0 / ((i + 1) as f32) } + + let mut t0: i64 = 0 + let mut t1: i64 = 0 + let mut t2: i64 = 0 + let mut t3: i64 = 0 + let mut reps: i32 = 1 + let mut sink: f32 = 0.0 + + # ---- transformed_target_regressor (regression) ---- + t0 = flow_now_ns() + let probe_transformed_target_regressor: TransformedTargetRegressor = transformed_target_regressor_fit(X_r, y_r, 0) + t1 = flow_now_ns() + transformed_target_regressor_free(probe_transformed_target_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_transformed_target_regressor: TransformedTargetRegressor = transformed_target_regressor_fit(X_r, y_r, 0) + transformed_target_regressor_free(m_transformed_target_regressor) + } + t1 = flow_now_ns() + let fitted_transformed_target_regressor: TransformedTargetRegressor = transformed_target_regressor_fit(X_r, y_r, 0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_transformed_target_regressor: ptr = transformed_target_regressor_predict(fitted_transformed_target_regressor, X_r) + sink = sink + o_transformed_target_regressor[0] + array_free_f32(o_transformed_target_regressor) + } + t3 = flow_now_ns() + transformed_target_regressor_free(fitted_transformed_target_regressor) + printf("ESTIMATOR|transformed_target_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- truncated_svd (unsupervised) ---- + t0 = flow_now_ns() + let probe_truncated_svd: TruncatedSVD = truncated_svd_fit(X_c, 2) + t1 = flow_now_ns() + truncated_svd_free(probe_truncated_svd) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_truncated_svd: TruncatedSVD = truncated_svd_fit(X_c, 2) + truncated_svd_free(m_truncated_svd) + } + t1 = flow_now_ns() + let fitted_truncated_svd: TruncatedSVD = truncated_svd_fit(X_c, 2) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_truncated_svd: Matrix = truncated_svd_transform(fitted_truncated_svd, X_c) + if o_truncated_svd.rows > 0 { + if o_truncated_svd.cols > 0 { sink = sink + o_truncated_svd.data[0] } + } + matrix_free(o_truncated_svd) + } + t3 = flow_now_ns() + truncated_svd_free(fitted_truncated_svd) + printf("ESTIMATOR|truncated_svd|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- tsne (unsupervised) ---- + t0 = flow_now_ns() + let probe_tsne: TSNE = tsne_fit(X_c, 2, 5.0, 0.1, 50, 42) + t1 = flow_now_ns() + tsne_free(probe_tsne) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_tsne: TSNE = tsne_fit(X_c, 2, 5.0, 0.1, 50, 42) + tsne_free(m_tsne) + } + t1 = flow_now_ns() + printf("ESTIMATOR|tsne|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + + # ---- tweedie_regressor (regression) ---- + t0 = flow_now_ns() + let probe_tweedie_regressor: TweedieRegressor = tweedie_regressor_fit(X_r, y_r, 1.0, 1.5, 100, 0.01) + t1 = flow_now_ns() + tweedie_regressor_free(probe_tweedie_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_tweedie_regressor: TweedieRegressor = tweedie_regressor_fit(X_r, y_r, 1.0, 1.5, 100, 0.01) + tweedie_regressor_free(m_tweedie_regressor) + } + t1 = flow_now_ns() + let fitted_tweedie_regressor: TweedieRegressor = tweedie_regressor_fit(X_r, y_r, 1.0, 1.5, 100, 0.01) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_tweedie_regressor: ptr = tweedie_regressor_predict(fitted_tweedie_regressor, X_r) + sink = sink + o_tweedie_regressor[0] + array_free_f32(o_tweedie_regressor) + } + t3 = flow_now_ns() + tweedie_regressor_free(fitted_tweedie_regressor) + printf("ESTIMATOR|tweedie_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- variance_threshold (unsupervised) ---- + t0 = flow_now_ns() + let probe_variance_threshold: VarianceThreshold = variance_threshold_fit(X_c, 0.5) + t1 = flow_now_ns() + variance_threshold_free(probe_variance_threshold) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_variance_threshold: VarianceThreshold = variance_threshold_fit(X_c, 0.5) + variance_threshold_free(m_variance_threshold) + } + t1 = flow_now_ns() + let fitted_variance_threshold: VarianceThreshold = variance_threshold_fit(X_c, 0.5) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_variance_threshold: Matrix = variance_threshold_transform(fitted_variance_threshold, X_c) + if o_variance_threshold.rows > 0 { + if o_variance_threshold.cols > 0 { sink = sink + o_variance_threshold.data[0] } + } + matrix_free(o_variance_threshold) + } + t3 = flow_now_ns() + variance_threshold_free(fitted_variance_threshold) + printf("ESTIMATOR|variance_threshold|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } + free(Y_label_rows as ptr) + matrix_free(Y_labels) + for i in 0 to n_r { array_free_f32(Y_rows[i]) } + free(Y_rows as ptr) + matrix_free(Y_multi) + free(yi_c as ptr) + free(yi_r as ptr) + array_free_f32(x1d_r) + array_free_f32(w_f) + # The sink is printed so the work above cannot be optimized away. The + # parser matches ESTIMATOR lines only, so this one is ignored. + printf("SINK|%.9f\n", sink) + return 0 +} diff --git a/lib/scikit/decomposition.flow b/lib/scikit/decomposition.flow index 524af63..ba28060 100644 --- a/lib/scikit/decomposition.flow +++ b/lib/scikit/decomposition.flow @@ -1367,7 +1367,13 @@ export function factor_analysis_fit(X: Matrix, n_components: i32, max_iter: i32, } } + # Every component row is allocated here. malloc does not zero what it + # hands back, so the loop below used to test components[c] against null and + # write through whatever the allocator left in place when the test failed. + # On macOS that memory happened to be zero and the fit worked. Under + # glibc's allocator it was not, and the fit wrote through a garbage pointer. let components: ptr > = malloc((n_components as i64) * 8) as ptr > + for c in 0 to n_components { components[c] = array_new_f32(p) } let noise_variance: ptr = array_new_f32(p) for j in 0 to p { @@ -1424,9 +1430,6 @@ export function factor_analysis_fit(X: Matrix, n_components: i32, max_iter: i32, } # Store component c - if components[c] == null { - components[c] = array_new_f32(p) - } for j in 0 to p { components[c][j] = w[j] } From 89b1dfe19ef38dab0667bcd57a7958f8cfd2fb74 Mon Sep 17 00:00:00 2001 From: godofecht Date: Sun, 27 Sep 2026 22:24:06 +0100 Subject: [PATCH 07/12] Race the last seven that can be raced, and fix the hash that aborted The composed rows and the input-shaped rows now carry their call in the registry the same way the other shaped rows do: Pipeline over a scaler and a logistic regression, ColumnTransformer over a scaler, FeatureUnion over a scaler and a passthrough, the two text vectorizers over one corpus, the dict vectorizer and the multi-label binarizer. That is 192 of 203 raced. The corpus lives in the registry, so the generated Flow file and the scikit-learn harness read one copy of it and cannot drift apart. A fit that takes a composed object built outside the timing frees once after the timing rather than on every repeat, because the object it hands back is the object the next repeat reads. _hash_string carried djb2 in an i32 and shifted the result left every character. Past the sixth character of a token the value had wrapped negative, and the runtime traps a left shift of a negative, so the process aborted inside count_vectorizer_fit. The four document fixture in the tests never reached that. A sixteen document corpus did. It is carried in 64 bits and reduced every step now. The five that are left are the ones where the Flow function does part of what the scikit-learn class does: the two voting estimators take fitted models, the stacking classifier takes base predictions, the self-training classifier takes probabilities, and incremental PCA takes one batch. Each says so in its own reason rather than in a generic message about its first argument. --- benchmarks/bench_estimators_sklearn.py | 21 +- benchmarks/estimator_coverage.json | 223 ++++++++++-- benchmarks/estimator_coverage.py | 133 +++++++ benchmarks/generate_estimator_bench.py | 21 +- benchmarks/generated/bench_estimators_00.flow | 49 ++- benchmarks/generated/bench_estimators_01.flow | 174 ++++++---- benchmarks/generated/bench_estimators_02.flow | 206 ++++++----- benchmarks/generated/bench_estimators_03.flow | 201 ++++++----- benchmarks/generated/bench_estimators_04.flow | 207 +++++------ benchmarks/generated/bench_estimators_05.flow | 256 ++++++++------ benchmarks/generated/bench_estimators_06.flow | 303 ++++++++-------- benchmarks/generated/bench_estimators_07.flow | 328 +++++++++--------- benchmarks/generated/bench_estimators_08.flow | 328 +++++++++--------- benchmarks/generated/bench_estimators_09.flow | 206 +++++++++++ lib/scikit/feature_extraction.flow | 13 +- 15 files changed, 1632 insertions(+), 1037 deletions(-) diff --git a/benchmarks/bench_estimators_sklearn.py b/benchmarks/bench_estimators_sklearn.py index f276f67..ed09341 100644 --- a/benchmarks/bench_estimators_sklearn.py +++ b/benchmarks/bench_estimators_sklearn.py @@ -33,7 +33,8 @@ # this data. A shallow tree keeps the wrapper's own overhead visible rather than # burying it under the base estimator's work. def _constructors() -> dict: - from sklearn.linear_model import Ridge + from sklearn.linear_model import LogisticRegression, Ridge + from sklearn.preprocessing import FunctionTransformer, StandardScaler from sklearn.tree import DecisionTreeClassifier, DecisionTreeRegressor import numpy as np @@ -59,6 +60,14 @@ def _constructors() -> dict: "NuSVC": lambda c: c(nu=0.1), # A selector needs something to read importances from. "SelectFromModel": lambda c: c(tree_c(), threshold=-np.inf, max_features=2), + # The composed rows, built to match what the Flow harness composes: + # a scaler and a logistic regression, a scaler over the columns, and a + # scaler beside a passthrough. + "Pipeline": lambda c: c([("scaler", StandardScaler()), + ("classifier", LogisticRegression(max_iter=50))]), + "ColumnTransformer": lambda c: c([("scaler", StandardScaler(), [0, 1, 2, 3])]), + "FeatureUnion": lambda c: c([("scaler", StandardScaler()), + ("passthrough", FunctionTransformer())]), } @@ -136,6 +145,16 @@ def main() -> int: first, second = y, None elif fit_input == "x1d": first, second = X[:, 0], y + elif fit_input == "docs": + # The corpus travels in the registry, so the Flow file and this one + # read one copy of it. + first, second = list(shape["corpus"]), None + elif fit_input == "dicts": + first = [{f"f{j}": float(v) for j, v in enumerate(row)} for row in X] + second = None + elif fit_input == "labelsets": + first = [tuple(int(v) for v in row) for row in y] + second = None else: first, second = X, y try: diff --git a/benchmarks/estimator_coverage.json b/benchmarks/estimator_coverage.json index f9eccb6..4fa9de1 100644 --- a/benchmarks/estimator_coverage.json +++ b/benchmarks/estimator_coverage.json @@ -3,10 +3,10 @@ "counts": { "estimators": 203, "runnable": 168, - "shaped": 13, - "different_shape": 12, + "shaped": 20, "simplified": 4, - "flow_only": 6 + "flow_only": 6, + "different_shape": 5 }, "sklearn_surface": 208, "entries": [ @@ -1162,9 +1162,28 @@ "ct", "X" ], - "bucket": "different_shape", - "reason": "takes ColumnTransformer first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "ColumnTransformer" + "bucket": "shaped", + "sklearn_estimator": "ColumnTransformer", + "shape": { + "dataset": "classification", + "flow_preamble": [ + "let ct_cols: ptr = malloc((f_c as i64) * 4) as ptr", + "for i in 0 to f_c { ct_cols[i] = i }", + "let ct_obj: ColumnTransformer = column_transformer_init(1)", + "# 0 is TRANSFORMER_STANDARD_SCALER. The constant is written out", + "# because an export const is not visible through the umbrella import.", + "column_transformer_set_spec(ct_obj, 0, 0, ct_cols, f_c)" + ], + "flow_fit": [ + "ct_obj", + "X_c" + ], + "flow_work": [ + "X_c" + ], + "flow_free": "after", + "sklearn_input": "X" + } }, { "flow_estimator": "complement_nb", @@ -1282,9 +1301,39 @@ "n_docs", "max_features" ], - "bucket": "different_shape", - "reason": "takes ptr first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "CountVectorizer" + "bucket": "shaped", + "sklearn_estimator": "CountVectorizer", + "shape": { + "dataset": "classification", + "corpus": [ + "the quick brown fox jumps over the lazy dog", + "a lazy dog sleeps in the warm sun", + "quick brown foxes are rare in the city", + "the dog and the fox share a field", + "warm sun and a cold river run together", + "a field of brown grass in the sun", + "the city river runs past the old field", + "old dogs sleep through a quick storm", + "a storm over the city wakes the dog", + "foxes hunt in the cold river valley", + "the valley holds a warm field of grass", + "grass grows where the river meets the sun", + "a rare fox crosses the old stone bridge", + "the stone bridge over the cold river", + "dogs and foxes keep their distance here", + "here the field the river and the city meet" + ], + "flow_fit": [ + "count_vectorizer_docs", + "16", + "50" + ], + "flow_work": [ + "count_vectorizer_docs", + "16" + ], + "sklearn_input": "docs" + } }, { "flow_estimator": "dbscan", @@ -1521,9 +1570,40 @@ "n_samples", "n_keys_per_sample" ], - "bucket": "different_shape", - "reason": "takes ptr > first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "DictVectorizer" + "bucket": "shaped", + "sklearn_estimator": "DictVectorizer", + "shape": { + "dataset": "classification", + "flow_preamble": [ + "let dv_counts: ptr = malloc((n_c as i64) * 4) as ptr", + "let dv_keys: ptr > = malloc((n_c as i64) * 8) as ptr >", + "let dv_vals: ptr > = malloc((n_c as i64) * 8) as ptr >", + "for i in 0 to n_c {", + " dv_counts[i] = f_c", + " let dv_kk: ptr = malloc((f_c as i64) * 4) as ptr", + " let dv_vv: ptr = array_new_f32(f_c)", + " for j in 0 to f_c {", + " dv_kk[j] = j", + " dv_vv[j] = matrix_at(X_c, i, j)", + " }", + " dv_keys[i] = dv_kk", + " dv_vals[i] = dv_vv", + "}" + ], + "flow_fit": [ + "dv_keys", + "dv_vals", + "n_c", + "dv_counts" + ], + "flow_work": [ + "dv_keys", + "dv_vals", + "n_c", + "dv_counts" + ], + "sklearn_input": "dicts" + } }, { "flow_estimator": "dictionary_learning", @@ -2536,9 +2616,27 @@ "fu", "X" ], - "bucket": "different_shape", - "reason": "takes FeatureUnion first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "FeatureUnion" + "bucket": "shaped", + "sklearn_estimator": "FeatureUnion", + "shape": { + "dataset": "classification", + "flow_preamble": [ + "let fu_obj: FeatureUnion = feature_union_init(2)", + "# 0 is FU_TRANSFORMER_STANDARD_SCALER and 2 is FU_TRANSFORMER_PASSTHROUGH,", + "# written out for the reason the column transformer above gives.", + "feature_union_set_transformer(fu_obj, 0, 0, 0)", + "feature_union_set_transformer(fu_obj, 1, 2, 0)" + ], + "flow_fit": [ + "fu_obj", + "X_c" + ], + "flow_work": [ + "X_c" + ], + "flow_free": "after", + "sklearn_input": "X" + } }, { "flow_estimator": "gamma_regressor", @@ -3471,7 +3569,7 @@ "X" ], "bucket": "different_shape", - "reason": "takes IncrementalPCA first, so it is not an estimator over a feature matrix and needs its own harness", + "reason": "one partial_fit step over one batch, where IncrementalPCA.fit walks the whole design in batches", "sklearn_estimator": "IncrementalPCA" }, { @@ -6879,9 +6977,27 @@ "n_labels_per_sample", "n_classes" ], - "bucket": "different_shape", - "reason": "takes ptr > first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "MultiLabelBinarizer" + "bucket": "shaped", + "sklearn_estimator": "MultiLabelBinarizer", + "shape": { + "dataset": "multioutput_class", + "flow_preamble": [ + "let mlb_counts: ptr = malloc((n_c as i64) * 4) as ptr", + "for i in 0 to n_c { mlb_counts[i] = 2 }" + ], + "flow_fit": [ + "Y_label_rows", + "n_c", + "mlb_counts", + "3" + ], + "flow_work": [ + "Y_label_rows", + "n_c", + "mlb_counts" + ], + "sklearn_input": "labelsets" + } }, { "flow_estimator": "multinomial_nb", @@ -8548,9 +8664,28 @@ "X", "y" ], - "bucket": "different_shape", - "reason": "takes Pipeline first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "Pipeline" + "bucket": "shaped", + "sklearn_estimator": "Pipeline", + "shape": { + "dataset": "classification", + "flow_preamble": [ + "let pipe_steps: array = [", + " step_standard_scaler(\"scaler\"),", + " step_logistic_regression(\"classifier\", 3, 50, 0.5, penalty_none())", + "]", + "let pipe_obj: Pipeline = pipeline_new(pipe_steps, 2)" + ], + "flow_fit": [ + "pipe_obj", + "X_c", + "y_c" + ], + "flow_work": [ + "X_c" + ], + "flow_free": "after", + "sklearn_input": "X" + } }, { "flow_estimator": "pls_canonical", @@ -10611,7 +10746,7 @@ "max_iter" ], "bucket": "different_shape", - "reason": "takes ptr first, so it is not an estimator over a feature matrix and needs its own harness", + "reason": "takes class probabilities as input, so it is the labelling loop alone while SelfTrainingClassifier also fits the base estimator on every round", "sklearn_estimator": "SelfTrainingClassifier" }, { @@ -11565,7 +11700,7 @@ "n_iter" ], "bucket": "different_shape", - "reason": "takes ptr > first, so it is not an estimator over a feature matrix and needs its own harness", + "reason": "takes the base estimators' predictions as input, so it is the meta-learner alone while StackingClassifier also fits the base estimators and cross-validates them", "sklearn_estimator": "StackingClassifier" }, { @@ -11968,9 +12103,39 @@ "n_docs", "max_features" ], - "bucket": "different_shape", - "reason": "takes ptr first, so it is not an estimator over a feature matrix and needs its own harness", - "sklearn_estimator": "TfidfVectorizer" + "bucket": "shaped", + "sklearn_estimator": "TfidfVectorizer", + "shape": { + "dataset": "classification", + "corpus": [ + "the quick brown fox jumps over the lazy dog", + "a lazy dog sleeps in the warm sun", + "quick brown foxes are rare in the city", + "the dog and the fox share a field", + "warm sun and a cold river run together", + "a field of brown grass in the sun", + "the city river runs past the old field", + "old dogs sleep through a quick storm", + "a storm over the city wakes the dog", + "foxes hunt in the cold river valley", + "the valley holds a warm field of grass", + "grass grows where the river meets the sun", + "a rare fox crosses the old stone bridge", + "the stone bridge over the cold river", + "dogs and foxes keep their distance here", + "here the field the river and the city meet" + ], + "flow_fit": [ + "tfidf_vectorizer_docs", + "16", + "50" + ], + "flow_work": [ + "tfidf_vectorizer_docs", + "16" + ], + "sklearn_input": "docs" + } }, { "flow_estimator": "theil_sen_regressor", @@ -12381,7 +12546,7 @@ "voting" ], "bucket": "different_shape", - "reason": "takes ptr first, so it is not an estimator over a feature matrix and needs its own harness", + "reason": "takes trees that are already fitted, so its fit is the vote alone while VotingClassifier fits every estimator it is given", "sklearn_estimator": "VotingClassifier" }, { @@ -12437,7 +12602,7 @@ "weights" ], "bucket": "different_shape", - "reason": "takes ptr first, so it is not an estimator over a feature matrix and needs its own harness", + "reason": "takes regressors that are already fitted, so its fit is the average alone while VotingRegressor fits every estimator it is given", "sklearn_estimator": "VotingRegressor" } ] diff --git a/benchmarks/estimator_coverage.py b/benchmarks/estimator_coverage.py index d969bee..5058c97 100644 --- a/benchmarks/estimator_coverage.py +++ b/benchmarks/estimator_coverage.py @@ -177,6 +177,45 @@ "n_categories": ("array", "[4, 4, 4, 4]"), } +# Flow functions whose scikit-learn namesake does more than they do. Timing +# them against that class would compare a part against the whole, which is the +# same objection the simplified rows carry. The wording says which part. +DIFFERENT_JOB: dict[str, str] = { + "stacking_classifier": "takes the base estimators' predictions as input, so it is the " + "meta-learner alone while StackingClassifier also fits the base " + "estimators and cross-validates them", + "self_training_classifier": "takes class probabilities as input, so it is the labelling " + "loop alone while SelfTrainingClassifier also fits the base " + "estimator on every round", + "voting_classifier": "takes trees that are already fitted, so its fit is the vote alone " + "while VotingClassifier fits every estimator it is given", + "voting_regressor": "takes regressors that are already fitted, so its fit is the average " + "alone while VotingRegressor fits every estimator it is given", + "incremental_pca_partial": "one partial_fit step over one batch, where IncrementalPCA.fit " + "walks the whole design in batches", +} + +# One corpus for the two text vectorizers, read by the Flow generator and by +# the scikit-learn harness, so the two sides cannot drift apart. +CORPUS: list[str] = [ + "the quick brown fox jumps over the lazy dog", + "a lazy dog sleeps in the warm sun", + "quick brown foxes are rare in the city", + "the dog and the fox share a field", + "warm sun and a cold river run together", + "a field of brown grass in the sun", + "the city river runs past the old field", + "old dogs sleep through a quick storm", + "a storm over the city wakes the dog", + "foxes hunt in the cold river valley", + "the valley holds a warm field of grass", + "grass grows where the river meets the sun", + "a rare fox crosses the old stone bridge", + "the stone bridge over the cold river", + "dogs and foxes keep their distance here", + "here the field the river and the city meet", +] + # Estimators whose fit does not begin with a feature matrix, and which race # scikit-learn perfectly well once the call is written out. Thirteen rows sat # in different_shape only because the generic path builds one call shape. @@ -272,6 +311,97 @@ "flow_work": ["x1d_r", "n_r"], "sklearn_input": "x1d", }, + "multilabel_binarizer": { + # The label rows, so the scikit-learn side gets sets of labels rather + # than one label per sample. + "dataset": "multioutput_class", + "flow_preamble": [ + "let mlb_counts: ptr = malloc((n_c as i64) * 4) as ptr", + "for i in 0 to n_c { mlb_counts[i] = 2 }", + ], + "flow_fit": ["Y_label_rows", "n_c", "mlb_counts", "3"], + "flow_work": ["Y_label_rows", "n_c", "mlb_counts"], + "sklearn_input": "labelsets", + }, + "dict_vectorizer": { + "dataset": "classification", + "flow_preamble": [ + "let dv_counts: ptr = malloc((n_c as i64) * 4) as ptr", + "let dv_keys: ptr > = malloc((n_c as i64) * 8) as ptr >", + "let dv_vals: ptr > = malloc((n_c as i64) * 8) as ptr >", + "for i in 0 to n_c {", + " dv_counts[i] = f_c", + " let dv_kk: ptr = malloc((f_c as i64) * 4) as ptr", + " let dv_vv: ptr = array_new_f32(f_c)", + " for j in 0 to f_c {", + " dv_kk[j] = j", + " dv_vv[j] = matrix_at(X_c, i, j)", + " }", + " dv_keys[i] = dv_kk", + " dv_vals[i] = dv_vv", + "}", + ], + "flow_fit": ["dv_keys", "dv_vals", "n_c", "dv_counts"], + "flow_work": ["dv_keys", "dv_vals", "n_c", "dv_counts"], + "sklearn_input": "dicts", + }, + "count_vectorizer": { + "dataset": "classification", + "corpus": CORPUS, + "flow_fit": ["count_vectorizer_docs", "16", "50"], + "flow_work": ["count_vectorizer_docs", "16"], + "sklearn_input": "docs", + }, + "tfidf_vectorizer": { + "dataset": "classification", + "corpus": CORPUS, + "flow_fit": ["tfidf_vectorizer_docs", "16", "50"], + "flow_work": ["tfidf_vectorizer_docs", "16"], + "sklearn_input": "docs", + }, + "pipeline": { + "dataset": "classification", + "flow_preamble": [ + "let pipe_steps: array = [", + ' step_standard_scaler("scaler"),', + ' step_logistic_regression("classifier", 3, 50, 0.5, penalty_none())', + "]", + "let pipe_obj: Pipeline = pipeline_new(pipe_steps, 2)", + ], + "flow_fit": ["pipe_obj", "X_c", "y_c"], + "flow_work": ["X_c"], + "flow_free": "after", + "sklearn_input": "X", + }, + "column_transformer": { + "dataset": "classification", + "flow_preamble": [ + "let ct_cols: ptr = malloc((f_c as i64) * 4) as ptr", + "for i in 0 to f_c { ct_cols[i] = i }", + "let ct_obj: ColumnTransformer = column_transformer_init(1)", + "# 0 is TRANSFORMER_STANDARD_SCALER. The constant is written out", + "# because an export const is not visible through the umbrella import.", + "column_transformer_set_spec(ct_obj, 0, 0, ct_cols, f_c)", + ], + "flow_fit": ["ct_obj", "X_c"], + "flow_work": ["X_c"], + "flow_free": "after", + "sklearn_input": "X", + }, + "feature_union": { + "dataset": "classification", + "flow_preamble": [ + "let fu_obj: FeatureUnion = feature_union_init(2)", + "# 0 is FU_TRANSFORMER_STANDARD_SCALER and 2 is FU_TRANSFORMER_PASSTHROUGH,", + "# written out for the reason the column transformer above gives.", + "feature_union_set_transformer(fu_obj, 0, 0, 0)", + "feature_union_set_transformer(fu_obj, 1, 2, 0)", + ], + "flow_fit": ["fu_obj", "X_c"], + "flow_work": ["X_c"], + "flow_free": "after", + "sklearn_input": "X", + }, } # Parameters whose value depends on the dataset rather than on a constant. @@ -404,6 +534,9 @@ def classify(base: str, spec: dict, known: set[str], exports: dict[str, dict]) - if base in SHAPED and sk is not None: entry.update(bucket="shaped", sklearn_estimator=sk, shape=SHAPED[base]) return entry + if base in DIFFERENT_JOB: + entry.update(bucket="different_shape", reason=DIFFERENT_JOB[base], sklearn_estimator=sk) + return entry if head != "Matrix": entry.update( bucket="different_shape", diff --git a/benchmarks/generate_estimator_bench.py b/benchmarks/generate_estimator_bench.py index ffb1651..2a582cc 100644 --- a/benchmarks/generate_estimator_bench.py +++ b/benchmarks/generate_estimator_bench.py @@ -223,16 +223,31 @@ def shaped_block(entry: dict) -> str: ret = entry["fit"]["returns"] call = f"{entry['fit']['name']}({', '.join(shape['flow_fit'])})" free = entry["companions"].get("free") + # A fit that takes a composed object built outside the timing mutates that + # object and hands it back, so freeing the result on every repeat would free + # what the next repeat is about to read. Those rows free once, after the + # timing, and leak the state a repeat leaves behind, which is a few hundred + # bytes per pass over a four-column design. + free_in_loop = shape.get("flow_free", "loop") == "loop" free_ok = free is not None and len(free["parameters"]) == 1 comp = entry["companions"].get("transform") or entry["companions"].get("predict") work = shape.get("flow_work") comp_ok = comp is not None and work is not None and comp["returns"] in ("Matrix", "ptr") lines = [f" # ---- {name} ({kind}, written out) ----"] + if shape.get("corpus"): + docs = shape["corpus"] + lines.append(f" let {name}_docs: array = [") + for i, doc in enumerate(docs): + tail = "," if i + 1 < len(docs) else "" + lines.append(f' "{doc}"{tail}') + lines.append(" ]") + for pre_line in shape.get("flow_preamble", []): + lines.append(" " + pre_line) lines.append(" t0 = flow_now_ns()") lines.append(f" let probe_{name}: {ret} = {call}") lines.append(" t1 = flow_now_ns()") - if free_ok: + if free_ok and free_in_loop: lines.append(f" {free['name']}(probe_{name})") lines.append(" reps = 1") lines.append(" if (t1 - t0) < 200000 { reps = 200 }") @@ -240,7 +255,7 @@ def shaped_block(entry: dict) -> str: lines.append(" t0 = flow_now_ns()") lines.append(" for rep in 0 to reps {") lines.append(f" let m_{name}: {ret} = {call}") - if free_ok: + if free_ok and free_in_loop: lines.append(f" {free['name']}(m_{name})") lines.append(" }") lines.append(" t1 = flow_now_ns()") @@ -261,6 +276,8 @@ def shaped_block(entry: dict) -> str: pred_expr = "ms_between(t2, t3) / (reps as f32)" else: pred_expr = "0.0" + if free_ok and not free_in_loop: + lines.append(f" {free['name']}(probe_{name})") lines.append( f' printf("ESTIMATOR|{name}|%.9f|%.9f|%d|ok\\n", ' diff --git a/benchmarks/generated/bench_estimators_00.flow b/benchmarks/generated/bench_estimators_00.flow index 3c09d03..364905e 100644 --- a/benchmarks/generated/bench_estimators_00.flow +++ b/benchmarks/generated/bench_estimators_00.flow @@ -512,47 +512,62 @@ function main() -> i32 { printf("ESTIMATOR|classifier_chain|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) fflush(null) - # ---- complement_nb (classification) ---- + # ---- column_transformer (classification, written out) ---- + let ct_cols: ptr = malloc((f_c as i64) * 4) as ptr + for i in 0 to f_c { ct_cols[i] = i } + let ct_obj: ColumnTransformer = column_transformer_init(1) + # 0 is TRANSFORMER_STANDARD_SCALER. The constant is written out + # because an export const is not visible through the umbrella import. + column_transformer_set_spec(ct_obj, 0, 0, ct_cols, f_c) t0 = flow_now_ns() - let probe_complement_nb: ComplementNB = complement_nb_fit(X_c, y_c, 3, 1.0) + let probe_column_transformer: ColumnTransformer = column_transformer_fit(ct_obj, X_c) t1 = flow_now_ns() - complement_nb_free(probe_complement_nb) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_complement_nb: ComplementNB = complement_nb_fit(X_c, y_c, 3, 1.0) - complement_nb_free(m_complement_nb) + let m_column_transformer: ColumnTransformer = column_transformer_fit(ct_obj, X_c) } t1 = flow_now_ns() - let fitted_complement_nb: ComplementNB = complement_nb_fit(X_c, y_c, 3, 1.0) + let fitted_column_transformer: ColumnTransformer = column_transformer_fit(ct_obj, X_c) t2 = flow_now_ns() for rep2 in 0 to reps { - let o_complement_nb: ptr = complement_nb_predict(fitted_complement_nb, X_c) - sink = sink + o_complement_nb[0] - array_free_f32(o_complement_nb) + let o_column_transformer: Matrix = column_transformer_transform(fitted_column_transformer, X_c) + if o_column_transformer.rows > 0 { + if o_column_transformer.cols > 0 { sink = sink + o_column_transformer.data[0] } + } + matrix_free(o_column_transformer) } t3 = flow_now_ns() - complement_nb_free(fitted_complement_nb) - printf("ESTIMATOR|complement_nb|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + column_transformer_free(fitted_column_transformer) + printf("ESTIMATOR|column_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- dbscan (unsupervised) ---- + # ---- complement_nb (classification) ---- t0 = flow_now_ns() - let probe_dbscan: DBSCAN = dbscan_fit(X_c, 0.5, 5) + let probe_complement_nb: ComplementNB = complement_nb_fit(X_c, y_c, 3, 1.0) t1 = flow_now_ns() - dbscan_free(probe_dbscan) + complement_nb_free(probe_complement_nb) reps = 1 if (t1 - t0) < 200000 { reps = 200 } elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_dbscan: DBSCAN = dbscan_fit(X_c, 0.5, 5) - dbscan_free(m_dbscan) + let m_complement_nb: ComplementNB = complement_nb_fit(X_c, y_c, 3, 1.0) + complement_nb_free(m_complement_nb) } t1 = flow_now_ns() - printf("ESTIMATOR|dbscan|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + let fitted_complement_nb: ComplementNB = complement_nb_fit(X_c, y_c, 3, 1.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_complement_nb: ptr = complement_nb_predict(fitted_complement_nb, X_c) + sink = sink + o_complement_nb[0] + array_free_f32(o_complement_nb) + } + t3 = flow_now_ns() + complement_nb_free(fitted_complement_nb) + printf("ESTIMATOR|complement_nb|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } diff --git a/benchmarks/generated/bench_estimators_01.flow b/benchmarks/generated/bench_estimators_01.flow index dcf0255..00810bc 100644 --- a/benchmarks/generated/bench_estimators_01.flow +++ b/benchmarks/generated/bench_estimators_01.flow @@ -82,6 +82,69 @@ function main() -> i32 { let mut reps: i32 = 1 let mut sink: f32 = 0.0 + # ---- count_vectorizer (classification, written out) ---- + let count_vectorizer_docs: array = [ + "the quick brown fox jumps over the lazy dog", + "a lazy dog sleeps in the warm sun", + "quick brown foxes are rare in the city", + "the dog and the fox share a field", + "warm sun and a cold river run together", + "a field of brown grass in the sun", + "the city river runs past the old field", + "old dogs sleep through a quick storm", + "a storm over the city wakes the dog", + "foxes hunt in the cold river valley", + "the valley holds a warm field of grass", + "grass grows where the river meets the sun", + "a rare fox crosses the old stone bridge", + "the stone bridge over the cold river", + "dogs and foxes keep their distance here", + "here the field the river and the city meet" + ] + t0 = flow_now_ns() + let probe_count_vectorizer: CountVectorizer = count_vectorizer_fit(count_vectorizer_docs, 16, 50) + t1 = flow_now_ns() + count_vectorizer_free(probe_count_vectorizer) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_count_vectorizer: CountVectorizer = count_vectorizer_fit(count_vectorizer_docs, 16, 50) + count_vectorizer_free(m_count_vectorizer) + } + t1 = flow_now_ns() + let fitted_count_vectorizer: CountVectorizer = count_vectorizer_fit(count_vectorizer_docs, 16, 50) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_count_vectorizer: Matrix = count_vectorizer_transform(fitted_count_vectorizer, count_vectorizer_docs, 16) + if o_count_vectorizer.rows > 0 { + if o_count_vectorizer.cols > 0 { sink = sink + o_count_vectorizer.data[0] } + } + matrix_free(o_count_vectorizer) + } + t3 = flow_now_ns() + count_vectorizer_free(fitted_count_vectorizer) + printf("ESTIMATOR|count_vectorizer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- dbscan (unsupervised) ---- + t0 = flow_now_ns() + let probe_dbscan: DBSCAN = dbscan_fit(X_c, 0.5, 5) + t1 = flow_now_ns() + dbscan_free(probe_dbscan) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_dbscan: DBSCAN = dbscan_fit(X_c, 0.5, 5) + dbscan_free(m_dbscan) + } + t1 = flow_now_ns() + printf("ESTIMATOR|dbscan|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + # ---- decision_tree_classifier (classification) ---- t0 = flow_now_ns() let probe_decision_tree_classifier: DecisionTreeClassifier = decision_tree_classifier_fit(X_c, y_c, 3, 5, 0) @@ -134,6 +197,48 @@ function main() -> i32 { printf("ESTIMATOR|decision_tree_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) + # ---- dict_vectorizer (classification, written out) ---- + let dv_counts: ptr = malloc((n_c as i64) * 4) as ptr + let dv_keys: ptr > = malloc((n_c as i64) * 8) as ptr > + let dv_vals: ptr > = malloc((n_c as i64) * 8) as ptr > + for i in 0 to n_c { + dv_counts[i] = f_c + let dv_kk: ptr = malloc((f_c as i64) * 4) as ptr + let dv_vv: ptr = array_new_f32(f_c) + for j in 0 to f_c { + dv_kk[j] = j + dv_vv[j] = matrix_at(X_c, i, j) + } + dv_keys[i] = dv_kk + dv_vals[i] = dv_vv + } + t0 = flow_now_ns() + let probe_dict_vectorizer: DictVectorizer = dict_vectorizer_fit(dv_keys, dv_vals, n_c, dv_counts) + t1 = flow_now_ns() + dict_vectorizer_free(probe_dict_vectorizer) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_dict_vectorizer: DictVectorizer = dict_vectorizer_fit(dv_keys, dv_vals, n_c, dv_counts) + dict_vectorizer_free(m_dict_vectorizer) + } + t1 = flow_now_ns() + let fitted_dict_vectorizer: DictVectorizer = dict_vectorizer_fit(dv_keys, dv_vals, n_c, dv_counts) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_dict_vectorizer: Matrix = dict_vectorizer_transform(fitted_dict_vectorizer, dv_keys, dv_vals, n_c, dv_counts) + if o_dict_vectorizer.rows > 0 { + if o_dict_vectorizer.cols > 0 { sink = sink + o_dict_vectorizer.data[0] } + } + matrix_free(o_dict_vectorizer) + } + t3 = flow_now_ns() + dict_vectorizer_free(fitted_dict_vectorizer) + printf("ESTIMATOR|dict_vectorizer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- dictionary_learning (unsupervised) ---- t0 = flow_now_ns() let probe_dictionary_learning: DictionaryLearning = dictionary_learning_fit(X_c, 2, 1.0, 100, 0.0001, 42) @@ -523,75 +628,6 @@ function main() -> i32 { printf("ESTIMATOR|feature_agglomeration|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- gamma_regressor (regression) ---- - t0 = flow_now_ns() - let probe_gamma_regressor: GammaRegressor = gamma_regressor_fit(X_r, y_r, 1.0, 100, 0.01) - t1 = flow_now_ns() - gamma_regressor_free(probe_gamma_regressor) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_gamma_regressor: GammaRegressor = gamma_regressor_fit(X_r, y_r, 1.0, 100, 0.01) - gamma_regressor_free(m_gamma_regressor) - } - t1 = flow_now_ns() - let fitted_gamma_regressor: GammaRegressor = gamma_regressor_fit(X_r, y_r, 1.0, 100, 0.01) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_gamma_regressor: ptr = gamma_regressor_predict(fitted_gamma_regressor, X_r) - sink = sink + o_gamma_regressor[0] - array_free_f32(o_gamma_regressor) - } - t3 = flow_now_ns() - gamma_regressor_free(fitted_gamma_regressor) - printf("ESTIMATOR|gamma_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- gaussian_mixture (unsupervised) ---- - t0 = flow_now_ns() - let probe_gaussian_mixture: GaussianMixture = gaussian_mixture_fit(X_c, 2, 100, 0.0001, 42) - t1 = flow_now_ns() - gaussian_mixture_free(probe_gaussian_mixture) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_gaussian_mixture: GaussianMixture = gaussian_mixture_fit(X_c, 2, 100, 0.0001, 42) - gaussian_mixture_free(m_gaussian_mixture) - } - t1 = flow_now_ns() - printf("ESTIMATOR|gaussian_mixture|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) - fflush(null) - - # ---- gaussian_nb (classification) ---- - t0 = flow_now_ns() - let probe_gaussian_nb: GaussianNB = gaussian_nb_fit(X_c, y_c, 3, 0.000000001) - t1 = flow_now_ns() - gaussian_nb_free(probe_gaussian_nb) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_gaussian_nb: GaussianNB = gaussian_nb_fit(X_c, y_c, 3, 0.000000001) - gaussian_nb_free(m_gaussian_nb) - } - t1 = flow_now_ns() - let fitted_gaussian_nb: GaussianNB = gaussian_nb_fit(X_c, y_c, 3, 0.000000001) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_gaussian_nb: ptr = gaussian_nb_predict(fitted_gaussian_nb, X_c) - sink = sink + o_gaussian_nb[0] - array_free_f32(o_gaussian_nb) - } - t3 = flow_now_ns() - gaussian_nb_free(fitted_gaussian_nb) - printf("ESTIMATOR|gaussian_nb|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) diff --git a/benchmarks/generated/bench_estimators_02.flow b/benchmarks/generated/bench_estimators_02.flow index 5d45b94..2bc53bb 100644 --- a/benchmarks/generated/bench_estimators_02.flow +++ b/benchmarks/generated/bench_estimators_02.flow @@ -82,6 +82,106 @@ function main() -> i32 { let mut reps: i32 = 1 let mut sink: f32 = 0.0 + # ---- feature_union (classification, written out) ---- + let fu_obj: FeatureUnion = feature_union_init(2) + # 0 is FU_TRANSFORMER_STANDARD_SCALER and 2 is FU_TRANSFORMER_PASSTHROUGH, + # written out for the reason the column transformer above gives. + feature_union_set_transformer(fu_obj, 0, 0, 0) + feature_union_set_transformer(fu_obj, 1, 2, 0) + t0 = flow_now_ns() + let probe_feature_union: FeatureUnion = feature_union_fit(fu_obj, X_c) + t1 = flow_now_ns() + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_feature_union: FeatureUnion = feature_union_fit(fu_obj, X_c) + } + t1 = flow_now_ns() + let fitted_feature_union: FeatureUnion = feature_union_fit(fu_obj, X_c) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_feature_union: Matrix = feature_union_transform(fitted_feature_union, X_c) + if o_feature_union.rows > 0 { + if o_feature_union.cols > 0 { sink = sink + o_feature_union.data[0] } + } + matrix_free(o_feature_union) + } + t3 = flow_now_ns() + feature_union_free(fitted_feature_union) + printf("ESTIMATOR|feature_union|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- gamma_regressor (regression) ---- + t0 = flow_now_ns() + let probe_gamma_regressor: GammaRegressor = gamma_regressor_fit(X_r, y_r, 1.0, 100, 0.01) + t1 = flow_now_ns() + gamma_regressor_free(probe_gamma_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_gamma_regressor: GammaRegressor = gamma_regressor_fit(X_r, y_r, 1.0, 100, 0.01) + gamma_regressor_free(m_gamma_regressor) + } + t1 = flow_now_ns() + let fitted_gamma_regressor: GammaRegressor = gamma_regressor_fit(X_r, y_r, 1.0, 100, 0.01) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_gamma_regressor: ptr = gamma_regressor_predict(fitted_gamma_regressor, X_r) + sink = sink + o_gamma_regressor[0] + array_free_f32(o_gamma_regressor) + } + t3 = flow_now_ns() + gamma_regressor_free(fitted_gamma_regressor) + printf("ESTIMATOR|gamma_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- gaussian_mixture (unsupervised) ---- + t0 = flow_now_ns() + let probe_gaussian_mixture: GaussianMixture = gaussian_mixture_fit(X_c, 2, 100, 0.0001, 42) + t1 = flow_now_ns() + gaussian_mixture_free(probe_gaussian_mixture) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_gaussian_mixture: GaussianMixture = gaussian_mixture_fit(X_c, 2, 100, 0.0001, 42) + gaussian_mixture_free(m_gaussian_mixture) + } + t1 = flow_now_ns() + printf("ESTIMATOR|gaussian_mixture|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + + # ---- gaussian_nb (classification) ---- + t0 = flow_now_ns() + let probe_gaussian_nb: GaussianNB = gaussian_nb_fit(X_c, y_c, 3, 0.000000001) + t1 = flow_now_ns() + gaussian_nb_free(probe_gaussian_nb) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_gaussian_nb: GaussianNB = gaussian_nb_fit(X_c, y_c, 3, 0.000000001) + gaussian_nb_free(m_gaussian_nb) + } + t1 = flow_now_ns() + let fitted_gaussian_nb: GaussianNB = gaussian_nb_fit(X_c, y_c, 3, 0.000000001) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_gaussian_nb: ptr = gaussian_nb_predict(fitted_gaussian_nb, X_c) + sink = sink + o_gaussian_nb[0] + array_free_f32(o_gaussian_nb) + } + t3 = flow_now_ns() + gaussian_nb_free(fitted_gaussian_nb) + printf("ESTIMATOR|gaussian_nb|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- gaussian_process_classifier (regression) ---- t0 = flow_now_ns() let probe_gaussian_process_classifier: GaussianProcessClassifier = gaussian_process_classifier_fit(X_r, y_r, 0.1, 100) @@ -459,112 +559,6 @@ function main() -> i32 { printf("ESTIMATOR|kernel_density|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) fflush(null) - # ---- kernel_pca (unsupervised) ---- - t0 = flow_now_ns() - let probe_kernel_pca: KernelPCA = kernel_pca_fit(X_c, 2, 0, 0.1, 2, 0.0) - t1 = flow_now_ns() - kernel_pca_free(probe_kernel_pca) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_kernel_pca: KernelPCA = kernel_pca_fit(X_c, 2, 0, 0.1, 2, 0.0) - kernel_pca_free(m_kernel_pca) - } - t1 = flow_now_ns() - let fitted_kernel_pca: KernelPCA = kernel_pca_fit(X_c, 2, 0, 0.1, 2, 0.0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_kernel_pca: Matrix = kernel_pca_transform(fitted_kernel_pca, X_c) - if o_kernel_pca.rows > 0 { - if o_kernel_pca.cols > 0 { sink = sink + o_kernel_pca.data[0] } - } - matrix_free(o_kernel_pca) - } - t3 = flow_now_ns() - kernel_pca_free(fitted_kernel_pca) - printf("ESTIMATOR|kernel_pca|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- kernel_ridge (regression) ---- - t0 = flow_now_ns() - let probe_kernel_ridge: KernelRidge = kernel_ridge_fit(X_r, y_r, 1.0, 0, 0.1, 2, 0.0) - t1 = flow_now_ns() - kernel_ridge_free(probe_kernel_ridge) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_kernel_ridge: KernelRidge = kernel_ridge_fit(X_r, y_r, 1.0, 0, 0.1, 2, 0.0) - kernel_ridge_free(m_kernel_ridge) - } - t1 = flow_now_ns() - let fitted_kernel_ridge: KernelRidge = kernel_ridge_fit(X_r, y_r, 1.0, 0, 0.1, 2, 0.0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_kernel_ridge: ptr = kernel_ridge_predict(fitted_kernel_ridge, X_r) - sink = sink + o_kernel_ridge[0] - array_free_f32(o_kernel_ridge) - } - t3 = flow_now_ns() - kernel_ridge_free(fitted_kernel_ridge) - printf("ESTIMATOR|kernel_ridge|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- kernel_svc (classification) ---- - t0 = flow_now_ns() - let probe_kernel_svc: KernelSVC = kernel_svc_fit(X_c, y_c, 3, 1.0, 0.1, 100) - t1 = flow_now_ns() - kernel_svc_free(probe_kernel_svc) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_kernel_svc: KernelSVC = kernel_svc_fit(X_c, y_c, 3, 1.0, 0.1, 100) - kernel_svc_free(m_kernel_svc) - } - t1 = flow_now_ns() - let fitted_kernel_svc: KernelSVC = kernel_svc_fit(X_c, y_c, 3, 1.0, 0.1, 100) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_kernel_svc: ptr = kernel_svc_predict(fitted_kernel_svc, X_c) - sink = sink + o_kernel_svc[0] - array_free_f32(o_kernel_svc) - } - t3 = flow_now_ns() - kernel_svc_free(fitted_kernel_svc) - printf("ESTIMATOR|kernel_svc|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- kernel_svc_multi (classification) ---- - t0 = flow_now_ns() - let probe_kernel_svc_multi: KernelSVCMulti = kernel_svc_multi_fit(X_c, y_c, 3, 0.1, 1.0, 100) - t1 = flow_now_ns() - kernel_svc_multi_free(probe_kernel_svc_multi) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_kernel_svc_multi: KernelSVCMulti = kernel_svc_multi_fit(X_c, y_c, 3, 0.1, 1.0, 100) - kernel_svc_multi_free(m_kernel_svc_multi) - } - t1 = flow_now_ns() - let fitted_kernel_svc_multi: KernelSVCMulti = kernel_svc_multi_fit(X_c, y_c, 3, 0.1, 1.0, 100) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_kernel_svc_multi: ptr = kernel_svc_multi_predict(fitted_kernel_svc_multi, X_c) - sink = sink + o_kernel_svc_multi[0] - array_free_f32(o_kernel_svc_multi) - } - t3 = flow_now_ns() - kernel_svc_multi_free(fitted_kernel_svc_multi) - printf("ESTIMATOR|kernel_svc_multi|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) diff --git a/benchmarks/generated/bench_estimators_03.flow b/benchmarks/generated/bench_estimators_03.flow index 3f797ba..c97541c 100644 --- a/benchmarks/generated/bench_estimators_03.flow +++ b/benchmarks/generated/bench_estimators_03.flow @@ -82,6 +82,112 @@ function main() -> i32 { let mut reps: i32 = 1 let mut sink: f32 = 0.0 + # ---- kernel_pca (unsupervised) ---- + t0 = flow_now_ns() + let probe_kernel_pca: KernelPCA = kernel_pca_fit(X_c, 2, 0, 0.1, 2, 0.0) + t1 = flow_now_ns() + kernel_pca_free(probe_kernel_pca) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_kernel_pca: KernelPCA = kernel_pca_fit(X_c, 2, 0, 0.1, 2, 0.0) + kernel_pca_free(m_kernel_pca) + } + t1 = flow_now_ns() + let fitted_kernel_pca: KernelPCA = kernel_pca_fit(X_c, 2, 0, 0.1, 2, 0.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_kernel_pca: Matrix = kernel_pca_transform(fitted_kernel_pca, X_c) + if o_kernel_pca.rows > 0 { + if o_kernel_pca.cols > 0 { sink = sink + o_kernel_pca.data[0] } + } + matrix_free(o_kernel_pca) + } + t3 = flow_now_ns() + kernel_pca_free(fitted_kernel_pca) + printf("ESTIMATOR|kernel_pca|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- kernel_ridge (regression) ---- + t0 = flow_now_ns() + let probe_kernel_ridge: KernelRidge = kernel_ridge_fit(X_r, y_r, 1.0, 0, 0.1, 2, 0.0) + t1 = flow_now_ns() + kernel_ridge_free(probe_kernel_ridge) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_kernel_ridge: KernelRidge = kernel_ridge_fit(X_r, y_r, 1.0, 0, 0.1, 2, 0.0) + kernel_ridge_free(m_kernel_ridge) + } + t1 = flow_now_ns() + let fitted_kernel_ridge: KernelRidge = kernel_ridge_fit(X_r, y_r, 1.0, 0, 0.1, 2, 0.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_kernel_ridge: ptr = kernel_ridge_predict(fitted_kernel_ridge, X_r) + sink = sink + o_kernel_ridge[0] + array_free_f32(o_kernel_ridge) + } + t3 = flow_now_ns() + kernel_ridge_free(fitted_kernel_ridge) + printf("ESTIMATOR|kernel_ridge|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- kernel_svc (classification) ---- + t0 = flow_now_ns() + let probe_kernel_svc: KernelSVC = kernel_svc_fit(X_c, y_c, 3, 1.0, 0.1, 100) + t1 = flow_now_ns() + kernel_svc_free(probe_kernel_svc) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_kernel_svc: KernelSVC = kernel_svc_fit(X_c, y_c, 3, 1.0, 0.1, 100) + kernel_svc_free(m_kernel_svc) + } + t1 = flow_now_ns() + let fitted_kernel_svc: KernelSVC = kernel_svc_fit(X_c, y_c, 3, 1.0, 0.1, 100) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_kernel_svc: ptr = kernel_svc_predict(fitted_kernel_svc, X_c) + sink = sink + o_kernel_svc[0] + array_free_f32(o_kernel_svc) + } + t3 = flow_now_ns() + kernel_svc_free(fitted_kernel_svc) + printf("ESTIMATOR|kernel_svc|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- kernel_svc_multi (classification) ---- + t0 = flow_now_ns() + let probe_kernel_svc_multi: KernelSVCMulti = kernel_svc_multi_fit(X_c, y_c, 3, 0.1, 1.0, 100) + t1 = flow_now_ns() + kernel_svc_multi_free(probe_kernel_svc_multi) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_kernel_svc_multi: KernelSVCMulti = kernel_svc_multi_fit(X_c, y_c, 3, 0.1, 1.0, 100) + kernel_svc_multi_free(m_kernel_svc_multi) + } + t1 = flow_now_ns() + let fitted_kernel_svc_multi: KernelSVCMulti = kernel_svc_multi_fit(X_c, y_c, 3, 0.1, 1.0, 100) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_kernel_svc_multi: ptr = kernel_svc_multi_predict(fitted_kernel_svc_multi, X_c) + sink = sink + o_kernel_svc_multi[0] + array_free_f32(o_kernel_svc_multi) + } + t3 = flow_now_ns() + kernel_svc_multi_free(fitted_kernel_svc_multi) + printf("ESTIMATOR|kernel_svc_multi|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- kmeans (unsupervised) ---- t0 = flow_now_ns() let probe_kmeans: KMeans = kmeans_fit(X_c, 3, 100, 0.0001, 42) @@ -479,101 +585,6 @@ function main() -> i32 { printf("ESTIMATOR|lasso_lars_ic|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- lda (unsupervised) ---- - t0 = flow_now_ns() - let probe_lda: LatentDirichletAllocation = lda_fit(X_c, 3, 50, 1.0, 0.1, 42) - t1 = flow_now_ns() - lda_free(probe_lda) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_lda: LatentDirichletAllocation = lda_fit(X_c, 3, 50, 1.0, 0.1, 42) - lda_free(m_lda) - } - t1 = flow_now_ns() - let fitted_lda: LatentDirichletAllocation = lda_fit(X_c, 3, 50, 1.0, 0.1, 42) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_lda: Matrix = lda_transform(fitted_lda, X_c) - if o_lda.rows > 0 { - if o_lda.cols > 0 { sink = sink + o_lda.data[0] } - } - matrix_free(o_lda) - } - t3 = flow_now_ns() - lda_free(fitted_lda) - printf("ESTIMATOR|lda|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- ledoit_wolf_estimator (unsupervised) ---- - t0 = flow_now_ns() - let probe_ledoit_wolf_estimator: ShrunkCovariance = ledoit_wolf_estimator_fit(X_c) - t1 = flow_now_ns() - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_ledoit_wolf_estimator: ShrunkCovariance = ledoit_wolf_estimator_fit(X_c) - } - t1 = flow_now_ns() - printf("ESTIMATOR|ledoit_wolf_estimator|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) - fflush(null) - - # ---- linear_regression (regression) ---- - t0 = flow_now_ns() - let probe_linear_regression: LinearRegression = linear_regression_fit(X_r, y_r, penalty_none()) - t1 = flow_now_ns() - linear_regression_free(probe_linear_regression) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_linear_regression: LinearRegression = linear_regression_fit(X_r, y_r, penalty_none()) - linear_regression_free(m_linear_regression) - } - t1 = flow_now_ns() - let fitted_linear_regression: LinearRegression = linear_regression_fit(X_r, y_r, penalty_none()) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_linear_regression: ptr = linear_regression_predict(fitted_linear_regression, X_r) - sink = sink + o_linear_regression[0] - array_free_f32(o_linear_regression) - } - t3 = flow_now_ns() - linear_regression_free(fitted_linear_regression) - printf("ESTIMATOR|linear_regression|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- linear_svc (classification) ---- - t0 = flow_now_ns() - let probe_linear_svc: LinearSVC = linear_svc_fit(X_c, y_c, 3, 1.0, 50, 0.01) - t1 = flow_now_ns() - linear_svc_free(probe_linear_svc) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_linear_svc: LinearSVC = linear_svc_fit(X_c, y_c, 3, 1.0, 50, 0.01) - linear_svc_free(m_linear_svc) - } - t1 = flow_now_ns() - let fitted_linear_svc: LinearSVC = linear_svc_fit(X_c, y_c, 3, 1.0, 50, 0.01) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_linear_svc: ptr = linear_svc_predict(fitted_linear_svc, X_c) - sink = sink + o_linear_svc[0] - array_free_f32(o_linear_svc) - } - t3 = flow_now_ns() - linear_svc_free(fitted_linear_svc) - printf("ESTIMATOR|linear_svc|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) diff --git a/benchmarks/generated/bench_estimators_04.flow b/benchmarks/generated/bench_estimators_04.flow index 1da5bfe..b57b814 100644 --- a/benchmarks/generated/bench_estimators_04.flow +++ b/benchmarks/generated/bench_estimators_04.flow @@ -82,6 +82,101 @@ function main() -> i32 { let mut reps: i32 = 1 let mut sink: f32 = 0.0 + # ---- lda (unsupervised) ---- + t0 = flow_now_ns() + let probe_lda: LatentDirichletAllocation = lda_fit(X_c, 3, 50, 1.0, 0.1, 42) + t1 = flow_now_ns() + lda_free(probe_lda) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_lda: LatentDirichletAllocation = lda_fit(X_c, 3, 50, 1.0, 0.1, 42) + lda_free(m_lda) + } + t1 = flow_now_ns() + let fitted_lda: LatentDirichletAllocation = lda_fit(X_c, 3, 50, 1.0, 0.1, 42) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_lda: Matrix = lda_transform(fitted_lda, X_c) + if o_lda.rows > 0 { + if o_lda.cols > 0 { sink = sink + o_lda.data[0] } + } + matrix_free(o_lda) + } + t3 = flow_now_ns() + lda_free(fitted_lda) + printf("ESTIMATOR|lda|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- ledoit_wolf_estimator (unsupervised) ---- + t0 = flow_now_ns() + let probe_ledoit_wolf_estimator: ShrunkCovariance = ledoit_wolf_estimator_fit(X_c) + t1 = flow_now_ns() + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_ledoit_wolf_estimator: ShrunkCovariance = ledoit_wolf_estimator_fit(X_c) + } + t1 = flow_now_ns() + printf("ESTIMATOR|ledoit_wolf_estimator|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + + # ---- linear_regression (regression) ---- + t0 = flow_now_ns() + let probe_linear_regression: LinearRegression = linear_regression_fit(X_r, y_r, penalty_none()) + t1 = flow_now_ns() + linear_regression_free(probe_linear_regression) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_linear_regression: LinearRegression = linear_regression_fit(X_r, y_r, penalty_none()) + linear_regression_free(m_linear_regression) + } + t1 = flow_now_ns() + let fitted_linear_regression: LinearRegression = linear_regression_fit(X_r, y_r, penalty_none()) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_linear_regression: ptr = linear_regression_predict(fitted_linear_regression, X_r) + sink = sink + o_linear_regression[0] + array_free_f32(o_linear_regression) + } + t3 = flow_now_ns() + linear_regression_free(fitted_linear_regression) + printf("ESTIMATOR|linear_regression|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- linear_svc (classification) ---- + t0 = flow_now_ns() + let probe_linear_svc: LinearSVC = linear_svc_fit(X_c, y_c, 3, 1.0, 50, 0.01) + t1 = flow_now_ns() + linear_svc_free(probe_linear_svc) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_linear_svc: LinearSVC = linear_svc_fit(X_c, y_c, 3, 1.0, 50, 0.01) + linear_svc_free(m_linear_svc) + } + t1 = flow_now_ns() + let fitted_linear_svc: LinearSVC = linear_svc_fit(X_c, y_c, 3, 1.0, 50, 0.01) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_linear_svc: ptr = linear_svc_predict(fitted_linear_svc, X_c) + sink = sink + o_linear_svc[0] + array_free_f32(o_linear_svc) + } + t3 = flow_now_ns() + linear_svc_free(fitted_linear_svc) + printf("ESTIMATOR|linear_svc|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- linear_svc_multi (classification) ---- t0 = flow_now_ns() let probe_linear_svc_multi: LinearSVCMulti = linear_svc_multi_fit(X_c, y_c, 3, 1.0, 50) @@ -447,118 +542,6 @@ function main() -> i32 { printf("ESTIMATOR|missing_indicator|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- mlp_classifier (classification) ---- - let hidden_sizes_mlp_classifier: array = [8] - t0 = flow_now_ns() - let probe_mlp_classifier: MLPClassifier = mlp_classifier_fit(X_c, y_c, 3, hidden_sizes_mlp_classifier, 1, 0, 50, 0.01, 0.9, 42) - t1 = flow_now_ns() - mlp_classifier_free(probe_mlp_classifier) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_mlp_classifier: MLPClassifier = mlp_classifier_fit(X_c, y_c, 3, hidden_sizes_mlp_classifier, 1, 0, 50, 0.01, 0.9, 42) - mlp_classifier_free(m_mlp_classifier) - } - t1 = flow_now_ns() - let fitted_mlp_classifier: MLPClassifier = mlp_classifier_fit(X_c, y_c, 3, hidden_sizes_mlp_classifier, 1, 0, 50, 0.01, 0.9, 42) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_mlp_classifier: Matrix = mlp_classifier_predict(fitted_mlp_classifier, X_c) - if o_mlp_classifier.rows > 0 { - if o_mlp_classifier.cols > 0 { sink = sink + o_mlp_classifier.data[0] } - } - matrix_free(o_mlp_classifier) - } - t3 = flow_now_ns() - mlp_classifier_free(fitted_mlp_classifier) - printf("ESTIMATOR|mlp_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- mlp_regressor (regression) ---- - let hidden_sizes_mlp_regressor: array = [8] - t0 = flow_now_ns() - let probe_mlp_regressor: MLPRegressor = mlp_regressor_fit(X_r, y_r, hidden_sizes_mlp_regressor, 1, 0, 50, 0.01, 0.9, 42) - t1 = flow_now_ns() - mlp_regressor_free(probe_mlp_regressor) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_mlp_regressor: MLPRegressor = mlp_regressor_fit(X_r, y_r, hidden_sizes_mlp_regressor, 1, 0, 50, 0.01, 0.9, 42) - mlp_regressor_free(m_mlp_regressor) - } - t1 = flow_now_ns() - let fitted_mlp_regressor: MLPRegressor = mlp_regressor_fit(X_r, y_r, hidden_sizes_mlp_regressor, 1, 0, 50, 0.01, 0.9, 42) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_mlp_regressor: ptr = mlp_regressor_predict(fitted_mlp_regressor, X_r) - sink = sink + o_mlp_regressor[0] - array_free_f32(o_mlp_regressor) - } - t3 = flow_now_ns() - mlp_regressor_free(fitted_mlp_regressor) - printf("ESTIMATOR|mlp_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- multi_output_classifier (multioutput_class) ---- - t0 = flow_now_ns() - let probe_multi_output_classifier: MultiOutputClassifier = multi_output_classifier_fit(X_c, Y_labels, 2, 1, 50, 0.01) - t1 = flow_now_ns() - multi_output_classifier_free(probe_multi_output_classifier) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_multi_output_classifier: MultiOutputClassifier = multi_output_classifier_fit(X_c, Y_labels, 2, 1, 50, 0.01) - multi_output_classifier_free(m_multi_output_classifier) - } - t1 = flow_now_ns() - let fitted_multi_output_classifier: MultiOutputClassifier = multi_output_classifier_fit(X_c, Y_labels, 2, 1, 50, 0.01) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_multi_output_classifier: Matrix = multi_output_classifier_predict(fitted_multi_output_classifier, X_c) - if o_multi_output_classifier.rows > 0 { - if o_multi_output_classifier.cols > 0 { sink = sink + o_multi_output_classifier.data[0] } - } - matrix_free(o_multi_output_classifier) - } - t3 = flow_now_ns() - multi_output_classifier_free(fitted_multi_output_classifier) - printf("ESTIMATOR|multi_output_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- multi_output_regressor (multioutput) ---- - t0 = flow_now_ns() - let probe_multi_output_regressor: MultiOutputRegressor = multi_output_regressor_fit(X_r, Y_multi, 2) - t1 = flow_now_ns() - multi_output_regressor_free(probe_multi_output_regressor) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_multi_output_regressor: MultiOutputRegressor = multi_output_regressor_fit(X_r, Y_multi, 2) - multi_output_regressor_free(m_multi_output_regressor) - } - t1 = flow_now_ns() - let fitted_multi_output_regressor: MultiOutputRegressor = multi_output_regressor_fit(X_r, Y_multi, 2) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_multi_output_regressor: Matrix = multi_output_regressor_predict(fitted_multi_output_regressor, X_r) - if o_multi_output_regressor.rows > 0 { - if o_multi_output_regressor.cols > 0 { sink = sink + o_multi_output_regressor.data[0] } - } - matrix_free(o_multi_output_regressor) - } - t3 = flow_now_ns() - multi_output_regressor_free(fitted_multi_output_regressor) - printf("ESTIMATOR|multi_output_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) diff --git a/benchmarks/generated/bench_estimators_05.flow b/benchmarks/generated/bench_estimators_05.flow index 0c54ec5..ff3e11c 100644 --- a/benchmarks/generated/bench_estimators_05.flow +++ b/benchmarks/generated/bench_estimators_05.flow @@ -82,6 +82,118 @@ function main() -> i32 { let mut reps: i32 = 1 let mut sink: f32 = 0.0 + # ---- mlp_classifier (classification) ---- + let hidden_sizes_mlp_classifier: array = [8] + t0 = flow_now_ns() + let probe_mlp_classifier: MLPClassifier = mlp_classifier_fit(X_c, y_c, 3, hidden_sizes_mlp_classifier, 1, 0, 50, 0.01, 0.9, 42) + t1 = flow_now_ns() + mlp_classifier_free(probe_mlp_classifier) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_mlp_classifier: MLPClassifier = mlp_classifier_fit(X_c, y_c, 3, hidden_sizes_mlp_classifier, 1, 0, 50, 0.01, 0.9, 42) + mlp_classifier_free(m_mlp_classifier) + } + t1 = flow_now_ns() + let fitted_mlp_classifier: MLPClassifier = mlp_classifier_fit(X_c, y_c, 3, hidden_sizes_mlp_classifier, 1, 0, 50, 0.01, 0.9, 42) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_mlp_classifier: Matrix = mlp_classifier_predict(fitted_mlp_classifier, X_c) + if o_mlp_classifier.rows > 0 { + if o_mlp_classifier.cols > 0 { sink = sink + o_mlp_classifier.data[0] } + } + matrix_free(o_mlp_classifier) + } + t3 = flow_now_ns() + mlp_classifier_free(fitted_mlp_classifier) + printf("ESTIMATOR|mlp_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- mlp_regressor (regression) ---- + let hidden_sizes_mlp_regressor: array = [8] + t0 = flow_now_ns() + let probe_mlp_regressor: MLPRegressor = mlp_regressor_fit(X_r, y_r, hidden_sizes_mlp_regressor, 1, 0, 50, 0.01, 0.9, 42) + t1 = flow_now_ns() + mlp_regressor_free(probe_mlp_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_mlp_regressor: MLPRegressor = mlp_regressor_fit(X_r, y_r, hidden_sizes_mlp_regressor, 1, 0, 50, 0.01, 0.9, 42) + mlp_regressor_free(m_mlp_regressor) + } + t1 = flow_now_ns() + let fitted_mlp_regressor: MLPRegressor = mlp_regressor_fit(X_r, y_r, hidden_sizes_mlp_regressor, 1, 0, 50, 0.01, 0.9, 42) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_mlp_regressor: ptr = mlp_regressor_predict(fitted_mlp_regressor, X_r) + sink = sink + o_mlp_regressor[0] + array_free_f32(o_mlp_regressor) + } + t3 = flow_now_ns() + mlp_regressor_free(fitted_mlp_regressor) + printf("ESTIMATOR|mlp_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- multi_output_classifier (multioutput_class) ---- + t0 = flow_now_ns() + let probe_multi_output_classifier: MultiOutputClassifier = multi_output_classifier_fit(X_c, Y_labels, 2, 1, 50, 0.01) + t1 = flow_now_ns() + multi_output_classifier_free(probe_multi_output_classifier) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_multi_output_classifier: MultiOutputClassifier = multi_output_classifier_fit(X_c, Y_labels, 2, 1, 50, 0.01) + multi_output_classifier_free(m_multi_output_classifier) + } + t1 = flow_now_ns() + let fitted_multi_output_classifier: MultiOutputClassifier = multi_output_classifier_fit(X_c, Y_labels, 2, 1, 50, 0.01) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_multi_output_classifier: Matrix = multi_output_classifier_predict(fitted_multi_output_classifier, X_c) + if o_multi_output_classifier.rows > 0 { + if o_multi_output_classifier.cols > 0 { sink = sink + o_multi_output_classifier.data[0] } + } + matrix_free(o_multi_output_classifier) + } + t3 = flow_now_ns() + multi_output_classifier_free(fitted_multi_output_classifier) + printf("ESTIMATOR|multi_output_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- multi_output_regressor (multioutput) ---- + t0 = flow_now_ns() + let probe_multi_output_regressor: MultiOutputRegressor = multi_output_regressor_fit(X_r, Y_multi, 2) + t1 = flow_now_ns() + multi_output_regressor_free(probe_multi_output_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_multi_output_regressor: MultiOutputRegressor = multi_output_regressor_fit(X_r, Y_multi, 2) + multi_output_regressor_free(m_multi_output_regressor) + } + t1 = flow_now_ns() + let fitted_multi_output_regressor: MultiOutputRegressor = multi_output_regressor_fit(X_r, Y_multi, 2) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_multi_output_regressor: Matrix = multi_output_regressor_predict(fitted_multi_output_regressor, X_r) + if o_multi_output_regressor.rows > 0 { + if o_multi_output_regressor.cols > 0 { sink = sink + o_multi_output_regressor.data[0] } + } + matrix_free(o_multi_output_regressor) + } + t3 = flow_now_ns() + multi_output_regressor_free(fitted_multi_output_regressor) + printf("ESTIMATOR|multi_output_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- multiclass_logistic (classification) ---- t0 = flow_now_ns() let probe_multiclass_logistic: MultiClassLogisticRegression = multiclass_logistic_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) @@ -110,6 +222,36 @@ function main() -> i32 { printf("ESTIMATOR|multiclass_logistic|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) + # ---- multilabel_binarizer (multioutput_class, written out) ---- + let mlb_counts: ptr = malloc((n_c as i64) * 4) as ptr + for i in 0 to n_c { mlb_counts[i] = 2 } + t0 = flow_now_ns() + let probe_multilabel_binarizer: MultiLabelBinarizer = multilabel_binarizer_fit(Y_label_rows, n_c, mlb_counts, 3) + t1 = flow_now_ns() + multilabel_binarizer_free(probe_multilabel_binarizer) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_multilabel_binarizer: MultiLabelBinarizer = multilabel_binarizer_fit(Y_label_rows, n_c, mlb_counts, 3) + multilabel_binarizer_free(m_multilabel_binarizer) + } + t1 = flow_now_ns() + let fitted_multilabel_binarizer: MultiLabelBinarizer = multilabel_binarizer_fit(Y_label_rows, n_c, mlb_counts, 3) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_multilabel_binarizer: Matrix = multilabel_binarizer_transform(fitted_multilabel_binarizer, Y_label_rows, n_c, mlb_counts) + if o_multilabel_binarizer.rows > 0 { + if o_multilabel_binarizer.cols > 0 { sink = sink + o_multilabel_binarizer.data[0] } + } + matrix_free(o_multilabel_binarizer) + } + t3 = flow_now_ns() + multilabel_binarizer_free(fitted_multilabel_binarizer) + printf("ESTIMATOR|multilabel_binarizer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- multinomial_nb (classification) ---- t0 = flow_now_ns() let probe_multinomial_nb: MultinomialNB = multinomial_nb_fit(X_c, y_c, 3, 1.0) @@ -468,120 +610,6 @@ function main() -> i32 { printf("ESTIMATOR|omp_cv|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- one_class_svm (unsupervised) ---- - t0 = flow_now_ns() - let probe_one_class_svm: OneClassSVM = one_class_svm_fit(X_c, 0.5, 0.1, 100) - t1 = flow_now_ns() - one_class_svm_free(probe_one_class_svm) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_one_class_svm: OneClassSVM = one_class_svm_fit(X_c, 0.5, 0.1, 100) - one_class_svm_free(m_one_class_svm) - } - t1 = flow_now_ns() - printf("ESTIMATOR|one_class_svm|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) - fflush(null) - - # ---- one_vs_one (classification) ---- - t0 = flow_now_ns() - let probe_one_vs_one: OneVsOneClassifier = one_vs_one_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) - t1 = flow_now_ns() - one_vs_one_free(probe_one_vs_one) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_one_vs_one: OneVsOneClassifier = one_vs_one_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) - one_vs_one_free(m_one_vs_one) - } - t1 = flow_now_ns() - let fitted_one_vs_one: OneVsOneClassifier = one_vs_one_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_one_vs_one: ptr = one_vs_one_predict(fitted_one_vs_one, X_c) - sink = sink + o_one_vs_one[0] - array_free_f32(o_one_vs_one) - } - t3 = flow_now_ns() - one_vs_one_free(fitted_one_vs_one) - printf("ESTIMATOR|one_vs_one|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- one_vs_rest (classification) ---- - t0 = flow_now_ns() - let probe_one_vs_rest: OneVsRestClassifier = one_vs_rest_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) - t1 = flow_now_ns() - one_vs_rest_free(probe_one_vs_rest) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_one_vs_rest: OneVsRestClassifier = one_vs_rest_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) - one_vs_rest_free(m_one_vs_rest) - } - t1 = flow_now_ns() - let fitted_one_vs_rest: OneVsRestClassifier = one_vs_rest_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_one_vs_rest: ptr = one_vs_rest_predict(fitted_one_vs_rest, X_c) - sink = sink + o_one_vs_rest[0] - array_free_f32(o_one_vs_rest) - } - t3 = flow_now_ns() - one_vs_rest_free(fitted_one_vs_rest) - printf("ESTIMATOR|one_vs_rest|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- onehot_encoder (unsupervised) ---- - t0 = flow_now_ns() - let probe_onehot_encoder: OneHotEncoder = onehot_encoder_fit(X_c, 0) - t1 = flow_now_ns() - onehot_encoder_free(probe_onehot_encoder) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_onehot_encoder: OneHotEncoder = onehot_encoder_fit(X_c, 0) - onehot_encoder_free(m_onehot_encoder) - } - t1 = flow_now_ns() - let fitted_onehot_encoder: OneHotEncoder = onehot_encoder_fit(X_c, 0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_onehot_encoder: Matrix = onehot_encoder_transform(fitted_onehot_encoder, X_c) - if o_onehot_encoder.rows > 0 { - if o_onehot_encoder.cols > 0 { sink = sink + o_onehot_encoder.data[0] } - } - matrix_free(o_onehot_encoder) - } - t3 = flow_now_ns() - onehot_encoder_free(fitted_onehot_encoder) - printf("ESTIMATOR|onehot_encoder|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- optics (unsupervised) ---- - t0 = flow_now_ns() - let probe_optics: OPTICS = optics_fit(X_c, 0.5, 5) - t1 = flow_now_ns() - optics_free(probe_optics) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_optics: OPTICS = optics_fit(X_c, 0.5, 5) - optics_free(m_optics) - } - t1 = flow_now_ns() - printf("ESTIMATOR|optics|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) diff --git a/benchmarks/generated/bench_estimators_06.flow b/benchmarks/generated/bench_estimators_06.flow index c443010..abf372f 100644 --- a/benchmarks/generated/bench_estimators_06.flow +++ b/benchmarks/generated/bench_estimators_06.flow @@ -82,6 +82,120 @@ function main() -> i32 { let mut reps: i32 = 1 let mut sink: f32 = 0.0 + # ---- one_class_svm (unsupervised) ---- + t0 = flow_now_ns() + let probe_one_class_svm: OneClassSVM = one_class_svm_fit(X_c, 0.5, 0.1, 100) + t1 = flow_now_ns() + one_class_svm_free(probe_one_class_svm) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_one_class_svm: OneClassSVM = one_class_svm_fit(X_c, 0.5, 0.1, 100) + one_class_svm_free(m_one_class_svm) + } + t1 = flow_now_ns() + printf("ESTIMATOR|one_class_svm|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + + # ---- one_vs_one (classification) ---- + t0 = flow_now_ns() + let probe_one_vs_one: OneVsOneClassifier = one_vs_one_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) + t1 = flow_now_ns() + one_vs_one_free(probe_one_vs_one) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_one_vs_one: OneVsOneClassifier = one_vs_one_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) + one_vs_one_free(m_one_vs_one) + } + t1 = flow_now_ns() + let fitted_one_vs_one: OneVsOneClassifier = one_vs_one_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_one_vs_one: ptr = one_vs_one_predict(fitted_one_vs_one, X_c) + sink = sink + o_one_vs_one[0] + array_free_f32(o_one_vs_one) + } + t3 = flow_now_ns() + one_vs_one_free(fitted_one_vs_one) + printf("ESTIMATOR|one_vs_one|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- one_vs_rest (classification) ---- + t0 = flow_now_ns() + let probe_one_vs_rest: OneVsRestClassifier = one_vs_rest_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) + t1 = flow_now_ns() + one_vs_rest_free(probe_one_vs_rest) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_one_vs_rest: OneVsRestClassifier = one_vs_rest_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) + one_vs_rest_free(m_one_vs_rest) + } + t1 = flow_now_ns() + let fitted_one_vs_rest: OneVsRestClassifier = one_vs_rest_fit(X_c, y_c, 3, 50, 0.01, penalty_none()) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_one_vs_rest: ptr = one_vs_rest_predict(fitted_one_vs_rest, X_c) + sink = sink + o_one_vs_rest[0] + array_free_f32(o_one_vs_rest) + } + t3 = flow_now_ns() + one_vs_rest_free(fitted_one_vs_rest) + printf("ESTIMATOR|one_vs_rest|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- onehot_encoder (unsupervised) ---- + t0 = flow_now_ns() + let probe_onehot_encoder: OneHotEncoder = onehot_encoder_fit(X_c, 0) + t1 = flow_now_ns() + onehot_encoder_free(probe_onehot_encoder) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_onehot_encoder: OneHotEncoder = onehot_encoder_fit(X_c, 0) + onehot_encoder_free(m_onehot_encoder) + } + t1 = flow_now_ns() + let fitted_onehot_encoder: OneHotEncoder = onehot_encoder_fit(X_c, 0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_onehot_encoder: Matrix = onehot_encoder_transform(fitted_onehot_encoder, X_c) + if o_onehot_encoder.rows > 0 { + if o_onehot_encoder.cols > 0 { sink = sink + o_onehot_encoder.data[0] } + } + matrix_free(o_onehot_encoder) + } + t3 = flow_now_ns() + onehot_encoder_free(fitted_onehot_encoder) + printf("ESTIMATOR|onehot_encoder|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- optics (unsupervised) ---- + t0 = flow_now_ns() + let probe_optics: OPTICS = optics_fit(X_c, 0.5, 5) + t1 = flow_now_ns() + optics_free(probe_optics) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_optics: OPTICS = optics_fit(X_c, 0.5, 5) + optics_free(m_optics) + } + t1 = flow_now_ns() + printf("ESTIMATOR|optics|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + fflush(null) + # ---- ordinal_encoder (unsupervised) ---- t0 = flow_now_ns() let probe_ordinal_encoder: OrdinalEncoder = ordinal_encoder_fit(X_c, 0) @@ -268,6 +382,35 @@ function main() -> i32 { printf("ESTIMATOR|perceptron|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) + # ---- pipeline (classification, written out) ---- + let pipe_steps: array = [ + step_standard_scaler("scaler"), + step_logistic_regression("classifier", 3, 50, 0.5, penalty_none()) + ] + let pipe_obj: Pipeline = pipeline_new(pipe_steps, 2) + t0 = flow_now_ns() + let probe_pipeline: Pipeline = pipeline_fit(pipe_obj, X_c, y_c) + t1 = flow_now_ns() + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_pipeline: Pipeline = pipeline_fit(pipe_obj, X_c, y_c) + } + t1 = flow_now_ns() + let fitted_pipeline: Pipeline = pipeline_fit(pipe_obj, X_c, y_c) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_pipeline: ptr = pipeline_predict(fitted_pipeline, X_c) + sink = sink + o_pipeline[0] + array_free_f32(o_pipeline) + } + t3 = flow_now_ns() + pipeline_free(fitted_pipeline) + printf("ESTIMATOR|pipeline|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- pls_canonical (multioutput) ---- t0 = flow_now_ns() let probe_pls_canonical: PLSCanonical = pls_canonical_fit(X_r, Y_multi, 2) @@ -462,166 +605,6 @@ function main() -> i32 { printf("ESTIMATOR|power_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- qda (classification) ---- - t0 = flow_now_ns() - let probe_qda: QuadraticDiscriminantAnalysis = qda_fit(X_c, y_c, 3) - t1 = flow_now_ns() - qda_free(probe_qda) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_qda: QuadraticDiscriminantAnalysis = qda_fit(X_c, y_c, 3) - qda_free(m_qda) - } - t1 = flow_now_ns() - let fitted_qda: QuadraticDiscriminantAnalysis = qda_fit(X_c, y_c, 3) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_qda: ptr = qda_predict(fitted_qda, X_c) - sink = sink + o_qda[0] - array_free_f32(o_qda) - } - t3 = flow_now_ns() - qda_free(fitted_qda) - printf("ESTIMATOR|qda|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- quantile_regressor (regression) ---- - t0 = flow_now_ns() - let probe_quantile_regressor: QuantileRegressor = quantile_regressor_fit(X_r, y_r, 0.5, 1.0, 50, 0.01) - t1 = flow_now_ns() - quantile_regressor_free(probe_quantile_regressor) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_quantile_regressor: QuantileRegressor = quantile_regressor_fit(X_r, y_r, 0.5, 1.0, 50, 0.01) - quantile_regressor_free(m_quantile_regressor) - } - t1 = flow_now_ns() - let fitted_quantile_regressor: QuantileRegressor = quantile_regressor_fit(X_r, y_r, 0.5, 1.0, 50, 0.01) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_quantile_regressor: ptr = quantile_regressor_predict(fitted_quantile_regressor, X_r) - sink = sink + o_quantile_regressor[0] - array_free_f32(o_quantile_regressor) - } - t3 = flow_now_ns() - quantile_regressor_free(fitted_quantile_regressor) - printf("ESTIMATOR|quantile_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- quantile_transformer (unsupervised) ---- - t0 = flow_now_ns() - let probe_quantile_transformer: QuantileTransformer = quantile_transformer_fit(X_c, 10, 0) - t1 = flow_now_ns() - quantile_transformer_free(probe_quantile_transformer) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_quantile_transformer: QuantileTransformer = quantile_transformer_fit(X_c, 10, 0) - quantile_transformer_free(m_quantile_transformer) - } - t1 = flow_now_ns() - let fitted_quantile_transformer: QuantileTransformer = quantile_transformer_fit(X_c, 10, 0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_quantile_transformer: Matrix = quantile_transformer_transform(fitted_quantile_transformer, X_c) - if o_quantile_transformer.rows > 0 { - if o_quantile_transformer.cols > 0 { sink = sink + o_quantile_transformer.data[0] } - } - matrix_free(o_quantile_transformer) - } - t3 = flow_now_ns() - quantile_transformer_free(fitted_quantile_transformer) - printf("ESTIMATOR|quantile_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- radius_neighbors_classifier (classification) ---- - t0 = flow_now_ns() - let probe_radius_neighbors_classifier: RadiusNeighborsClassifier = radius_neighbors_classifier_fit(X_c, y_c, 1.0, 3, 0.0) - t1 = flow_now_ns() - radius_neighbors_classifier_free(probe_radius_neighbors_classifier) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_radius_neighbors_classifier: RadiusNeighborsClassifier = radius_neighbors_classifier_fit(X_c, y_c, 1.0, 3, 0.0) - radius_neighbors_classifier_free(m_radius_neighbors_classifier) - } - t1 = flow_now_ns() - let fitted_radius_neighbors_classifier: RadiusNeighborsClassifier = radius_neighbors_classifier_fit(X_c, y_c, 1.0, 3, 0.0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_radius_neighbors_classifier: ptr = radius_neighbors_classifier_predict(fitted_radius_neighbors_classifier, X_c) - sink = sink + o_radius_neighbors_classifier[0] - array_free_f32(o_radius_neighbors_classifier) - } - t3 = flow_now_ns() - radius_neighbors_classifier_free(fitted_radius_neighbors_classifier) - printf("ESTIMATOR|radius_neighbors_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- radius_neighbors_regressor (regression) ---- - t0 = flow_now_ns() - let probe_radius_neighbors_regressor: RadiusNeighborsRegressor = radius_neighbors_regressor_fit(X_r, y_r, 1.0) - t1 = flow_now_ns() - radius_neighbors_regressor_free(probe_radius_neighbors_regressor) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_radius_neighbors_regressor: RadiusNeighborsRegressor = radius_neighbors_regressor_fit(X_r, y_r, 1.0) - radius_neighbors_regressor_free(m_radius_neighbors_regressor) - } - t1 = flow_now_ns() - let fitted_radius_neighbors_regressor: RadiusNeighborsRegressor = radius_neighbors_regressor_fit(X_r, y_r, 1.0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_radius_neighbors_regressor: ptr = radius_neighbors_regressor_predict(fitted_radius_neighbors_regressor, X_r) - sink = sink + o_radius_neighbors_regressor[0] - array_free_f32(o_radius_neighbors_regressor) - } - t3 = flow_now_ns() - radius_neighbors_regressor_free(fitted_radius_neighbors_regressor) - printf("ESTIMATOR|radius_neighbors_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- radius_neighbors_transformer (unsupervised) ---- - t0 = flow_now_ns() - let probe_radius_neighbors_transformer: RadiusNeighborsTransformer = radius_neighbors_transformer_fit(X_c, 1.0, 0) - t1 = flow_now_ns() - radius_neighbors_transformer_free(probe_radius_neighbors_transformer) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_radius_neighbors_transformer: RadiusNeighborsTransformer = radius_neighbors_transformer_fit(X_c, 1.0, 0) - radius_neighbors_transformer_free(m_radius_neighbors_transformer) - } - t1 = flow_now_ns() - let fitted_radius_neighbors_transformer: RadiusNeighborsTransformer = radius_neighbors_transformer_fit(X_c, 1.0, 0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_radius_neighbors_transformer: Matrix = radius_neighbors_transformer_transform(fitted_radius_neighbors_transformer, X_c) - if o_radius_neighbors_transformer.rows > 0 { - if o_radius_neighbors_transformer.cols > 0 { sink = sink + o_radius_neighbors_transformer.data[0] } - } - matrix_free(o_radius_neighbors_transformer) - } - t3 = flow_now_ns() - radius_neighbors_transformer_free(fitted_radius_neighbors_transformer) - printf("ESTIMATOR|radius_neighbors_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) diff --git a/benchmarks/generated/bench_estimators_07.flow b/benchmarks/generated/bench_estimators_07.flow index cbba33a..fc7b530 100644 --- a/benchmarks/generated/bench_estimators_07.flow +++ b/benchmarks/generated/bench_estimators_07.flow @@ -82,6 +82,166 @@ function main() -> i32 { let mut reps: i32 = 1 let mut sink: f32 = 0.0 + # ---- qda (classification) ---- + t0 = flow_now_ns() + let probe_qda: QuadraticDiscriminantAnalysis = qda_fit(X_c, y_c, 3) + t1 = flow_now_ns() + qda_free(probe_qda) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_qda: QuadraticDiscriminantAnalysis = qda_fit(X_c, y_c, 3) + qda_free(m_qda) + } + t1 = flow_now_ns() + let fitted_qda: QuadraticDiscriminantAnalysis = qda_fit(X_c, y_c, 3) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_qda: ptr = qda_predict(fitted_qda, X_c) + sink = sink + o_qda[0] + array_free_f32(o_qda) + } + t3 = flow_now_ns() + qda_free(fitted_qda) + printf("ESTIMATOR|qda|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- quantile_regressor (regression) ---- + t0 = flow_now_ns() + let probe_quantile_regressor: QuantileRegressor = quantile_regressor_fit(X_r, y_r, 0.5, 1.0, 50, 0.01) + t1 = flow_now_ns() + quantile_regressor_free(probe_quantile_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_quantile_regressor: QuantileRegressor = quantile_regressor_fit(X_r, y_r, 0.5, 1.0, 50, 0.01) + quantile_regressor_free(m_quantile_regressor) + } + t1 = flow_now_ns() + let fitted_quantile_regressor: QuantileRegressor = quantile_regressor_fit(X_r, y_r, 0.5, 1.0, 50, 0.01) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_quantile_regressor: ptr = quantile_regressor_predict(fitted_quantile_regressor, X_r) + sink = sink + o_quantile_regressor[0] + array_free_f32(o_quantile_regressor) + } + t3 = flow_now_ns() + quantile_regressor_free(fitted_quantile_regressor) + printf("ESTIMATOR|quantile_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- quantile_transformer (unsupervised) ---- + t0 = flow_now_ns() + let probe_quantile_transformer: QuantileTransformer = quantile_transformer_fit(X_c, 10, 0) + t1 = flow_now_ns() + quantile_transformer_free(probe_quantile_transformer) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_quantile_transformer: QuantileTransformer = quantile_transformer_fit(X_c, 10, 0) + quantile_transformer_free(m_quantile_transformer) + } + t1 = flow_now_ns() + let fitted_quantile_transformer: QuantileTransformer = quantile_transformer_fit(X_c, 10, 0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_quantile_transformer: Matrix = quantile_transformer_transform(fitted_quantile_transformer, X_c) + if o_quantile_transformer.rows > 0 { + if o_quantile_transformer.cols > 0 { sink = sink + o_quantile_transformer.data[0] } + } + matrix_free(o_quantile_transformer) + } + t3 = flow_now_ns() + quantile_transformer_free(fitted_quantile_transformer) + printf("ESTIMATOR|quantile_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- radius_neighbors_classifier (classification) ---- + t0 = flow_now_ns() + let probe_radius_neighbors_classifier: RadiusNeighborsClassifier = radius_neighbors_classifier_fit(X_c, y_c, 1.0, 3, 0.0) + t1 = flow_now_ns() + radius_neighbors_classifier_free(probe_radius_neighbors_classifier) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_radius_neighbors_classifier: RadiusNeighborsClassifier = radius_neighbors_classifier_fit(X_c, y_c, 1.0, 3, 0.0) + radius_neighbors_classifier_free(m_radius_neighbors_classifier) + } + t1 = flow_now_ns() + let fitted_radius_neighbors_classifier: RadiusNeighborsClassifier = radius_neighbors_classifier_fit(X_c, y_c, 1.0, 3, 0.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_radius_neighbors_classifier: ptr = radius_neighbors_classifier_predict(fitted_radius_neighbors_classifier, X_c) + sink = sink + o_radius_neighbors_classifier[0] + array_free_f32(o_radius_neighbors_classifier) + } + t3 = flow_now_ns() + radius_neighbors_classifier_free(fitted_radius_neighbors_classifier) + printf("ESTIMATOR|radius_neighbors_classifier|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- radius_neighbors_regressor (regression) ---- + t0 = flow_now_ns() + let probe_radius_neighbors_regressor: RadiusNeighborsRegressor = radius_neighbors_regressor_fit(X_r, y_r, 1.0) + t1 = flow_now_ns() + radius_neighbors_regressor_free(probe_radius_neighbors_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_radius_neighbors_regressor: RadiusNeighborsRegressor = radius_neighbors_regressor_fit(X_r, y_r, 1.0) + radius_neighbors_regressor_free(m_radius_neighbors_regressor) + } + t1 = flow_now_ns() + let fitted_radius_neighbors_regressor: RadiusNeighborsRegressor = radius_neighbors_regressor_fit(X_r, y_r, 1.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_radius_neighbors_regressor: ptr = radius_neighbors_regressor_predict(fitted_radius_neighbors_regressor, X_r) + sink = sink + o_radius_neighbors_regressor[0] + array_free_f32(o_radius_neighbors_regressor) + } + t3 = flow_now_ns() + radius_neighbors_regressor_free(fitted_radius_neighbors_regressor) + printf("ESTIMATOR|radius_neighbors_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- radius_neighbors_transformer (unsupervised) ---- + t0 = flow_now_ns() + let probe_radius_neighbors_transformer: RadiusNeighborsTransformer = radius_neighbors_transformer_fit(X_c, 1.0, 0) + t1 = flow_now_ns() + radius_neighbors_transformer_free(probe_radius_neighbors_transformer) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_radius_neighbors_transformer: RadiusNeighborsTransformer = radius_neighbors_transformer_fit(X_c, 1.0, 0) + radius_neighbors_transformer_free(m_radius_neighbors_transformer) + } + t1 = flow_now_ns() + let fitted_radius_neighbors_transformer: RadiusNeighborsTransformer = radius_neighbors_transformer_fit(X_c, 1.0, 0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_radius_neighbors_transformer: Matrix = radius_neighbors_transformer_transform(fitted_radius_neighbors_transformer, X_c) + if o_radius_neighbors_transformer.rows > 0 { + if o_radius_neighbors_transformer.cols > 0 { sink = sink + o_radius_neighbors_transformer.data[0] } + } + matrix_free(o_radius_neighbors_transformer) + } + t3 = flow_now_ns() + radius_neighbors_transformer_free(fitted_radius_neighbors_transformer) + printf("ESTIMATOR|radius_neighbors_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- random_forest_classifier (classification) ---- t0 = flow_now_ns() let probe_random_forest_classifier: RandomForestClassifier = random_forest_classifier_fit(X_c, y_c, 3, 10, 5, 42) @@ -438,174 +598,6 @@ function main() -> i32 { printf("ESTIMATOR|select_fdr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- select_fpr (regression) ---- - t0 = flow_now_ns() - let probe_select_fpr: SelectFpr = select_fpr_fit(X_r, y_r, 1.0) - t1 = flow_now_ns() - select_fpr_free(probe_select_fpr) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_select_fpr: SelectFpr = select_fpr_fit(X_r, y_r, 1.0) - select_fpr_free(m_select_fpr) - } - t1 = flow_now_ns() - let fitted_select_fpr: SelectFpr = select_fpr_fit(X_r, y_r, 1.0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_select_fpr: Matrix = select_fpr_transform(fitted_select_fpr, X_r) - if o_select_fpr.rows > 0 { - if o_select_fpr.cols > 0 { sink = sink + o_select_fpr.data[0] } - } - matrix_free(o_select_fpr) - } - t3 = flow_now_ns() - select_fpr_free(fitted_select_fpr) - printf("ESTIMATOR|select_fpr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- select_from_model (classification, written out) ---- - t0 = flow_now_ns() - let probe_select_from_model: SelectFromModel = select_from_model_fit(w_f, f_c, 0.5) - t1 = flow_now_ns() - select_from_model_free(probe_select_from_model) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_select_from_model: SelectFromModel = select_from_model_fit(w_f, f_c, 0.5) - select_from_model_free(m_select_from_model) - } - t1 = flow_now_ns() - let fitted_select_from_model: SelectFromModel = select_from_model_fit(w_f, f_c, 0.5) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_select_from_model: Matrix = select_from_model_transform(fitted_select_from_model, X_c) - if o_select_from_model.rows > 0 { - if o_select_from_model.cols > 0 { sink = sink + o_select_from_model.data[0] } - } - matrix_free(o_select_from_model) - } - t3 = flow_now_ns() - select_from_model_free(fitted_select_from_model) - printf("ESTIMATOR|select_from_model|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- select_fwe (regression) ---- - t0 = flow_now_ns() - let probe_select_fwe: SelectFwe = select_fwe_fit(X_r, y_r, 1.0) - t1 = flow_now_ns() - select_fwe_free(probe_select_fwe) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_select_fwe: SelectFwe = select_fwe_fit(X_r, y_r, 1.0) - select_fwe_free(m_select_fwe) - } - t1 = flow_now_ns() - let fitted_select_fwe: SelectFwe = select_fwe_fit(X_r, y_r, 1.0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_select_fwe: Matrix = select_fwe_transform(fitted_select_fwe, X_r) - if o_select_fwe.rows > 0 { - if o_select_fwe.cols > 0 { sink = sink + o_select_fwe.data[0] } - } - matrix_free(o_select_fwe) - } - t3 = flow_now_ns() - select_fwe_free(fitted_select_fwe) - printf("ESTIMATOR|select_fwe|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- select_k_best (regression) ---- - t0 = flow_now_ns() - let probe_select_k_best: SelectKBest = select_k_best_fit(X_r, y_r, 3, 0) - t1 = flow_now_ns() - select_k_best_free(probe_select_k_best) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_select_k_best: SelectKBest = select_k_best_fit(X_r, y_r, 3, 0) - select_k_best_free(m_select_k_best) - } - t1 = flow_now_ns() - let fitted_select_k_best: SelectKBest = select_k_best_fit(X_r, y_r, 3, 0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_select_k_best: Matrix = select_k_best_transform(fitted_select_k_best, X_r) - if o_select_k_best.rows > 0 { - if o_select_k_best.cols > 0 { sink = sink + o_select_k_best.data[0] } - } - matrix_free(o_select_k_best) - } - t3 = flow_now_ns() - select_k_best_free(fitted_select_k_best) - printf("ESTIMATOR|select_k_best|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- select_percentile (regression) ---- - t0 = flow_now_ns() - let probe_select_percentile: SelectPercentile = select_percentile_fit(X_r, y_r, 50) - t1 = flow_now_ns() - select_percentile_free(probe_select_percentile) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_select_percentile: SelectPercentile = select_percentile_fit(X_r, y_r, 50) - select_percentile_free(m_select_percentile) - } - t1 = flow_now_ns() - let fitted_select_percentile: SelectPercentile = select_percentile_fit(X_r, y_r, 50) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_select_percentile: Matrix = select_percentile_transform(fitted_select_percentile, X_r) - if o_select_percentile.rows > 0 { - if o_select_percentile.cols > 0 { sink = sink + o_select_percentile.data[0] } - } - matrix_free(o_select_percentile) - } - t3 = flow_now_ns() - select_percentile_free(fitted_select_percentile) - printf("ESTIMATOR|select_percentile|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- sequential_feature_selector (regression) ---- - t0 = flow_now_ns() - let probe_sequential_feature_selector: SequentialFeatureSelector = sequential_feature_selector_fit(X_r, y_r, 2, 0, n_r, f_r) - t1 = flow_now_ns() - sequential_feature_selector_free(probe_sequential_feature_selector) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_sequential_feature_selector: SequentialFeatureSelector = sequential_feature_selector_fit(X_r, y_r, 2, 0, n_r, f_r) - sequential_feature_selector_free(m_sequential_feature_selector) - } - t1 = flow_now_ns() - let fitted_sequential_feature_selector: SequentialFeatureSelector = sequential_feature_selector_fit(X_r, y_r, 2, 0, n_r, f_r) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_sequential_feature_selector: Matrix = sequential_feature_selector_transform(fitted_sequential_feature_selector, X_r) - if o_sequential_feature_selector.rows > 0 { - if o_sequential_feature_selector.cols > 0 { sink = sink + o_sequential_feature_selector.data[0] } - } - matrix_free(o_sequential_feature_selector) - } - t3 = flow_now_ns() - sequential_feature_selector_free(fitted_sequential_feature_selector) - printf("ESTIMATOR|sequential_feature_selector|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) diff --git a/benchmarks/generated/bench_estimators_08.flow b/benchmarks/generated/bench_estimators_08.flow index 3311597..27f1bdc 100644 --- a/benchmarks/generated/bench_estimators_08.flow +++ b/benchmarks/generated/bench_estimators_08.flow @@ -82,6 +82,174 @@ function main() -> i32 { let mut reps: i32 = 1 let mut sink: f32 = 0.0 + # ---- select_fpr (regression) ---- + t0 = flow_now_ns() + let probe_select_fpr: SelectFpr = select_fpr_fit(X_r, y_r, 1.0) + t1 = flow_now_ns() + select_fpr_free(probe_select_fpr) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_select_fpr: SelectFpr = select_fpr_fit(X_r, y_r, 1.0) + select_fpr_free(m_select_fpr) + } + t1 = flow_now_ns() + let fitted_select_fpr: SelectFpr = select_fpr_fit(X_r, y_r, 1.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_select_fpr: Matrix = select_fpr_transform(fitted_select_fpr, X_r) + if o_select_fpr.rows > 0 { + if o_select_fpr.cols > 0 { sink = sink + o_select_fpr.data[0] } + } + matrix_free(o_select_fpr) + } + t3 = flow_now_ns() + select_fpr_free(fitted_select_fpr) + printf("ESTIMATOR|select_fpr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- select_from_model (classification, written out) ---- + t0 = flow_now_ns() + let probe_select_from_model: SelectFromModel = select_from_model_fit(w_f, f_c, 0.5) + t1 = flow_now_ns() + select_from_model_free(probe_select_from_model) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_select_from_model: SelectFromModel = select_from_model_fit(w_f, f_c, 0.5) + select_from_model_free(m_select_from_model) + } + t1 = flow_now_ns() + let fitted_select_from_model: SelectFromModel = select_from_model_fit(w_f, f_c, 0.5) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_select_from_model: Matrix = select_from_model_transform(fitted_select_from_model, X_c) + if o_select_from_model.rows > 0 { + if o_select_from_model.cols > 0 { sink = sink + o_select_from_model.data[0] } + } + matrix_free(o_select_from_model) + } + t3 = flow_now_ns() + select_from_model_free(fitted_select_from_model) + printf("ESTIMATOR|select_from_model|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- select_fwe (regression) ---- + t0 = flow_now_ns() + let probe_select_fwe: SelectFwe = select_fwe_fit(X_r, y_r, 1.0) + t1 = flow_now_ns() + select_fwe_free(probe_select_fwe) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_select_fwe: SelectFwe = select_fwe_fit(X_r, y_r, 1.0) + select_fwe_free(m_select_fwe) + } + t1 = flow_now_ns() + let fitted_select_fwe: SelectFwe = select_fwe_fit(X_r, y_r, 1.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_select_fwe: Matrix = select_fwe_transform(fitted_select_fwe, X_r) + if o_select_fwe.rows > 0 { + if o_select_fwe.cols > 0 { sink = sink + o_select_fwe.data[0] } + } + matrix_free(o_select_fwe) + } + t3 = flow_now_ns() + select_fwe_free(fitted_select_fwe) + printf("ESTIMATOR|select_fwe|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- select_k_best (regression) ---- + t0 = flow_now_ns() + let probe_select_k_best: SelectKBest = select_k_best_fit(X_r, y_r, 3, 0) + t1 = flow_now_ns() + select_k_best_free(probe_select_k_best) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_select_k_best: SelectKBest = select_k_best_fit(X_r, y_r, 3, 0) + select_k_best_free(m_select_k_best) + } + t1 = flow_now_ns() + let fitted_select_k_best: SelectKBest = select_k_best_fit(X_r, y_r, 3, 0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_select_k_best: Matrix = select_k_best_transform(fitted_select_k_best, X_r) + if o_select_k_best.rows > 0 { + if o_select_k_best.cols > 0 { sink = sink + o_select_k_best.data[0] } + } + matrix_free(o_select_k_best) + } + t3 = flow_now_ns() + select_k_best_free(fitted_select_k_best) + printf("ESTIMATOR|select_k_best|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- select_percentile (regression) ---- + t0 = flow_now_ns() + let probe_select_percentile: SelectPercentile = select_percentile_fit(X_r, y_r, 50) + t1 = flow_now_ns() + select_percentile_free(probe_select_percentile) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_select_percentile: SelectPercentile = select_percentile_fit(X_r, y_r, 50) + select_percentile_free(m_select_percentile) + } + t1 = flow_now_ns() + let fitted_select_percentile: SelectPercentile = select_percentile_fit(X_r, y_r, 50) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_select_percentile: Matrix = select_percentile_transform(fitted_select_percentile, X_r) + if o_select_percentile.rows > 0 { + if o_select_percentile.cols > 0 { sink = sink + o_select_percentile.data[0] } + } + matrix_free(o_select_percentile) + } + t3 = flow_now_ns() + select_percentile_free(fitted_select_percentile) + printf("ESTIMATOR|select_percentile|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- sequential_feature_selector (regression) ---- + t0 = flow_now_ns() + let probe_sequential_feature_selector: SequentialFeatureSelector = sequential_feature_selector_fit(X_r, y_r, 2, 0, n_r, f_r) + t1 = flow_now_ns() + sequential_feature_selector_free(probe_sequential_feature_selector) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_sequential_feature_selector: SequentialFeatureSelector = sequential_feature_selector_fit(X_r, y_r, 2, 0, n_r, f_r) + sequential_feature_selector_free(m_sequential_feature_selector) + } + t1 = flow_now_ns() + let fitted_sequential_feature_selector: SequentialFeatureSelector = sequential_feature_selector_fit(X_r, y_r, 2, 0, n_r, f_r) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_sequential_feature_selector: Matrix = sequential_feature_selector_transform(fitted_sequential_feature_selector, X_r) + if o_sequential_feature_selector.rows > 0 { + if o_sequential_feature_selector.cols > 0 { sink = sink + o_sequential_feature_selector.data[0] } + } + matrix_free(o_sequential_feature_selector) + } + t3 = flow_now_ns() + sequential_feature_selector_free(fitted_sequential_feature_selector) + printf("ESTIMATOR|sequential_feature_selector|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- sgd_classifier (classification) ---- t0 = flow_now_ns() let probe_sgd_classifier: SGDClassifier = sgd_classifier_fit(X_c, y_c, 3, 0, 1.0, 50, 0.01) @@ -413,166 +581,6 @@ function main() -> i32 { printf("ESTIMATOR|spline_transformer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- stacking_regressor (regression) ---- - t0 = flow_now_ns() - let probe_stacking_regressor: StackingRegressor = stacking_regressor_fit(X_r, y_r, 3, 5, 42) - t1 = flow_now_ns() - stacking_regressor_free(probe_stacking_regressor) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_stacking_regressor: StackingRegressor = stacking_regressor_fit(X_r, y_r, 3, 5, 42) - stacking_regressor_free(m_stacking_regressor) - } - t1 = flow_now_ns() - let fitted_stacking_regressor: StackingRegressor = stacking_regressor_fit(X_r, y_r, 3, 5, 42) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_stacking_regressor: ptr = stacking_regressor_predict(fitted_stacking_regressor, X_r) - sink = sink + o_stacking_regressor[0] - array_free_f32(o_stacking_regressor) - } - t3 = flow_now_ns() - stacking_regressor_free(fitted_stacking_regressor) - printf("ESTIMATOR|stacking_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- standard_scaler (unsupervised) ---- - t0 = flow_now_ns() - let probe_standard_scaler: StandardScaler = standard_scaler_fit(X_c) - t1 = flow_now_ns() - standard_scaler_free(probe_standard_scaler) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_standard_scaler: StandardScaler = standard_scaler_fit(X_c) - standard_scaler_free(m_standard_scaler) - } - t1 = flow_now_ns() - let fitted_standard_scaler: StandardScaler = standard_scaler_fit(X_c) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_standard_scaler: Matrix = standard_scaler_transform(fitted_standard_scaler, X_c) - if o_standard_scaler.rows > 0 { - if o_standard_scaler.cols > 0 { sink = sink + o_standard_scaler.data[0] } - } - matrix_free(o_standard_scaler) - } - t3 = flow_now_ns() - standard_scaler_free(fitted_standard_scaler) - printf("ESTIMATOR|standard_scaler|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- svc (classification) ---- - t0 = flow_now_ns() - let probe_svc: SVC = svc_fit(X_c, y_c, 3, 1.0, 0, 0.1, 2, 0.0) - t1 = flow_now_ns() - svc_free(probe_svc) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_svc: SVC = svc_fit(X_c, y_c, 3, 1.0, 0, 0.1, 2, 0.0) - svc_free(m_svc) - } - t1 = flow_now_ns() - let fitted_svc: SVC = svc_fit(X_c, y_c, 3, 1.0, 0, 0.1, 2, 0.0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_svc: ptr = svc_predict(fitted_svc, X_c) - sink = sink + o_svc[0] - array_free_f32(o_svc) - } - t3 = flow_now_ns() - svc_free(fitted_svc) - printf("ESTIMATOR|svc|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- svr (regression) ---- - t0 = flow_now_ns() - let probe_svr: SVR = svr_fit(X_r, y_r, 1.0, 0.1, 0, 0.1, 2, 0.0) - t1 = flow_now_ns() - svr_free(probe_svr) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_svr: SVR = svr_fit(X_r, y_r, 1.0, 0.1, 0, 0.1, 2, 0.0) - svr_free(m_svr) - } - t1 = flow_now_ns() - let fitted_svr: SVR = svr_fit(X_r, y_r, 1.0, 0.1, 0, 0.1, 2, 0.0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_svr: ptr = svr_predict(fitted_svr, X_r) - sink = sink + o_svr[0] - array_free_f32(o_svr) - } - t3 = flow_now_ns() - svr_free(fitted_svr) - printf("ESTIMATOR|svr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- target_encoder (regression) ---- - t0 = flow_now_ns() - let probe_target_encoder: TargetEncoder = target_encoder_fit(X_r, y_r, n_r, 1.0) - t1 = flow_now_ns() - target_encoder_free(probe_target_encoder) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_target_encoder: TargetEncoder = target_encoder_fit(X_r, y_r, n_r, 1.0) - target_encoder_free(m_target_encoder) - } - t1 = flow_now_ns() - let fitted_target_encoder: TargetEncoder = target_encoder_fit(X_r, y_r, n_r, 1.0) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_target_encoder: Matrix = target_encoder_transform(fitted_target_encoder, X_r) - if o_target_encoder.rows > 0 { - if o_target_encoder.cols > 0 { sink = sink + o_target_encoder.data[0] } - } - matrix_free(o_target_encoder) - } - t3 = flow_now_ns() - target_encoder_free(fitted_target_encoder) - printf("ESTIMATOR|target_encoder|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - - # ---- theil_sen_regressor (regression) ---- - t0 = flow_now_ns() - let probe_theil_sen_regressor: TheilSenRegressor = theil_sen_regressor_fit(X_r, y_r, 10, 100, 42) - t1 = flow_now_ns() - theil_sen_regressor_free(probe_theil_sen_regressor) - reps = 1 - if (t1 - t0) < 200000 { reps = 200 } - elif (t1 - t0) < 2000000 { reps = 20 } - t0 = flow_now_ns() - for rep in 0 to reps { - let m_theil_sen_regressor: TheilSenRegressor = theil_sen_regressor_fit(X_r, y_r, 10, 100, 42) - theil_sen_regressor_free(m_theil_sen_regressor) - } - t1 = flow_now_ns() - let fitted_theil_sen_regressor: TheilSenRegressor = theil_sen_regressor_fit(X_r, y_r, 10, 100, 42) - t2 = flow_now_ns() - for rep2 in 0 to reps { - let o_theil_sen_regressor: ptr = theil_sen_regressor_predict(fitted_theil_sen_regressor, X_r) - sink = sink + o_theil_sen_regressor[0] - array_free_f32(o_theil_sen_regressor) - } - t3 = flow_now_ns() - theil_sen_regressor_free(fitted_theil_sen_regressor) - printf("ESTIMATOR|theil_sen_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) - fflush(null) - for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } free(Y_label_rows as ptr) matrix_free(Y_labels) diff --git a/benchmarks/generated/bench_estimators_09.flow b/benchmarks/generated/bench_estimators_09.flow index e29c98f..f8f386a 100644 --- a/benchmarks/generated/bench_estimators_09.flow +++ b/benchmarks/generated/bench_estimators_09.flow @@ -82,6 +82,212 @@ function main() -> i32 { let mut reps: i32 = 1 let mut sink: f32 = 0.0 + # ---- stacking_regressor (regression) ---- + t0 = flow_now_ns() + let probe_stacking_regressor: StackingRegressor = stacking_regressor_fit(X_r, y_r, 3, 5, 42) + t1 = flow_now_ns() + stacking_regressor_free(probe_stacking_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_stacking_regressor: StackingRegressor = stacking_regressor_fit(X_r, y_r, 3, 5, 42) + stacking_regressor_free(m_stacking_regressor) + } + t1 = flow_now_ns() + let fitted_stacking_regressor: StackingRegressor = stacking_regressor_fit(X_r, y_r, 3, 5, 42) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_stacking_regressor: ptr = stacking_regressor_predict(fitted_stacking_regressor, X_r) + sink = sink + o_stacking_regressor[0] + array_free_f32(o_stacking_regressor) + } + t3 = flow_now_ns() + stacking_regressor_free(fitted_stacking_regressor) + printf("ESTIMATOR|stacking_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- standard_scaler (unsupervised) ---- + t0 = flow_now_ns() + let probe_standard_scaler: StandardScaler = standard_scaler_fit(X_c) + t1 = flow_now_ns() + standard_scaler_free(probe_standard_scaler) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_standard_scaler: StandardScaler = standard_scaler_fit(X_c) + standard_scaler_free(m_standard_scaler) + } + t1 = flow_now_ns() + let fitted_standard_scaler: StandardScaler = standard_scaler_fit(X_c) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_standard_scaler: Matrix = standard_scaler_transform(fitted_standard_scaler, X_c) + if o_standard_scaler.rows > 0 { + if o_standard_scaler.cols > 0 { sink = sink + o_standard_scaler.data[0] } + } + matrix_free(o_standard_scaler) + } + t3 = flow_now_ns() + standard_scaler_free(fitted_standard_scaler) + printf("ESTIMATOR|standard_scaler|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- svc (classification) ---- + t0 = flow_now_ns() + let probe_svc: SVC = svc_fit(X_c, y_c, 3, 1.0, 0, 0.1, 2, 0.0) + t1 = flow_now_ns() + svc_free(probe_svc) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_svc: SVC = svc_fit(X_c, y_c, 3, 1.0, 0, 0.1, 2, 0.0) + svc_free(m_svc) + } + t1 = flow_now_ns() + let fitted_svc: SVC = svc_fit(X_c, y_c, 3, 1.0, 0, 0.1, 2, 0.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_svc: ptr = svc_predict(fitted_svc, X_c) + sink = sink + o_svc[0] + array_free_f32(o_svc) + } + t3 = flow_now_ns() + svc_free(fitted_svc) + printf("ESTIMATOR|svc|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- svr (regression) ---- + t0 = flow_now_ns() + let probe_svr: SVR = svr_fit(X_r, y_r, 1.0, 0.1, 0, 0.1, 2, 0.0) + t1 = flow_now_ns() + svr_free(probe_svr) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_svr: SVR = svr_fit(X_r, y_r, 1.0, 0.1, 0, 0.1, 2, 0.0) + svr_free(m_svr) + } + t1 = flow_now_ns() + let fitted_svr: SVR = svr_fit(X_r, y_r, 1.0, 0.1, 0, 0.1, 2, 0.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_svr: ptr = svr_predict(fitted_svr, X_r) + sink = sink + o_svr[0] + array_free_f32(o_svr) + } + t3 = flow_now_ns() + svr_free(fitted_svr) + printf("ESTIMATOR|svr|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- target_encoder (regression) ---- + t0 = flow_now_ns() + let probe_target_encoder: TargetEncoder = target_encoder_fit(X_r, y_r, n_r, 1.0) + t1 = flow_now_ns() + target_encoder_free(probe_target_encoder) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_target_encoder: TargetEncoder = target_encoder_fit(X_r, y_r, n_r, 1.0) + target_encoder_free(m_target_encoder) + } + t1 = flow_now_ns() + let fitted_target_encoder: TargetEncoder = target_encoder_fit(X_r, y_r, n_r, 1.0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_target_encoder: Matrix = target_encoder_transform(fitted_target_encoder, X_r) + if o_target_encoder.rows > 0 { + if o_target_encoder.cols > 0 { sink = sink + o_target_encoder.data[0] } + } + matrix_free(o_target_encoder) + } + t3 = flow_now_ns() + target_encoder_free(fitted_target_encoder) + printf("ESTIMATOR|target_encoder|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- tfidf_vectorizer (classification, written out) ---- + let tfidf_vectorizer_docs: array = [ + "the quick brown fox jumps over the lazy dog", + "a lazy dog sleeps in the warm sun", + "quick brown foxes are rare in the city", + "the dog and the fox share a field", + "warm sun and a cold river run together", + "a field of brown grass in the sun", + "the city river runs past the old field", + "old dogs sleep through a quick storm", + "a storm over the city wakes the dog", + "foxes hunt in the cold river valley", + "the valley holds a warm field of grass", + "grass grows where the river meets the sun", + "a rare fox crosses the old stone bridge", + "the stone bridge over the cold river", + "dogs and foxes keep their distance here", + "here the field the river and the city meet" + ] + t0 = flow_now_ns() + let probe_tfidf_vectorizer: TfidfVectorizer = tfidf_vectorizer_fit(tfidf_vectorizer_docs, 16, 50) + t1 = flow_now_ns() + tfidf_vectorizer_free(probe_tfidf_vectorizer) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_tfidf_vectorizer: TfidfVectorizer = tfidf_vectorizer_fit(tfidf_vectorizer_docs, 16, 50) + tfidf_vectorizer_free(m_tfidf_vectorizer) + } + t1 = flow_now_ns() + let fitted_tfidf_vectorizer: TfidfVectorizer = tfidf_vectorizer_fit(tfidf_vectorizer_docs, 16, 50) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_tfidf_vectorizer: Matrix = tfidf_vectorizer_transform(fitted_tfidf_vectorizer, tfidf_vectorizer_docs, 16) + if o_tfidf_vectorizer.rows > 0 { + if o_tfidf_vectorizer.cols > 0 { sink = sink + o_tfidf_vectorizer.data[0] } + } + matrix_free(o_tfidf_vectorizer) + } + t3 = flow_now_ns() + tfidf_vectorizer_free(fitted_tfidf_vectorizer) + printf("ESTIMATOR|tfidf_vectorizer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + + # ---- theil_sen_regressor (regression) ---- + t0 = flow_now_ns() + let probe_theil_sen_regressor: TheilSenRegressor = theil_sen_regressor_fit(X_r, y_r, 10, 100, 42) + t1 = flow_now_ns() + theil_sen_regressor_free(probe_theil_sen_regressor) + reps = 1 + if (t1 - t0) < 200000 { reps = 200 } + elif (t1 - t0) < 2000000 { reps = 20 } + t0 = flow_now_ns() + for rep in 0 to reps { + let m_theil_sen_regressor: TheilSenRegressor = theil_sen_regressor_fit(X_r, y_r, 10, 100, 42) + theil_sen_regressor_free(m_theil_sen_regressor) + } + t1 = flow_now_ns() + let fitted_theil_sen_regressor: TheilSenRegressor = theil_sen_regressor_fit(X_r, y_r, 10, 100, 42) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_theil_sen_regressor: ptr = theil_sen_regressor_predict(fitted_theil_sen_regressor, X_r) + sink = sink + o_theil_sen_regressor[0] + array_free_f32(o_theil_sen_regressor) + } + t3 = flow_now_ns() + theil_sen_regressor_free(fitted_theil_sen_regressor) + printf("ESTIMATOR|theil_sen_regressor|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) + fflush(null) + # ---- transformed_target_regressor (regression) ---- t0 = flow_now_ns() let probe_transformed_target_regressor: TransformedTargetRegressor = transformed_target_regressor_fit(X_r, y_r, 0) diff --git a/lib/scikit/feature_extraction.flow b/lib/scikit/feature_extraction.flow index 45ac2cd..f368448 100644 --- a/lib/scikit/feature_extraction.flow +++ b/lib/scikit/feature_extraction.flow @@ -41,15 +41,20 @@ export struct CountVectorizer { fitted: bool } +# djb2, carried in 64 bits and reduced every step. +# +# The i32 version overflowed into a negative h on the sixth character of a +# token and then shifted it, which the runtime traps as a left shift of a +# negative. A four document fixture never reached that; a sixteen document one +# did, and the process aborted inside the fit. function _hash_string(s: string) -> i32 { - let h: i32 = 5381 + let mut h: i64 = 5381 let len: i32 = _string_length(s) for i in 0 to len { let c: i32 = _string_char_at(s, i) - h = ((h << 5) + h) + c + h = (h * 33 + (c as i64)) % 2147483647 } - if h < 0 { h = 0 - h } - return h + return h as i32 } function _tokenize(doc: string, out_tokens: ptr, max_tokens: i32) -> i32 { From 6e12da3945b78dfd626e43d732a1c97fb62ba1d6 Mon Sep 17 00:00:00 2001 From: godofecht Date: Sun, 27 Sep 2026 22:36:49 +0100 Subject: [PATCH 08/12] Fix the four rows CI found, and rank the two the clock could not see The first CI run of the wide matrix put 175 of 179 ranked rows in front of scikit-learn and named the four that were behind. Every one of them is a different story from the macOS numbers, which is the point of gating on the runner rather than on a developer machine. multitask_lasso_cv, 109.9 ms against 59.7. It now runs the same block coordinate descent as the elastic net, warm started down the alpha path. The sweep loop also gained a real convergence test: the coefficient move it used was weak at a small penalty, where the iterate creeps and no single coefficient moves far in one sweep, so a fit spent its whole budget on an objective that had stopped falling. The objective comes out of the residual the sweep already maintains, for about half a percent of a sweep. 109.9 ms to 3.6 ms on this machine, and the elastic net path gets the same test. affinity_propagation, 15.5 ms against 4.9. The three n by n matrices are one allocation each rather than n row allocations, the similarity is built off row pointers and mirrored, and the label pass is folded into the availability pass that already has both terms in hand. tweedie_regressor, 1.31 ms against 0.98, with the Poisson and gamma fits on the same loop. All three walked the design twice per row through matrix_at and allocated a gradient buffer inside the iteration. One sgemv for the linear predictor, one transposed sgemv for the gradient, buffers allocated once. linear_svr, 0.66 ms against 0.42. The dot product that its epoch is made of accumulated into one register, so the loop was a chain of dependent additions. Four accumulators. kernel_density and nearest_neighbors sat unranked because their work function is called score_samples and kneighbors, so the generic path found nothing to time and the fit alone fell under the clock's floor. A recipe can now name the work function and the scikit-learn method to match it against. That puts every row the matrix can rank in the ranking. --- benchmarks/bench_estimators_sklearn.py | 6 +- benchmarks/estimator_coverage.json | 42 +- benchmarks/estimator_coverage.py | 22 + benchmarks/generate_estimator_bench.py | 16 +- benchmarks/generated/bench_estimators_02.flow | 17 +- benchmarks/generated/bench_estimators_05.flow | 12 +- docs/benchmarks.html | 2 +- lib/scikit/cluster.flow | 107 ++-- lib/scikit/linear.flow | 578 +++++++++--------- lib/scikit/svm.flow | 23 +- 10 files changed, 459 insertions(+), 366 deletions(-) diff --git a/benchmarks/bench_estimators_sklearn.py b/benchmarks/bench_estimators_sklearn.py index ed09341..a516f87 100644 --- a/benchmarks/bench_estimators_sklearn.py +++ b/benchmarks/bench_estimators_sklearn.py @@ -168,7 +168,11 @@ def main() -> int: args.repeats, ) pred_ms = 0.0 - for method in ("predict", "transform"): + # A recipe can name the method to time, for a class whose work + # is called something other than predict or transform. + methods = [shape["sklearn_work"]] if shape and shape.get("sklearn_work") \ + else ["predict", "transform"] + for method in methods: if hasattr(model, method): try: pred_ms = timed(lambda m=method: getattr(model, m)(first), args.repeats) diff --git a/benchmarks/estimator_coverage.json b/benchmarks/estimator_coverage.json index 4fa9de1..17e8d94 100644 --- a/benchmarks/estimator_coverage.json +++ b/benchmarks/estimator_coverage.json @@ -2,8 +2,8 @@ "schema_version": 1, "counts": { "estimators": 203, - "runnable": 168, - "shaped": 20, + "runnable": 166, + "shaped": 22, "simplified": 4, "flow_only": 6, "different_shape": 5 @@ -3922,8 +3922,23 @@ "bandwidth", "kernel" ], - "bucket": "runnable", - "sklearn_estimator": "KernelDensity" + "bucket": "shaped", + "sklearn_estimator": "KernelDensity", + "shape": { + "dataset": "classification", + "flow_fit": [ + "X_c", + "0.5", + "0" + ], + "flow_work": [ + "X_c" + ], + "flow_work_fn": "kernel_density_score_samples", + "flow_work_returns": "ptr", + "sklearn_input": "X", + "sklearn_work": "score_samples" + } }, { "flow_estimator": "kernel_pca", @@ -7472,8 +7487,23 @@ "X", "n_neighbors" ], - "bucket": "runnable", - "sklearn_estimator": "NearestNeighbors" + "bucket": "shaped", + "sklearn_estimator": "NearestNeighbors", + "shape": { + "dataset": "classification", + "flow_fit": [ + "X_c", + "5" + ], + "flow_work": [ + "X_c" + ], + "flow_work_fn": "nearest_neighbors_kneighbors", + "flow_work_returns": "ptr", + "flow_work_release": "nearest_neighbors_free_results({var}, n_c)", + "sklearn_input": "X", + "sklearn_work": "kneighbors" + } }, { "flow_estimator": "nmf", diff --git a/benchmarks/estimator_coverage.py b/benchmarks/estimator_coverage.py index 5058c97..bc3499f 100644 --- a/benchmarks/estimator_coverage.py +++ b/benchmarks/estimator_coverage.py @@ -311,6 +311,28 @@ "flow_work": ["x1d_r", "n_r"], "sklearn_input": "x1d", }, + # Two rows whose work function is named for what it returns rather than + # predict or transform, so the generic path found nothing to time and the + # fit alone fell under the clock's floor. + "kernel_density": { + "dataset": "classification", + "flow_fit": ["X_c", "0.5", "0"], + "flow_work": ["X_c"], + "flow_work_fn": "kernel_density_score_samples", + "flow_work_returns": "ptr", + "sklearn_input": "X", + "sklearn_work": "score_samples", + }, + "nearest_neighbors": { + "dataset": "classification", + "flow_fit": ["X_c", "5"], + "flow_work": ["X_c"], + "flow_work_fn": "nearest_neighbors_kneighbors", + "flow_work_returns": "ptr", + "flow_work_release": "nearest_neighbors_free_results({var}, n_c)", + "sklearn_input": "X", + "sklearn_work": "kneighbors", + }, "multilabel_binarizer": { # The label rows, so the scikit-learn side gets sets of labels rather # than one label per sample. diff --git a/benchmarks/generate_estimator_bench.py b/benchmarks/generate_estimator_bench.py index 2a582cc..9f2a808 100644 --- a/benchmarks/generate_estimator_bench.py +++ b/benchmarks/generate_estimator_bench.py @@ -232,7 +232,13 @@ def shaped_block(entry: dict) -> str: free_ok = free is not None and len(free["parameters"]) == 1 comp = entry["companions"].get("transform") or entry["companions"].get("predict") work = shape.get("flow_work") - comp_ok = comp is not None and work is not None and comp["returns"] in ("Matrix", "ptr") + # A recipe can name the work function itself, for an estimator whose + # companion is called something other than predict or transform. + if shape.get("flow_work_fn"): + comp = {"name": shape["flow_work_fn"], "returns": shape["flow_work_returns"]} + comp_ok = comp is not None and work is not None and ( + comp["returns"] in ("Matrix", "ptr") or shape.get("flow_work_release") + ) lines = [f" # ---- {name} ({kind}, written out) ----"] if shape.get("corpus"): @@ -267,8 +273,12 @@ def shaped_block(entry: dict) -> str: lines.append(" for rep2 in 0 to reps {") lines.append(f" let o_{name}: {comp['returns']} = " f"{comp['name']}(fitted_{name}, {', '.join(work)})") - lines += sink_line(f"o_{name}", comp["returns"]) - lines.append(f" {release}(o_{name})") + if comp["returns"] in ("Matrix", "ptr"): + lines += sink_line(f"o_{name}", comp["returns"]) + if shape.get("flow_work_release"): + lines.append(" " + shape["flow_work_release"].format(var=f"o_{name}")) + else: + lines.append(f" {release}(o_{name})") lines.append(" }") lines.append(" t3 = flow_now_ns()") if free_ok: diff --git a/benchmarks/generated/bench_estimators_02.flow b/benchmarks/generated/bench_estimators_02.flow index 2bc53bb..b5be195 100644 --- a/benchmarks/generated/bench_estimators_02.flow +++ b/benchmarks/generated/bench_estimators_02.flow @@ -542,9 +542,9 @@ function main() -> i32 { printf("ESTIMATOR|kbins_discretizer|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- kernel_density (unsupervised) ---- + # ---- kernel_density (classification, written out) ---- t0 = flow_now_ns() - let probe_kernel_density: KernelDensity = kernel_density_fit(X_c, 1.0, 0) + let probe_kernel_density: KernelDensity = kernel_density_fit(X_c, 0.5, 0) t1 = flow_now_ns() kernel_density_free(probe_kernel_density) reps = 1 @@ -552,11 +552,20 @@ function main() -> i32 { elif (t1 - t0) < 2000000 { reps = 20 } t0 = flow_now_ns() for rep in 0 to reps { - let m_kernel_density: KernelDensity = kernel_density_fit(X_c, 1.0, 0) + let m_kernel_density: KernelDensity = kernel_density_fit(X_c, 0.5, 0) kernel_density_free(m_kernel_density) } t1 = flow_now_ns() - printf("ESTIMATOR|kernel_density|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + let fitted_kernel_density: KernelDensity = kernel_density_fit(X_c, 0.5, 0) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_kernel_density: ptr = kernel_density_score_samples(fitted_kernel_density, X_c) + sink = sink + o_kernel_density[0] + array_free_f32(o_kernel_density) + } + t3 = flow_now_ns() + kernel_density_free(fitted_kernel_density) + printf("ESTIMATOR|kernel_density|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) for i in 0 to n_c { array_free_f32(Y_label_rows[i]) } diff --git a/benchmarks/generated/bench_estimators_05.flow b/benchmarks/generated/bench_estimators_05.flow index ff3e11c..85601b1 100644 --- a/benchmarks/generated/bench_estimators_05.flow +++ b/benchmarks/generated/bench_estimators_05.flow @@ -444,7 +444,7 @@ function main() -> i32 { printf("ESTIMATOR|nearest_centroid|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) - # ---- nearest_neighbors (unsupervised) ---- + # ---- nearest_neighbors (classification, written out) ---- t0 = flow_now_ns() let probe_nearest_neighbors: NearestNeighbors = nearest_neighbors_fit(X_c, 5) t1 = flow_now_ns() @@ -458,7 +458,15 @@ function main() -> i32 { nearest_neighbors_free(m_nearest_neighbors) } t1 = flow_now_ns() - printf("ESTIMATOR|nearest_neighbors|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), 0.0, reps) + let fitted_nearest_neighbors: NearestNeighbors = nearest_neighbors_fit(X_c, 5) + t2 = flow_now_ns() + for rep2 in 0 to reps { + let o_nearest_neighbors: ptr = nearest_neighbors_kneighbors(fitted_nearest_neighbors, X_c) + nearest_neighbors_free_results(o_nearest_neighbors, n_c) + } + t3 = flow_now_ns() + nearest_neighbors_free(fitted_nearest_neighbors) + printf("ESTIMATOR|nearest_neighbors|%.9f|%.9f|%d|ok\n", ms_between(t0, t1) / (reps as f32), ms_between(t2, t3) / (reps as f32), reps) fflush(null) # ---- nmf (unsupervised) ---- diff --git a/docs/benchmarks.html b/docs/benchmarks.html index f7754de..cf2a688 100644 --- a/docs/benchmarks.html +++ b/docs/benchmarks.html @@ -1,4 +1,4 @@ -Benchmarks, flow-scikit

canonical v2 / parity + disparity benchmark

Eligibility never means identity.

All 19 canonical rows are measured and currently eligible for comparison, but numerical, semantic and runtime disparities remain first-class evidence. This page renders the committed benchmark and disparity artifacts directly so differences cannot disappear merely because a row passes its contract.

Flow wins...

End-to-end fit + predict comparisons won by Flow.

sklearn wins...

End-to-end comparisons won by scikit-learn.

parity eligible...

Rows admitted to the competitive denominator.

substantive disparities...

Rows whose fitted state, score, configuration or semantics genuinely diverge, above float-noise floors. Runtime differences are tracked per row but not counted here.

TIMING_UNIT|msend-to-endseed=4280/20 persisted split2% practical tie thresholddisparity retained after eligibility
KMeans note: Digits KMeans is eligible under the same declared contract as every other clustering row. Its seeded k-means++ initialization now matches scikit-learn's, so the strict diagnostic and the final eligibility decision agree. The convergence statistic, the point at which inertia is reported, empty-cluster relocation and the n_init selection rule still differ and stay visible in the disparity artifact.

runtime overview

The plots are generated from the canonical JSON.

Each runtime plot shows end-to-end fit + predict time on a log scale. The plots use the same rows as the table below and therefore update whenever the frozen canonical result changes.

All 19 speed ratios

scikit-learn total time divided by Flow total time. The vertical 1× line separates Flow wins from scikit-learn wins.

Iris total runtime

scikit-learnFlow

Digits total runtime

scikit-learnFlow

Diabetes total runtime

scikit-learnFlow

persistent disparity

Passing parity does not erase the gap.

The disparity plot normalizes each row's principal numerical difference against its effective tolerance where a tolerance is available. A value near 1 means the row is close to the acceptance boundary. Semantic/configuration differences are tracked in the same artifact and remain visible in the table.

Numerical disparity relative to tolerance

The dashed line is the acceptance boundary. Values can remain non-zero even for eligible rows.

all canonical rows

No selected-win table.

Every row is shown below. Speedup is sklearn_ms / flow_ms; values above 1× favor Flow. Strict diagnostic status is kept separate from final eligibility.

AlgorithmDatasetFinal parityStrict diagnosticWinnerscore |Δ|sklearn msFlow msspeedup

larger data

The canonical rows all fit in 1797 samples.

A separate matrix runs five estimators at 100, 1000 and 10000 rows against 8 and 32 features. One run of it does not settle a row: at 100 and 1000 samples a fit finishes in well under a millisecond, and the CI runner moves that by more than the difference being measured. Lasso at 1000 rows and 32 features was recorded at 3.83x and at 0.92x on code that differs in nothing touching Lasso. The table is therefore the spread across consecutive runs rather than one run's number, sorted with the narrowest margins first.

Algorithmsamplesfeaturesruns wonmedianrange

This matrix is reported without gating the build. What does gate is benchmarks/scaled_flow_baseline.json, a Flow-against-itself comparison refreshed from a CI artifact.

the rest of the library

The library is 203 estimators.

The canonical rows above race twelve estimators under a parity contract. lib/scikit exports 203, and a statement about Flow against scikit-learn covers six percent of it while the rest go unmeasured. A separate registry maps every exported fit to its scikit-learn counterpart and times both sides on the same data.

raced...

Ranked against a named scikit-learn class.

Flow faster...

Of the rows that produce a ratio.

own harness needed...

Takes a pipeline, a vectorizer input or a list of fitted models first.

no counterpart...

Flow implements it and scikit-learn has no equivalent.

Read this for what it is. These rows carry no parity contract, no declared tolerances and no disparity report. Each library runs its own defaults over the same data, which answers whether an implementation is in the same performance league and says nothing about whether it computes the same thing. The canonical rows above are where numerical equivalence is established. These timings also come from a developer machine rather than from CI, and the machine was not idle.

The widest ratios say more about the defaults than about the code. Each side runs its own. A row can differ by three orders of magnitude simply because one library does far more work at its defaults. That is a difference in the job, and the ratio does not measure how fast either one is at the same job. Iris also carries no missing values, which leaves the imputers nothing to impute on the Flow side while scikit-learn still runs its full round robin. The narrow rows at the top of the table are the informative ones.

Flow estimatorscikit-learnFlow msscikit-learn msspeedup

methodology

Correctness, disparity and timing are separate dimensions.

The benchmark consumes the same persisted train/test indices in Python and Flow. Python uses high-resolution adaptive timing and the canonical runner aggregates repeated process measurements with medians and IQR. Flow timings are emitted in milliseconds and aggregated by the same runner.

Supervised rows compare predictive metrics under declared tolerances. PCA additionally checks explained variance, singular values, reconstruction error and sign-aligned components. KMeans uses permutation-invariant clustering quality and inertia. The persistent disparity artifact preserves raw numerical gaps and known semantic/configuration differences even after the estimator-specific eligibility contract succeeds.

historical deployment evidence

Footprint and startup remain separate experiments.

The repository also contains a historical deployment comparison recording a roughly 1.4 MB Flow native executable and a roughly 65× cold-start advantage (33 ms versus 2160 ms). Those figures come from a different deployment experiment and are intentionally not mixed into the canonical estimator timing denominator.

trajectory

Flow versus Python, across freezes.

Each row's speedup at the previous freeze and at the latest one. A speedup can move because Flow changed or because scikit-learn's side changed on that runner. When a row moves by more than 10%, the last column names which side's own time moved more, from the committed absolute timings.

reproduce

Read the source artifacts.

Canonical result ↗ Disparity report ↗