Skip to content

saloon: Householder QR and least squares; maroon: linear and ridge regression (Hoon) - #91

Open
sigilante wants to merge 1 commit into
mainfrom
sigilante/qr-regression
Open

sigilante wants to merge 1 commit into
mainfrom
sigilante/qr-regression

Conversation

@sigilante

Copy link
Copy Markdown
Collaborator

Step 4 of the foundation work list in #86: QR and the first regression models. Hoon only, no jet hints, off main (the #87/#88/#90 stack is merged, so nothing is stacked).

Saloon — %linalg

  • +qr — the thin Householder factorization A = Q*R for rows >= cols. Q has orthonormal columns; R is upper triangular with an exactly-zero subdiagonal rather than rounding residue. A zero column needs no reflection and is skipped.
  • +trsv-r — back substitution against R, read as it is (the companion of +qr, as +trsv-lo/+trsv-up are of +chol).
  • +lstsq — min |A*x - v| as R^-1 * Q^T*v. Solving through QR avoids the squared condition number of the normal equations. Needs full column rank.
  • +gram (x^T*x) and +matvec-t (m^T*x), neither materializing a transpose.
  • +house, +reflect, +sum-sq — the pieces +qr is built from.

Sign convention. +house uses LAPACK's choice, alpha = -sign(x0)*|x|, so v0 = x0 - alpha never cancels. A side effect is that the factors match numpy.linalg.qr including signs — negative diagonal entries in R and all — which the oracle verifies rather than assumes.

Maroon — linear models

  • +linreg — ordinary least squares through Saloon's +lstsq.
  • +ridge — solves (Xc^T*Xc + alpha*I) coef = Xc^T*yc with +chol-solve. Any alpha > 0 makes that positive definite even when x is rank-deficient, which is half the point of ridge; alpha = 0 reduces to OLS.
  • Both centre the data and recover the intercept afterwards (intercept = mean(y) - mean(x).coef), as scikit-learn does with fit_intercept=True — which also keeps the intercept out of the ridge penalty. Both return [coef intercept].
  • +predict, +mse, +r2 — r2 crashes on a constant target, where it is 0/0.

No transpose jet anywhere

The lagoon transpose jet crashes on every runtime older than urbit/vere#1057, which is most deployed ships today. So +lstsq uses +matvec-t and +ridge uses +gram instead of mmul of a transpose — and this PR also moves Maroon's existing +cov onto +gram, the one place #90 used transpose. +gram is exactly symmetric, since entries (i,j) and (j,i) multiply the same pairs.

Tests: 36 arms

saloon-qr.hoon (19) and maroon-regression.hoon (17).

Exact, compared as whole rays — inputs whose column norms are powers of two, so every Householder vector, scale and reflection is representable:

  • diag(2,2) → Q = -I, R = -2I (the LAPACK signs); the same on a tall 3×2
  • R exactly -5 for [3 4]^T; a zero column skipped, not divided by
  • lstsq = [2 3] on a tall system whose third row is unreachable, and on a square one
  • y = 3x + 1 on x = [0 0 2 2]: the centred feature [-1 -1 1 1] has norm exactly 2, giving coef 3, intercept 1, mse 0, r2 1
  • ridge with alpha = 12 makes the system [[16]], so coef 0.75, intercept 3.25, mse 5.0625 and r2 0.4375 are all exact; alpha = 0 equals OLS

Approximate, within a tolerance: R and Q against numpy.linalg.qr (signs included), Q*R ≈ A, Q^T*Q ≈ I (via +gram), lstsq against numpy.linalg.lstsq and against the normal equations, and two-feature fits against sklearn's LinearRegression and Ridge(alpha=1.0), plus the fit's R².

Oracles: saloon/tools/linalg_check.py (56 checks, now with a Householder mirror in exact rationals) and maroon/tools/ml_check.py (35 checks, with sklearn).

Verification

Fresh fakeship at hoon-135/zuse-408 on the vere #1057 build (with the @rq fix, since merged upstream as urbit/vere#1112):

python3 saloon/tools/linalg_check.py        # 56 checks, 0 failures
python3 maroon/tools/ml_check.py            # 35 checks, 0 failures
-test /=num=/tests/lib/saloon-qr            # 19, ok=%.y
-test /=num=/tests/lib/maroon-regression    # 17, ok=%.y
-test /=num=/tests/lib                      # 438 OK / 4 failed  (was 402 / 4)

+36 OK and the failures unchanged; maroon-pca (19) and saloon-linalg (44) stay green after the +cov change. The 4 are the pre-existing saloon-unum posit transcendentals, which fail on every runtime.

Not in this PR

LU with pivoting (not needed yet — Cholesky and QR cover the solves the models use), jets, and the next #86 step: gradient-descent optimizers, then logistic regression.

Refs #86.

🤖 Generated with Claude Code

https://claude.ai/code/session_01TgBnKsUPYzkoPZonjePnaq

…gression

Step 4 of #86 (Cholesky landed in #88).  Hoon only, no jet
hints, off main.

Saloon (%linalg):

  +qr        thin Householder QR, A = Q*R for rows >= cols.  Q has orthonormal
             columns; R is upper triangular with an EXACTLY-zero subdiagonal.
             The reflector uses LAPACK's convention, alpha = -sign(x0)*|x|, so
             v0 never cancels -- and the factors match numpy.linalg.qr's signs,
             negative diagonal entries included (the oracle checks this).  A
             zero column needs no reflection and is skipped.
  +trsv-r    back substitution against R, read as it is.
  +lstsq     min |A*x - v| as R^-1 * Q^T*v, avoiding the squared condition
             number of the normal equations.
  +gram      x^T*x, and +matvec-t, m^T*x, neither materializing a transpose.
  +house, +reflect, +sum-sq   the helpers +qr is built from.

Maroon:

  +linreg    OLS through +lstsq.
  +ridge     (Xc^T*Xc + alpha*I) coef = Xc^T*yc through +chol-solve; any
             alpha > 0 makes that positive definite even for rank-deficient x.
  Both centre the data and recover the intercept afterwards, as scikit-learn
  does with fit_intercept=True, which keeps the intercept out of the penalty.
  +predict, +mse, +r2   the model's predictions and the two metrics.

Nothing here routes through the lagoon transpose jet, which crashes on
runtimes older than urbit/vere#1057: +lstsq uses +matvec-t and +ridge uses
+gram.  The same change moves Maroon's existing +cov onto +gram.

Tests: 36 arms (saloon-qr 19, maroon-regression 17).  The exact cases pick
inputs whose column norms are powers of two, so every Householder vector,
scale and reflection is representable -- Q = -I and R = -2I for diag(2,2);
lstsq [2 3] on a tall system with an unreachable row; and y = 3x+1 on
x = [0 0 2 2], whose centred feature has norm exactly 2, giving coef 3,
intercept 1, and with alpha = 12 a ridge fit of 0.75/3.25 whose MSE (5.0625)
and R^2 (0.4375) are exact too.  The rest compare against numpy.linalg.qr/lstsq
and sklearn's LinearRegression/Ridge within a tolerance.  The two oracles
re-derive all of it: saloon/tools/linalg_check.py (56 checks, with a Householder
mirror in exact rationals) and maroon/tools/ml_check.py (35 checks).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TgBnKsUPYzkoPZonjePnaq
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant