Skip to content

Repository files navigation

NumberRecognitionAI

A neural network I wrote by hand in NumPy — no PyTorch, no scikit-learn — trained on MNIST and running in your browser. Draw a digit, watch what the network sees, and when it gets one wrong, tell it what the digit was: it takes a gradient step there and then.

Try it · no sign-up, no backend, nothing you draw leaves your browser.

97.87% on the 10,000 MNIST test images · 784 → 64 → 10 · 50,890 parameters · 418 KB of weights · zero runtime dependencies.

Drawing a digit and seeing it recognised

Why I built it

I wanted to know whether I could write backpropagation, not call it. Two lines of PyTorch give you a better model than this one and teach you nothing about why it works.

So everything that makes this a neural network is in net/: the forward pass, the backward pass, He initialisation, the softmax and its gradient, SGD with momentum. NumPy contributes the matrix multiply and nothing else. The proof that the derivation is right is not that the model trains — a subtly wrong gradient still trains, just worse. It is tests/test_gradients.py, which nudges every weight by 1e-6 and checks the analytic gradient predicted the change in the loss.

The part that actually matters

Most browser MNIST demos are disappointing, and they all fail the same way: they send the canvas straight to the model. MNIST digits were not scanned and left alone. Each one was cropped to its ink, scaled so its longest side was 20 pixels, and placed in a 28×28 field centred on its centre of mass — not on the middle of its bounding box.

Feed a model trained on that a raw drawing and it reads it at around 60%. Do the same three steps first and it reads it at the accuracy it was measured at. That is why the page shows the preparation instead of hiding it: crop box on your stroke, the 20-pixel version, and the centred result with a crosshair on the centre of mass.

The other detail worth having: the downscale is an area average, not nearest-neighbour sampling. MNIST strokes have soft grey edges because they were averaged down, and a model trained on soft edges does not like hard black-and-white ones.

Teaching it something wrong

Correcting the network

Correcting a digit runs the same backward pass used in training, on one example, with a learning rate of 0.002. Six corrections are enough to turn a 7 it was 99.6% sure about into a 9. The corrected weights are kept in localStorage, so the network you meet on your second visit is the one you taught, not the one everybody else gets.

Keep going and it starts forgetting the other digits. That is not a bug in the demo, it is what online SGD on a single example with no replay does, and a test asserts it happens. The restore button is there for that reason.

What the layers are doing

Results

Test accuracy 97.87% — 213 errors out of 10,000
Validation accuracy 98.12% (best of 38 epochs, early stopped)
Training time about 90 seconds on a laptop CPU
Inference in the browser 0.5 ms
Average confidence when wrong 72%

Its favourite mistakes are 9 read as 4 (15 times), 5 read as 3 (14) and 7 read as 2 (12). Full confusion matrix on the page.

What is wrong with it

  • It is confident when it is wrong. A miss still averages 72% confidence. The model is not calibrated and I have not calibrated it.
  • A mouse is not a pen. MNIST is handwriting on paper. Mouse strokes are thicker and shakier, so what you experience will be worse than 97.87%. Shift and rotation augmentation during training closes part of that gap, not all of it.
  • It is a multilayer perceptron, not a convolutional network. It has no idea that pixels next to each other are related. Move a digit a few pixels without recentring and it loses track of it, which is exactly why the preparation step is mandatory rather than decorative.
  • 64 hidden units is small. A wider layer or a convolution would get past 99%, at the cost of a slower correction step in the browser and a lot more code.

Running it

python -m venv .venv && .venv/Scripts/activate     # source .venv/bin/activate on macOS/Linux
pip install -r requirements.txt
npm --prefix web ci

npm run data        # downloads MNIST (11 MB) and verifies its checksums
npm run train       # about 90 seconds; writes web/public/weights.json
npm run evaluate    # metrics and the confusion matrix

npm run dev         # the demo at http://localhost:5173

The trained weights are committed, so npm run dev works without training anything.

npm run lint        # ruff + eslint
npm run typecheck   # tsc
npm test            # 26 pytest + 56 vitest
npm run test:e2e    # 7 Playwright tests against the production build

How it is laid out

net/                 Python: the network and the training
  data.py            MNIST download, checksum check, IDX parser
  layers.py          dense, ReLU, softmax + cross entropy
  network.py         forward, backward, SGD with momentum
  augment.py         random shifts and rotations
  train.py           the loop, early stopping, learning-rate decay
  evaluate.py        accuracy, confusion matrix
  export.py          weights.json and the golden fixture
web/                 TypeScript: the demo
  src/net.ts         the same network, in the browser
  src/preprocess.ts  crop, scale to 20 px, centre by mass
tests/               pytest
web/tests/           vitest and Playwright

The maths exists twice, in Python and in TypeScript, which is the one structural risk here. Nothing stops the two from drifting apart except the golden fixture: Python writes down its activations and gradients for twenty fixed images, and a Vitest test requires TypeScript to reproduce them. Both are float32, so the comparison is relative — two implementations that add 784 products in a different order cannot agree more closely than the arithmetic allows.

Credits

MNIST is by Yann LeCun, Corinna Cortes and Christopher Burges. The dataset is downloaded from the CVDF mirror; the original site refuses automated requests now. Twenty of its test images are committed as a test fixture.

The typeface is IBM Plex, under the SIL Open Font License, kept in web/public/fonts/ rather than loaded from Google so the page makes no third-party requests at all.

About

A neural network written by hand in NumPy, trained on MNIST, that learns from your corrections in the browser

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages