Skip to content

Repository files navigation

Laya on WebGPU: typed decisions running in the browser

Laya on WebGPU

Play it: stevenli-phoenix-work.itch.io/laya-webgpu (desktop Chrome / Edge)

A support ticket, an email or a chat message goes in, together with a few typed questions: which team? how urgent? is the customer about to leave? Calibrated probabilities come out in your browser. Nothing is generated and nothing is sent to a server.

The model is Laya, an open-source "System One" decision model. Every choice / score / noul (yes-no) question is answered in one forward pass. Here it runs on WebGPU through laya-ts and ONNX Runtime Web.

model (pick it on the page) download languages WebGPU p50* WASM p50*
typed-decisions · q4 (default) 467 MB English 361 ms ~2.8 s
multilingual · fp32 1.3 GB 100+ 148 ms 808 ms

*Apple M4 Pro, Chrome, 3 questions over a 163-token ticket, 20 warm runs. The first load downloads the weights from Hugging Face (q4, multilingual). After that they come from the browser cache.

Run it locally

python3 -m http.server 8765        # then open http://127.0.0.1:8765/

The page picks a hosted model by default. To use your own export instead:

./export_model.sh typed-decisions --q4   # clone laya @6d942c9, export fp32 (verified vs torch), quantize → ./model
./export_model.sh multilingual           # fp32, 100+ languages

On localhost a third "本地 ./model/" option appears. Any CORS-enabled export also works with ?model=https://…/resolve/main/.

How the 4-bit model was made, and what it costs

tools/quantize_q4.py runs ONNX Runtime's MatMulNBitsQuantizer (4 bit, block 32, symmetric) over every weight MatMul of the encoder (112) and the head (8). The token embedding stays fp32. Size drops from 1.69 GB to 467 MB.

Fidelity, measured with tools/eval_onnx.mjs on 40 labelled tickets × 3 questions (pnpm install && pnpm eval):

typed-decisions size agrees with fp32 max prob drift dept / urgency / churn accuracy WebGPU p50
fp32 1.69 GB — — 0.875 / 0.425 / 0.700 265 ms
q4 enc + head 467 MB 110 / 120 0.22 0.875 / 0.450 / 0.625 361 ms
q4 asym / block 16 539–588 MB 110–114 / 120 0.14–0.25 similar —
q8 705 MB 119 / 120 0.013 0.900 / 0.425 / 0.700 ≈ WASM
  • q8 is not usable in the browser. ONNX Runtime Web has no 8-bit MatMulNBits kernel for WebGPU, so those nodes quietly run on the CPU (2.8 s) while the page still reports WebGPU. Use q8 only for CPU / Node.
  • q4 changes about 8 % of decisions relative to fp32. Recalibrate thresholds on your own held-out data.
  • q4 runs slower than fp32 on the GPU (dequantization), but the download is 3.6× smaller. On a first visit that matters more.

Verify it (headless Chrome, own temp profile)

uv run verify_webgpu.py                 # load ./model, 20 warm runs, JSON report + screenshot
uv run verify_webgpu.py --force-wasm    # hide navigator.gpu: the WASM baseline. A big gap proves the GPU path ran
uv run verify_iframe.py                 # itch-style cross-site iframe; bare iframe; model host without CORS (must fail)
MODEL_URL=https://huggingface.co/Steven10429/laya-typed-decisions-webgpu-q4/resolve/main/ uv run verify_iframe.py --scenario itch

Things that matter when embedding (all measured):

  • WebGPU works inside a cross-site iframe with no extra allow.
  • The weight host must send CORS headers.
  • itch caps uploads at 200 MB per file, so the weights live on Hugging Face.
  • itch embeds with scrolling="no", so the page fits one viewport and each panel scrolls on its own.

Layout

index.html, app.js      the page (no build step; onnxruntime-web from jsDelivr via import map)
vendor/laya-ts/         laya-ts dist (ES modules) built from NandhaKishorM/laya @ 6d942c9, Apache-2.0
export_model.sh         checkpoint → split ONNX (→ q4)
tools/quantize_q4.py    4-bit MatMulNBits quantizer for laya-ts exports
tools/eval_onnx.mjs     accuracy / agreement / drift of one or more exports on a labelled JSONL (Node, CPU)
verify_*.py             headless-Chrome checks (PEP 723, run with uv)

Credits and license

The model is by ConvAI Innovations / Nandakishor M (NandhaKishorM/laya), Apache-2.0. vendor/laya-ts is theirs. This repository is also Apache-2.0; see NOTICE. The same page is packaged as part of a Claude Code skill in StevenLi-phoenix/laya-skill.

About

Laya typed decisions (choice / score / yes-no) running in the browser on WebGPU — itch.io demo

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages