Play it: stevenli-phoenix-work.itch.io/laya-webgpu (desktop Chrome / Edge)
A support ticket, an email or a chat message goes in, together with a few typed questions: which team? how urgent? is the customer about to leave? Calibrated probabilities come out in your browser. Nothing is generated and nothing is sent to a server.
The model is Laya, an open-source "System One" decision model. Every
choice / score / noul (yes-no) question is answered in one forward pass. Here it runs on WebGPU through
laya-ts and ONNX Runtime Web.
| model (pick it on the page) | download | languages | WebGPU p50* | WASM p50* |
|---|---|---|---|---|
| typed-decisions · q4 (default) | 467 MB | English | 361 ms | ~2.8 s |
| multilingual · fp32 | 1.3 GB | 100+ | 148 ms | 808 ms |
*Apple M4 Pro, Chrome, 3 questions over a 163-token ticket, 20 warm runs. The first load downloads the weights from Hugging Face (q4, multilingual). After that they come from the browser cache.
python3 -m http.server 8765 # then open http://127.0.0.1:8765/The page picks a hosted model by default. To use your own export instead:
./export_model.sh typed-decisions --q4 # clone laya @6d942c9, export fp32 (verified vs torch), quantize → ./model
./export_model.sh multilingual # fp32, 100+ languagesOn localhost a third "本地 ./model/" option appears. Any CORS-enabled export also works with
?model=https://…/resolve/main/.
tools/quantize_q4.py runs ONNX Runtime's MatMulNBitsQuantizer (4 bit, block 32, symmetric) over every weight
MatMul of the encoder (112) and the head (8). The token embedding stays fp32. Size drops from 1.69 GB to 467 MB.
Fidelity, measured with tools/eval_onnx.mjs on 40 labelled tickets × 3 questions (pnpm install && pnpm eval):
| typed-decisions | size | agrees with fp32 | max prob drift | dept / urgency / churn accuracy | WebGPU p50 |
|---|---|---|---|---|---|
| fp32 | 1.69 GB | — | — | 0.875 / 0.425 / 0.700 | 265 ms |
| q4 enc + head | 467 MB | 110 / 120 | 0.22 | 0.875 / 0.450 / 0.625 | 361 ms |
| q4 asym / block 16 | 539–588 MB | 110–114 / 120 | 0.14–0.25 | similar | — |
| q8 | 705 MB | 119 / 120 | 0.013 | 0.900 / 0.425 / 0.700 | ≈ WASM |
- q8 is not usable in the browser. ONNX Runtime Web has no 8-bit MatMulNBits kernel for WebGPU, so those nodes quietly run on the CPU (2.8 s) while the page still reports WebGPU. Use q8 only for CPU / Node.
- q4 changes about 8 % of decisions relative to fp32. Recalibrate thresholds on your own held-out data.
- q4 runs slower than fp32 on the GPU (dequantization), but the download is 3.6× smaller. On a first visit that matters more.
uv run verify_webgpu.py # load ./model, 20 warm runs, JSON report + screenshot
uv run verify_webgpu.py --force-wasm # hide navigator.gpu: the WASM baseline. A big gap proves the GPU path ran
uv run verify_iframe.py # itch-style cross-site iframe; bare iframe; model host without CORS (must fail)
MODEL_URL=https://huggingface.co/Steven10429/laya-typed-decisions-webgpu-q4/resolve/main/ uv run verify_iframe.py --scenario itchThings that matter when embedding (all measured):
- WebGPU works inside a cross-site iframe with no extra
allow. - The weight host must send CORS headers.
- itch caps uploads at 200 MB per file, so the weights live on Hugging Face.
- itch embeds with
scrolling="no", so the page fits one viewport and each panel scrolls on its own.
index.html, app.js the page (no build step; onnxruntime-web from jsDelivr via import map)
vendor/laya-ts/ laya-ts dist (ES modules) built from NandhaKishorM/laya @ 6d942c9, Apache-2.0
export_model.sh checkpoint → split ONNX (→ q4)
tools/quantize_q4.py 4-bit MatMulNBits quantizer for laya-ts exports
tools/eval_onnx.mjs accuracy / agreement / drift of one or more exports on a labelled JSONL (Node, CPU)
verify_*.py headless-Chrome checks (PEP 723, run with uv)
The model is by ConvAI Innovations / Nandakishor M (NandhaKishorM/laya),
Apache-2.0. vendor/laya-ts is theirs. This repository is also Apache-2.0; see NOTICE.
The same page is packaged as part of a Claude Code skill in
StevenLi-phoenix/laya-skill.
