Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
2a47358
Serve Laya on Apple Silicon: worker, benchmarks and recipe
cacheline999 Sep 28, 2026
326838a
Add measured Laya reports and a paired frontend-overhead probe
cacheline999 Sep 28, 2026
a2b8b29
Warm every loaded Laya model and harden the HTTP benchmark
cacheline999 Sep 28, 2026
d44d3fe
Move the Laya benchmark scripts under recipe/laya/bench
cacheline999 Sep 28, 2026
d3d0817
Add LAYA_WORKER_WEIGHTS=fp16 for the compiled worker
cacheline999 Sep 30, 2026
e10cf06
Compile the encoder for multi-question batches; make compile on/off
cacheline999 Sep 30, 2026
b00b285
Recipe: M5 results, test dependencies, first-request note
cacheline999 Sep 30, 2026
3dc662b
Say that fp16 weights are for MPS and warn on CPU
cacheline999 Sep 30, 2026
6311ae8
Read /health from the live model; correct the Apple Silicon recipe
cacheline999 Sep 30, 2026
9ad04e3
Run plain fp32 laya after a fallback to the CPU
cacheline999 Sep 30, 2026
5e260a1
Split the GPU optimizations out of worker.py; correct the Laya README
cacheline999 Sep 30, 2026
01db6e7
Merge remote-tracking branch 'origin/main' into laya-apple-silicon
cacheline999 Sep 30, 2026
005a874
Lay the Laya worker out like the Cua-S1 worker (#19)
cacheline999 Oct 1, 2026
3d6cb57
Merge remote-tracking branch 'origin/main' into laya-apple-silicon
cacheline999 Oct 1, 2026
9e8ddb1
Laya worker: keep measurements out of comments, report compile.active…
cacheline999 Oct 1, 2026
3a78507
Bench scripts: mark the imports that follow the sys.path setup (E402)
cacheline999 Oct 1, 2026
584680d
Laya worker: prepare checkpoints loaded while serving; fix report par…
cacheline999 Oct 2, 2026
7dc6b68
Merge remote-tracking branch 'origin/main' into laya-apple-silicon
cacheline999 Oct 2, 2026
87c920f
Laya worker: restore checkpoints evicted for a failed late load; say …
cacheline999 Oct 2, 2026
1450884
Laya recipe: say what a late checkpoint load costs through the fronte…
cacheline999 Oct 2, 2026
543acef
Laya worker: drop comments that repeat the code or the test name
cacheline999 Oct 2, 2026
f2b7c6a
Laya: third review fixes, requirement-derived tests on CPU and GPU, a…
cacheline999 Oct 2, 2026
668f331
Laya: keep profile runs out of the parity section; header test off ma…
cacheline999 Oct 2, 2026
d1bef5d
Laya bench and tests: two-sided answer check in paired summaries; rea…
cacheline999 Oct 2, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ LAYA can run as an external Python worker for text requests; its in-repository m

| Model | Status |
| --- | --- |
| LAYA | [External worker](recipe/laya/README.md); [CPU checkpoint reader](src/models/laya/README.md); model execution planned |
| LAYA | [External worker](recipe/laya/README.md); [Python worker on Apple Silicon (MPS) and CPU](recipe/laya/apple-silicon.md); [CPU checkpoint reader](src/models/laya/README.md); model execution planned |
| Cua-S1 4B 0.2 (`text` adapter) | [Python worker](recipe/cua_s1/text.md); [native worker](recipe/cua_s1/native.md), CUDA, run on sm_89 |

CUDA and Metal coverage will be documented per model as implementations are added and validated.
Expand Down
2 changes: 2 additions & 0 deletions recipe/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@

- [Laya text worker](laya/README.md): start the external Python worker, connect the
Rust frontend and compare direct and proxied responses.
- [Laya on Apple Silicon](laya/apple-silicon.md): serve Laya on the Mac GPU with the Laya
worker, put the frontend in front of it and run the benchmarks.
- [Cua-S1 4B 0.2 text worker](cua_s1/text.md): download the pinned weights, start
the worker and connect the Rust frontend.
- [Cua-S1 4B 0.2 native text worker](cua_s1/native.md): build the CUDA library and
Expand Down
3 changes: 2 additions & 1 deletion recipe/laya/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,8 @@
This recipe runs the external Laya Python package behind the Rust frontend.
It validates text decisions; image, audio and video inference are not covered.

Run all commands from the repository root.
Run all commands from the repository root. To serve on the GPU of an Apple Silicon Mac, see
[Laya on Apple Silicon](apple-silicon.md).

## Start the worker

Expand Down
192 changes: 192 additions & 0 deletions recipe/laya/apple-silicon.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,192 @@
# Laya on Apple Silicon

This recipe serves Laya on the GPU of an Apple Silicon Mac (PyTorch MPS) with the worker in
[`src/frontend/laya_mps.py`](../../src/frontend/laya_mps.py), puts the Rust frontend in front of it and
runs the benchmarks. The model-side code is in [`src/models/laya/`](../../src/models/laya/). The
[Laya text worker](README.md) recipe covers plain laya-serve on the CPU.

Validated on an M1 Pro (16 GB, 16-core GPU), macOS 26.1, Python 3.12, `laya[serve]==0.3.20`,
torch 2.14.0 and the `english` checkpoint (`convaiinnovations/laya` at `55cf4c4`), and by another
contributor on an M5 (10-core GPU, 32 GB, macOS 26.5.2). Other M-series Macs have not been tested.

Run all commands from the repository root.

## Install

Use Python 3.12. If `python3.12` is not on your `PATH` and you have uv, replace the first command below
with `uv venv --python 3.12 --seed .venv` (`--seed` puts pip in the environment).

```sh
python3.12 -m venv .venv
.venv/bin/python -m pip install -r recipe/laya/requirements-mps.txt
.venv/bin/python -c "import torch; print(torch.backends.mps.is_available())"
```

The last command must print `True`. The standard macOS arm64 wheel of torch includes MPS.

## Start the worker

```sh
PYTHONPATH=src .venv/bin/python -m frontend.laya_mps --device mps --model english --require-device --port 8000
```

First startup downloads the checkpoint (846 MB; 97 s into an empty cache at 8.7 MB/s when measured). The
worker loads the model, runs a warmup
over short, long and multi-question requests, and only then listens on port 8000, so the first
request it accepts is already warm: on an M1 Pro the first request after ready took 70–81 ms, against
0.7–1.1 s from plain laya-serve. `--require-device` makes it exit instead of silently serving on the CPU
when the model cannot be placed on MPS; without it the worker logs a warning and serves from the CPU.
Of laya-serve's environment variables, `LAYA_API_KEY` (bearer authentication) still applies. Those its
launcher reads do not: device, model, host, port and log level are the flags above, and `LAYA_THREADS` and
`LAYA_AUTO_TASK` are not read. The worker warns at startup if any of them is set.

Laya loads another checkpoint when a request names it (`"model": "multilingual"`) or its routing picks it
(a non-English state). The worker prepares that checkpoint the same way inside that first request; other
requests wait behind it, `/health` names it under `preparing` meanwhile and lists it afterwards. On the
M1 Pro that first request took about 5–10 s without the options and about 70 s with `--compile` (plus the
download the first time, 680 MB for `multilingual`). The frontend gives a backend 60 s, so with `--compile`
it answered that request with 504 while the worker finished preparing; the same request sent again then
took 35 ms. To avoid that, send one request for each further checkpoint straight to the worker after
startup. Each resident checkpoint needs its own memory (see Troubleshooting).

With `--require-device`, a checkpoint that does not land on the requested device is unloaded again and
the request fails with 500. If Laya evicted another checkpoint to make room for it (it keeps two by
default), the worker loads that one again.

Check what it is running on:

```sh
curl -s http://127.0.0.1:8000/health
```

`device` must be `mps` and `device_mismatch` `false`. The response also names the checkpoint and
revision, the weight dtype (`torch.float32`; Laya upcasts the fp16 checkpoint on MPS), the autocast
dtype Laya uses for requests with at least `mps_amp_min_rows` questions, and the warmup time. The device
and dtypes are read on every call: if a request runs out of GPU memory, Laya moves the model to the CPU
and keeps serving, and `/health` then shows `device: cpu` and `device_mismatch: true`. Triggered on the
M1 Pro by lowering PyTorch's MPS memory limit: the request that ran out of memory still returned 200
after about 30 s, and later 68-token requests took 140–270 ms from the CPU.

### Faster: compile and fp16 weights

```sh
PYTHONPATH=src .venv/bin/python -m frontend.laya_mps --device mps --model english --require-device \
--compile --weights fp16 --port 8000
```

`--compile` compiles the model during warmup: one-question requests run the whole model compiled,
requests with several questions run the encoder compiled and Laya's decision head as it is.
`--weights fp16` keeps the checkpoint's fp16 weights instead of Laya's fp32 upcast on MPS.

On the M1 Pro, with both workers running and every request sent to each back to back, the two options
together lowered warm p50 against the worker without them by 37–38% for a 68-token one-question
request (about 57 → 35 ms in those runs), 17–20% at 198–484 tokens, 14% for three questions and 18% for
six. Answers stayed within 0.0031 of the fp32 worker's. A worker running on its own uses about 3 GB
with the options instead of 4.2 GB (2.8 GB against 3.5 GB in those paired runs, where the two workers
shared the machine), measured on the six benchmark inputs; see below for how it grows. The price is startup: the worker became ready after 35–39 s instead of 8–10 s, and
its first request after that took 62–78 ms.

On an M5 the same paired comparison gave median ratios of 0.51–0.53 for one-question requests at
47–68 tokens, 0.30–0.33 at 198–484 tokens, 0.37 for three questions and 0.60 for six, most of it from
the fp16 weights, which on that GPU speed up every input even without compile. There the worker was
ready after 19 s instead of 3 s; its first request took 21–36 ms in 21 of 23 fresh starts and 327 and
409 ms in the other two, not yet explained (132–143 ms from plain laya-serve).

### What the warm numbers leave out

The latencies above are for requests sent back to back. Measured on the M1 Pro:

- **Idle gaps.** A request that follows a pause is slower, with or without the options, because the GPU
has slowed down in the meantime. For a short one-question request (25 ms back to back with the options,
40 ms without) it took about 50 ms after 0.2–1 s of idle and 105–115 ms after 2–5 s (60–68 ms and
114–127 ms without the options). This is also why the first request after ready costs more than a warm
one. An agent that asks once every few seconds sees these numbers, not the back-to-back ones. A
heartbeat of one forward pass every 0.5 s held it at about 45 ms in a probe, for 8% GPU load; the worker
does not do this.
- **New input lengths.** The first request of a length the worker has not seen costs about 15 ms more
once with the options (6 ms without). It is not a recompile (`recompiled_after_ready` stays `false`).
- **Memory grows with the lengths seen.** With `--compile`, PyTorch keeps host memory for every input
length the compiled model has run, about 5 MB each (fp16 weights alone add little): the 3 GB above became
3.3 GB after 100 new lengths and 5.3 GB after all 477, more than the 4.0 GB of a worker without the
options. `torch.mps.empty_cache()` releases most of it, and those lengths then pay their first-request
cost again.

`/health` reports under `compile` how many graphs existed when the worker became ready and how many
exist now; `recompiled_after_ready: true` means a request shape was not covered by the warmup.
`active` is `false` once no model runs the compiled path any more, i.e. after a fallback to the CPU.

Both options apply on the GPU only. After a fallback to the CPU the worker runs Laya's fp32 model
uncompiled, like a worker started without them.

## Start the frontend

In another terminal:

```sh
cargo build --release --locked
OMNI_JEV_BIND=127.0.0.1:8080 \
OMNI_JEV_BACKEND_URL=http://127.0.0.1:8000 \
./target/release/omni-jev
```

## Send a request

```sh
curl http://127.0.0.1:8080/v1/systemone \
-H 'Content-Type: application/json' \
-d '{"model":"english","state":"Please refund the duplicate charge.","questions":{"refund":{"type":"noul","instructions":"Does the customer ask for a refund?"}}}'
```

The frontend forwards the worker's response unchanged; `compare_with_backend.py` from the
[Laya text worker](README.md#compare-responses) recipe checks that against this setup as well.

## Test

The tests need `pytest` and `httpx2` (Starlette's `TestClient`; `httpx` works with a deprecation warning):

```sh
.venv/bin/python -m pip install pytest httpx2
PYTHONPATH=src .venv/bin/python -m pytest tests/laya # unit tests, no model
LAYA_CONTRACT=1 PYTHONPATH=src .venv/bin/python -m pytest tests/laya # plus contract tests against a CPU worker
```

The contract tests start a real worker and check readiness, the three decision types, error responses, and
that its answers match Laya run directly in fp32 on the CPU. On an Apple Silicon Mac, run them against the
GPU as well, without and with the options:

```sh
LAYA_CONTRACT=1 LAYA_CONTRACT_DEVICE=mps PYTHONPATH=src .venv/bin/python -m pytest tests/laya/test_contract.py
LAYA_CONTRACT=1 LAYA_CONTRACT_DEVICE=mps LAYA_CONTRACT_FLAGS="--compile --weights fp16" \
PYTHONPATH=src .venv/bin/python -m pytest tests/laya/test_contract.py
```

## Benchmark

Stop the worker and frontend first; the benchmark starts its own. The scripts are listed in
[`bench/`](bench/README.md). A first pass that checks everything runs:

```sh
.venv/bin/python recipe/laya/bench/bench_inproc.py --device mps --config C2 --run feasibility
.venv/bin/python recipe/laya/bench/bench_http.py --config C3 --run feasibility --spawn .venv/bin/laya-serve
.venv/bin/python recipe/laya/bench/bench_http.py --config C4 --run feasibility \
--url http://127.0.0.1:8080 --frontend target/release/omni-jev --spawn .venv/bin/laya-serve
.venv/bin/python recipe/laya/bench/paired.py --run feasibility --a "" --b "--compile --weights fp16"
.venv/bin/python recipe/laya/bench/report.py recipe/laya/bench/results/*_feasibility.jsonl --ref C2
.venv/bin/python recipe/laya/bench/paired.py --summarize recipe/laya/bench/results/paired_feasibility.jsonl
```

Runs labelled anything other than `feasibility` refuse to start on battery power or when the
1-minute load average is above 2, so close other heavy applications and plug the Mac in first.

## Troubleshooting

- `device_mismatch: true` at startup, or with `--require-device` the worker exits with
`asked for mps, english is on cpu`: MPS is not available to this Python. Check the `torch.backends.mps.is_available()` line above (an x86_64
Python under Rosetta, for example, has no MPS).
- `device_mismatch: true` on a worker that started on MPS: Laya fell back to the CPU after a GPU
out-of-memory error. It keeps answering, several times slower; free memory and restart the worker
to get back on the GPU.
- The worker process uses about 4 GB, or 3 GB with fp16 weights (Activity Monitor's Memory column, which
counts MPS allocations), with one checkpoint loaded; with `--compile` it grows towards 5 GB as it sees
more input lengths (see "What the warm numbers leave out"); a second one Laya loads later adds its own. On a 16 GB Mac, close other large applications before benchmarking.
- `Address already in use`: another worker or frontend still holds port 8000 or 8080.
60 changes: 60 additions & 0 deletions recipe/laya/bench/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
# Laya benchmark scripts

Scripts behind the numbers in the [Apple Silicon recipe](../apple-silicon.md). Each run writes raw
JSONL to `results/` (kept out of the repository); `report.py` builds the tables from it.

| file | purpose |
| --- | --- |
| `workloads.jsonl` | the fixed inputs: W1–W6 are timed, P* are for answer comparison only. Tokens per row with Laya's tokenizer: W1 68, W2 198, W3 484, W4 68/48/47, W5 40–68, W6 47 |
| `bench_inproc.py` | Laya in-process (no HTTP): load, warmup, first request, warm latency, memory |
| `bench_http.py` | a `/v1/systemone` server, optionally started by the script and optionally behind the frontend: time to ready, first request, warm latency, throughput |
| `paired.py` | two configurations compared request by request, both alive at once: two worker flag sets, or two running servers (e.g. a worker directly and through the frontend) |
| `profile_mps.py` | where a request's time goes on MPS |
| `report.py` | tables from the JSONL, including the run-to-run gate and the answer comparison against a reference config |
| `env.py` | shared: versions, checkpoint, hardware, power and load recorded with each run; memory footprint |

## Run

From the repository root, in the recipe's environment (`.venv`), with the frontend built:

```sh
python recipe/laya/bench/bench_inproc.py --device cpu --config C1 --run m1
python recipe/laya/bench/bench_inproc.py --device mps --config C2 --run m1
python recipe/laya/bench/bench_http.py --config C3 --run m1 --spawn .venv/bin/laya-serve
python recipe/laya/bench/bench_http.py --config C4 --run m1 --url http://127.0.0.1:8080 \
--frontend target/release/omni-jev --spawn .venv/bin/laya-serve
python recipe/laya/bench/bench_http.py --config C3w --run m1 \
--spawn .venv/bin/python -m frontend.laya_mps --device {device} --model {model} --port {port}
python recipe/laya/bench/bench_http.py --config C3o --run m1 \
--spawn .venv/bin/python -m frontend.laya_mps --device {device} --model {model} --compile --weights fp16 --port {port}
python recipe/laya/bench/report.py recipe/laya/bench/results/*_m[0-9].jsonl --ref C1
```

Repeat with `--run m2` for a second measured run. Runs refuse to start on battery power or above a
1-minute load average of `--max-load` (default 2) unless labelled `--run feasibility`. Memory is the
process's physical footprint, which on Apple Silicon includes MPS allocations.

Two configurations compared request by request, which holds up under background load better than
separate runs:

```sh
python recipe/laya/bench/paired.py --run p1 --a "" --b "--compile --weights fp16"
python recipe/laya/bench/paired.py --run f1 --a-url http://127.0.0.1:8000 --b-url http://127.0.0.1:8080
python recipe/laya/bench/paired.py --summarize recipe/laya/bench/results/paired_p1.jsonl
```

## Results

The measured runs on an M1 Pro are published as assets of one release on the fork,
<https://github.com/cacheline999/system1-omni/releases/tag/laya-mps-results-2026-09-28>:

| asset | contents | sha256 |
| --- | --- | --- |
| `laya-mps-reports-2026-10-01.tar.gz` | the tables: baseline report and parity, frontend overhead, paired fp16, paired all optimizations | `e857da5082da4983e07a91104b20f84fb1ffd7994c56123e0a832e0d7a870cea` |
| `laya-mps-results-2026-09-28.tar.gz` | raw JSONL of the baseline runs (C1–C4, C3w, C3s) | `611ed30707ac8c98875b5aa5382360b5a7d760da166d61c626eb07ebe1ee6404` |
| `laya-mps-paired-fp16-2026-09-30.tar.gz` | raw JSONL of the paired fp16 runs | `cc0d6f5bda6e3e0ee1f40c6966f84429e902a2f585b8e1ee33658a9be139326e` |
| `laya-mps-paired-all-2026-09-30.tar.gz` | raw JSONL of the paired all-optimizations runs | `25d1b7bc9d6dff7173f9b972ebd0204f8fb4e2089fe2d27924d48ba6e589780c` |

Extract the raw JSONL into `results/` and run `report.py` or `paired.py --summarize` on it to rebuild
the tables. `C3s` in the baseline runs is an earlier compile mode that compiled one-question requests
only; `--compile` does the same for them and adds the encoder for several questions.
Loading