Skip to content

Laya on Apple Silicon: worker, benchmarks and recipe - #30

Open
cacheline999 wants to merge 23 commits into
ThinkFlowLab:mainfrom
cacheline999:laya-apple-silicon
Open

cacheline999 wants to merge 23 commits into
ThinkFlowLab:mainfrom
cacheline999:laya-apple-silicon

Conversation

@cacheline999

@cacheline999 cacheline999 commented Sep 28, 2026 •

Copy link
Copy Markdown

Purpose

Closes #3. Serves Laya on Apple Silicon (PyTorch MPS) behind /v1/systemone and measures it.

Laid out like the Cua-S1 worker (#19): the HTTP worker is src/frontend/laya_mps.py, started as
PYTHONPATH=src python -m frontend.laya_mps --device mps --model english; the model-side code is
src/models/laya/{engine,optimize}.py; tests are in tests/laya/; the recipe pins its environment in
recipe/laya/requirements-mps.txt. The worker runs laya-serve (laya[serve]==0.3.20) with its request
handling unchanged and adds three optimizations and one fix:

  1. Compile (--compile): one-question requests run the whole model through
    torch.compile; requests with several questions compile only the encoder, because Laya's decision
    head is slower compiled on MPS. 68-token request on an M1 Pro: ~46 → ~33 ms (−29%).
  2. fp16 weights (--weights fp16): keep the checkpoint's fp16 weights instead of Laya's
    fp32 upcast on MPS (act_head stays fp32). A further −14% on top of compile, and
    4.2 → about 3 GB of memory for a worker on its own, on the six benchmark inputs; with --compile it
    grows with the input lengths seen (see "What the warm numbers leave out").
  3. Warmup before readiness: the worker binds only after every loaded model has run short, long and
    multi-question requests. laya-serve answers /health before any forward pass, so the first request
    after it took 0.7–1.1 s; through the worker it takes 70–81 ms (62–78 ms with both options), still more
    than a warm request because the GPU has been idle since the warmup. A checkpoint
    Laya loads later (a request that names another model, or routes to it) is prepared the same way inside
    its first request, and /health lists it under preparing meanwhile. With --compile that request
    took about 70 s on the M1 Pro, past the frontend's 60 s backend timeout: through the frontend it got 504
    and the same request sent again took 35 ms. The recipe says how to avoid it.
  4. Fix: /health tells the truth. It reports, read on every call, the device the model is actually on,
    dtypes, checkpoint and the revision the weights were loaded from. laya-serve echoes LAYA_DEVICE and keeps reporting
    mps after Laya falls back to CPU; --require-device makes the worker exit at startup instead, and
    after a fallback while serving the worker runs Laya's plain fp32 model and says so (compile.active turns
    false).

Together, 1 and 2 bring a 68-token one-question request to 0.62 of its time on the M1 Pro and 0.51 on an
M5, with every input faster and answers within 0.0031 of the unoptimized worker's (0.0054 on the M5);
results below.

  • recipe/laya/apple-silicon.md: setup, /health, frontend, tests and benchmark commands on a Mac.
    The six scripts behind the numbers below are in recipe/laya/bench/. The result tables and the raw
    JSONL are release assets on my fork
    (checksums in recipe/laya/bench/README.md), so no data files sit in the repository.

Test Plan

System1-Omni Version / Commit: based on 3062243 (main after #2). Each measurement ran on the working tree
committed right after it: baseline and one-question compile as 2a47358, fp16 weights as d3d0817, all
optimizations together as e10cf06. Each result file records the commit, versions, power and load.

  • Unit tests with a fake router: PYTHONPATH=src python -m pytest tests/laya (26).
  • Contract tests against a real CPU worker: LAYA_CONTRACT=1 PYTHONPATH=src python -m pytest tests/laya
    (13 more: all decision types, 400/401/413/422, first request after ready within 2× warm p50).
    Plain laya-serve fails exactly the two tests covering the readiness and device issues.
  • Recipe walked end to end in a clean copy with a fresh Python 3.12 venv.
  • Benchmarks on an M1 Pro (16 GB, macOS 26.1, torch 2.14.0, checkpoint 55cf4c4), on AC power, in two designs:
    • Separate runs (baseline, frontend, one-question compile): one feasibility run, then two measured runs
      (three for the configs that failed the gate after two) of 300 requests per input after 20 discarded,
      configs interleaved in the same rounds. A result counts only if all its measured runs agree on p50
      within 10%; no run was dropped. These started at a 1-minute load average ≤ 5 (the plan said ≤ 2, which
      this machine never reached), so absolute numbers may be a little high.
    • Paired runs (fp16 weights, all optimizations together): both workers alive, every request sent to
      each back to back in alternating order, 300 pairs per input after 20 discarded, two runs; the
      statistic is the median per-pair ratio with a 95% bootstrap interval. Separate runs could not
      resolve a 10% difference on this machine, so these had no load gate: they started at load 19–31
      (fp16) and 11–15 (all together), with intervals within ±3% of the median.
      Criteria were written down before each measured run.
  • ruff check / ruff format on the Python files (CI runs Rust checks only).

Test Result

All optimizations together

The worker with --compile --weights fp16 against the worker with neither,
both running at once, every request sent to each back to back (paired.py, 300 pairs per input). Declared
before measuring: a one-question median ratio ≤ 0.70 with the 95% interval below 0.75, no input slower,
answers within 1e-2, no graph compiled after ready. All held in both runs.

median optimized / baseline time run 1 run 2
1 question, 68 tokens 0.626 (0.620–0.633) 0.616 (0.609–0.628)
1 question, 47 tokens 0.616 (0.608–0.625) 0.584 (0.569–0.600)
1 question, 198 tokens 0.798 (0.792–0.803) 0.795 (0.778–0.814)
1 question, 484 tokens 0.833 (0.826–0.842) 0.832 (0.825–0.838)
3 questions 0.859 (0.853–0.868) 0.863 (0.857–0.868)
6 questions 0.818 (0.808–0.825) 0.824 (0.818–0.834)

In absolute terms the 68-token request went from 55–59 ms to 35–36 ms in these runs (both workers share the
GPU and the machine was loaded; a worker on its own measured ~46 ms before and ~28 ms after, the latter in
feasibility runs only). Answers within 0.0031 of the
baseline's, no decision changes. Worker memory 3.5 GB → 2.8 GB in these runs, where two workers shared the
machine; on its own a worker uses 4.2 GB without the options and about 3 GB with them on these six inputs. Ready after 35–39 s
instead of 8–10 s (two fresh starts each). Separately, the worker's warmup before readiness brings the first
request after /health from 0.7–1.1 s (plain laya-serve) to 70–81 ms, or 62–78 ms with both options. On an M5, 2 of 23 fresh starts still
gave a slow first request (327 and 409 ms, against 21–36 ms otherwise); not explained yet.

What the warm numbers leave out

Everything above is for requests sent back to back on six fixed inputs. Measured afterwards on the M1 Pro:

  • Idle gaps. A request that follows a pause is slower, with or without the options. A short
    one-question request takes 25 ms back to back with the options and 40 ms without; after 0.2–1 s of idle
    it took about 50 ms and 60–68 ms, after 2–5 s 105–115 ms and 114–127 ms. An agent that asks once every
    few seconds sees those numbers, and there the options are worth about 10%, not 37%. This is also why the
    first request after ready is slower than a warm one.
  • New input lengths. The first request of a length the worker has not seen costs about 15 ms more once
    with the options, 6 ms without. Not a recompile.
  • Memory grows with the lengths seen. With --compile, PyTorch keeps host memory for every input length
    the compiled model has run, about 5 MB each (fp16 weights alone add little): 2.8 → 3.3 GB after 100 new
    lengths, 5.3 GB after all 477, against 4.0 GB without the options. torch.mps.empty_cache() releases
    most of it and those lengths are cold again.
  • Download. The checkpoint took 97 s to download into an empty cache (846 MB at 8.7 MB/s); the ready
    times above are with it cached.

No regression from the review rounds

The worker was restructured several times during review, so the measured code (b00b285) was compared
with the reviewed one on the M1 Pro (AC power, load 7.6–8.1, 300 pairs per input, both workers alive):

input new / old, compile + fp16 new / old, no options options / none at the new commit (2 runs)
W1 1.001 (0.998–1.005) 0.999 (0.996–1.002) 0.620, 0.620
W2 0.999 (0.997–1.001) 1.000 (0.999–1.003) 0.801, 0.796
W3 1.002 (1.000–1.003) 1.000 (0.998–1.004) 0.826, 0.825
W4 1.002 (0.999–1.004) 0.999 (0.996–1.001) 0.856, 0.854
W5 1.000 (0.999–1.002) 1.000 (0.998–1.002) 0.811, 0.811
W6 1.005 (1.001–1.008) 1.001 (0.998–1.005) 0.608, 0.608

Old and new gave identical answers. The last column reproduces the paired results above within 0.03.
bench_http.py at f2b7c6a: ready after 10.6 s without the options and 39.3 s with them; first request
after ready 86 and 72 ms; footprint after warmup 4.2 and 2.8 GB.

What each part contributes

  • Compiling one-question requests. The profile (profile_mps.py) showed short requests bound by a
    fixed per-request cost from ~4,000 small ops; compiling cut the 68-token request from ~46 to ~33 ms
    (−29%) in two measured runs against the uncompiled worker, with multi-question requests unchanged.
  • fp16 weights (the checkpoint's own precision, so the conversion is exact). On top of the compiled
    path, two paired runs gave a 0.855–0.857 median ratio for the 68-token request. On the M1 Pro it did
    not help without compiling: fp16 GEMMs are ~25% faster there, but the fixed cost hid that. On an M5 it
    speeds up every input on its own (below).
  • Compiling the encoder for several questions. Compiling the whole model made multi-question requests
    slower: Laya's decision head (two nn.TransformerEncoderLayer with a padding mask) loses PyTorch's fused
    path when compiled and took ~55 ms instead of 3–12 ms. Compiling only the encoder made them ~14%
    faster in a same-process probe; the combined run above measures it together with the rest.

On an M5

@twu3202 ran the recipe and paired.py on an Apple M5 (10-core GPU, 32 GB, macOS 26.5.2) at e10cf06
(comment): all tests pass,
and against the worker with neither option the median ratios were 0.51–0.53 for one-question requests at
47–68 tokens, 0.30–0.33 at 198–484 tokens, 0.37 for three questions and 0.60 for six. There fp16 weights
alone speed up every input (0.43–0.81). The declared criteria hold on that machine too.

Tried and dropped

  • Compiling every batch: multi-question requests up to 65% slower, over a minute to start.
  • Static-shape compile with padded buckets (rows 1/2/4/8, lengths 64–512): multi-question requests 2–4x
    slower, one-question slower than dynamic.
  • ModernBERT's reference_compile: transformers 5.x no longer reads it (0 graphs, same latency).
  • fp16 weights without compile: no gain for one-question requests; one input deviated 0.0146 (over the
    1e-2 tolerance). Keeping encoder layer 3 in fp32 fixed that input but not held-out ones.
  • MLX or a Metal port: after compiling, the one-question path is close to the GEMM floor of this GPU
    through MPS (a 28-layer chain of Laya's four GEMMs alone takes 28.5 ms at 68 tokens, fp32, against
    ~33 ms for the whole request). MLX 0.32.3 was within 3% of PyTorch MPS or slower on 22 of 24 GEMM
    shape/precision cells (faster by 5% and 10% on two fp16 cells) and on the chain (32.7 vs 28.5 ms fp32, 25.2 vs 21.0 ms fp16), measured in one
    process at load ~9. A port would not buy speed on this hardware.
    This is about speed for the 0.4B Laya, which fits in memory; a larger model that has to be quantized
    to fit may decide differently ([RFC] Support emerging System 1 models, with CLM as the second model after LAYA #9).
  • A custom Metal GEMM for small M (torch.mps.compile_shader, 8×8 simdgroup matrices, fp16 weights, fp32
    accumulate, exact against fp32): timed alone it beat mm on three of Laya's four GEMM shapes, but in a
    28-layer chain at 68 tokens the best variant came to 0.997 of the fp16 mm chain and the others to
    1.05–1.29, with or without repacking the weights per column tile. The bar set beforehand was 0.90.
    Split-K through bmm did not help either (−4% at best). At 8 rows the chain still takes ~12 ms, of which
    dispatch is ~2 ms; the rest scales with the weights walked and is the same for mm and the custom
    kernel. So what remains at short inputs is not MPSGraph overhead but the cost of walking the weights,
    and --weights fp16 already runs them at the checkpoint's own precision. Measured at load 17–24, so
    ratios only.

Baseline

Warm p50 in ms at concurrency 1, range over the measured runs. * means the runs differ by more than
10%, so that number is indicative only.

68 tok, 1 q 198 tok 484 tok 3 q 6 q 47 tok noul
CPU in-process (3 runs) 134–144 212–238* 381–427* 198–226* 290–400* 108–123*
MPS in-process (2 runs) 44–45 68–69 145–150 71–72 139–145 37–38
MPS, laya-serve (3 runs) 45–46 68–72 147–156 72–76 139–141 37–42*
MPS, laya-serve + frontend (3 runs) 46–51* 72–87* 154–190* 72–90* 139–150 38–42
MPS, this worker (3 runs) 46–47 68–72 145–159 70–74 139–140 38–40
  • MPS is about 3x faster than CPU on the short requests, and HTTP through laya-serve adds about 1 ms
    over in-process. CPU numbers are unstable on this machine (other applications were running).
  • laya-serve runs one request at a time: at 4 concurrent clients throughput stays at ~22 req/s for
    the 68-token request and latency grows 4x.
  • The frontend runs failed the gate, but whole blocks of requests shifted together, which pointed at
    background load rather than the frontend. Sending every request directly and through the frontend
    back to back (now paired.py --a-url/--b-url; 120 pairs per input): the frontend adds +0.1 to +0.5 ms
    at the median, and 1 of 720 pairs differed by more than 10 ms.
  • Ready after 7.3–8.7 s for laya-serve and 7.9–9.3 s for the worker (checkpoint cached). First request
    after ready: 697–1135 ms for laya-serve, 75–81 ms for the worker.
  • Memory (physical footprint, includes MPS allocations): about 4.2 GB for the worker process on MPS,
    2.0 GB on CPU; the frontend uses 3 MB.

Output parity

Every configuration and run against CPU fp32: 330/330 questions with the same decision and |Δp|
within 1e-3 (1e-2 where laya autocasts to fp16).

Follow-ups (not in this PR)

  • Packing several questions into one row. A multi-question request pads every row to the longest
    (20% of the GEMM rows in the 3-question input are padding) and runs the decision head eagerly. A
    block-diagonal mask with per-row positions is exact and would make every request a single row on the
    compiled path; estimated −12 to −17% for the 3-question input, nothing for one question.
  • Idle gaps. The largest thing left for real use (numbers above). In a probe, one forward pass every
    0.5 s held a request after any pause at about 45 ms for 8% GPU load, and every 0.2 s at 31–41 ms for 18%;
    a tiny GPU operation as heartbeat did nothing. Not in this PR: the worker reuses laya-serve's request
    handling, whose inference thread it cannot reach, and driving MPS from a second thread aborted the process.
    The slow first requests on the M5 (327 and 409 ms in 2 of 23 starts) are probably the same effect.

Not verified: M-series chips other than the M1 Pro and M5.

Self-review

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

Done at 005a874, after laying the files out like #19 (worker in src/frontend/, model code in
src/models/laya/, tests in tests/laya/, flags instead of environment variables, pinned requirements,
six benchmark scripts instead of ten, no result files in the repository). The diff review found and fixed: warmup and /health covering only one of several loaded
models; the revision read from the cache ref instead of the loaded snapshot; three error-path bugs in
bench_http.py; fp16 weights described as suitable for CPU (same answers there, but 334 ms instead of
138 ms; now documented and warned about); and /health describing the model as it was at startup, so a
later fallback to the CPU stayed invisible (now read on every call).

The recipe was then checked line by line against the measurements and a real worker. Corrected: ready
time with both options (35–39 s, not the 20–30 s of the earlier one-question-only compile), the memory
baseline, which configuration the first-request numbers belong to, the uv route to Python 3.12, and the
first-pass report command, which failed once paired results were in the directory. --require-device
was exercised on a real worker with MPS reported unavailable: exit code 1 with the documented message.

The second review's findings (checkpoints loaded while serving, the report's parity section on paired
files, revisions of bundled checkpoints) are fixed in 584680d. Checked on the M1 Pro with the real
multilingual checkpoint: loaded by an explicit model and by a German state with --compile --weights fp16,
both checkpoints described in /health with revision 55cf4c4; and with --require-device and the MPS
limit lowered so the late checkpoint falls back to CPU, it is unloaded, that request gets 500 and the
preloaded checkpoint keeps serving.

A further pass over the whole diff, each finding reproduced by a test before fixing (87c920f): a late
checkpoint that fails preparation no longer costs the checkpoint Laya evicted for it; the docs said
laya-serve's environment variables still apply, but only LAYA_API_KEY does, and the worker now warns
about the others; paired.py timed requests with retries and its summary did not show a side that fell
back to CPU; two cleanup and polling faults in the benchmark scripts. The evicted-checkpoint fix is covered
by a unit test on Laya's real Router with stub agents only: it needs three checkpoints to trigger.

A second pass, again reproduced before fixing (1450884): the late-load cost above, which the recipe had
described as seconds; benchmark scripts leaving a worker running when the next process failed to start, and
not passing --model/--device to a spawned frontend.laya_mps; /health reporting recompiles forever
after a late checkpoint failed its warmup; and the first-pass report having no parity reference.

The third review's findings (runs whose answer probes all failed vanishing from the report, and a
checkpoint's revision changing when another one is downloaded from the same repository) are fixed in
f2b7c6a, together with four more crashes of the result readers on damaged files that a self-audit found.
The contract tests now also run against a worker on the GPU (LAYA_CONTRACT_DEVICE=mps, with and without
LAYA_CONTRACT_FLAGS="--compile --weights fp16") and compare its answers with Laya run directly; 20 passed
in each of the three configurations.

A further self-review fixed in 668f331: a profile run's results no longer show as parity failures in
report.py, the run-header test no longer assumes macOS tools, and the memory growth above is attributed
to --compile rather than to the fp16 weights.

Checks at 668f331: 72 unit tests, 92 with LAYA_CONTRACT=1, ruff check and ruff format --check on
the Python files. cargo fmt --all --check, cargo clippy --workspace --locked --all-targets -- -D warnings
and cargo test --workspace --locked passed at 87c920f, after merging main; no Rust file is changed by this PR. Laya's fallback to CPU
was triggered on the real GPU by lowering PyTorch's MPS memory limit (PYTORCH_MPS_HIGH_WATERMARK_RATIO)
and sending a 64-question request. The request still returned 200 and /health switched to device: cpu,
device_mismatch: true. The first run of this showed a worker with compile and fp16 weights degrading
badly afterwards (336 s for the request that fell back, a 28 s recompile, then 370–550 ms per request),
so the worker now runs Laya's fp32 model uncompiled whenever the inputs are on the CPU. Rerun with both
options, fp16 only and compile only: 32–40 s for the request that fell back, then 140–270 ms, no graph
recompiled, answers within 0.0002 of those before the fallback.

Add a Laya worker (src/models/laya/worker.py) that runs laya-serve unchanged
except for startup and /health: it binds only after a warmup, and /health
reports the device, dtypes, checkpoint and revision the model actually uses.
LAYA_REQUIRE_DEVICE=1 exits instead of serving on the wrong device, and
LAYA_WORKER_COMPILE=single compiles the one-question path.

Add benchmarks/laya: fixed inputs, in-process and HTTP benchmarks (laya-serve
directly and behind the Rust frontend), output parity against CPU with
tolerances fixed in advance, an MPS profile and a report built from raw JSONL.

Add recipe/laya/apple-silicon.md with setup, frontend, tests and benchmark
commands on a Mac.
Reports built from the measured runs on an M1 Pro (CPU and MPS in-process,
laya-serve, laya-serve behind the Rust frontend, the worker with and
without LAYA_WORKER_COMPILE=single): latency tables, output parity and
the paired frontend overhead. The raw JSONL is published as a release
asset of the fork; benchmarks/laya/README.md has the link and checksum.

frontend_overhead.py sends each request directly and through the frontend
back to back, so background load that shifts separate runs cancels out.
Worker: warm up, compile and describe every model the router has loaded,
not only LAYA_WORKER_MODEL. With LAYA_MODELS unset laya-serve preloads all
checkpoints, and a request auto-routed to one of the others was served
cold. A model that was not preloaded is no longer loaded just for warmup.
/health lists each model under `models`; device_mismatch is set if any
model is off the requested device.

Worker: report the revision the weights were actually downloaded from,
recorded from snapshot_download's path, instead of reading refs/main from
the cache, which can move or not apply to a local checkpoint.

bench_http: close the connection on any failed request so the next one
reconnects (a timeout used to turn every later request on that thread
into CannotSendRequest), record a failed answers request instead of
aborting the run, and count only successful requests in req/s.
Keep the repository to the src/ and recipe/ layout: the scripts that
reproduce the recipe's numbers now live next to it, like the Cua-S1
benchmarks under recipe/cua_s1/. The README describes only how to run
them and where the results are.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed commit d44d3fe31834b00f517f386e0081f3eea3cb84bc.

No actionable source findings. Reviewed warmup across loaded models, actual-device health reporting, compile wrapper, upstream route reuse, HTTP/CPU benchmarks and parity checks. Unit-test collection blocked: available Laya environments lack the httpx2 dependency required by their installed Starlette; no Apple Silicon execution or checkpoint contract run.

Keep the checkpoint's fp16 weights instead of laya's fp32 upcast on MPS
(act_head stays fp32, since laya feeds it fp32 features). With
LAYA_WORKER_COMPILE=single, two paired runs with both workers alive
lowered warm p50 by about 14% for one-question requests (median ratio
0.855-0.857 at 68 tokens), 9-11% for longer and six-question requests,
raised it 4% for three questions, and cut the worker's memory from
3.6 GB to 2.7 GB. Answers stayed within 0.0031 of the fp32 worker's.

paired.py runs two worker configurations at once and sends each request
to both back to back; separate runs could not resolve a 10% difference
under background load. parity.py now applies the fp16 tolerance to every
request of a worker whose weights are fp16.
LAYA_WORKER_COMPILE is now off or on. On compiles one-question batches
end to end, as before, and for several questions compiles only the
encoder: Laya's decision head is two nn.TransformerEncoderLayer with a
padding mask, which lose PyTorch's fused path when compiled and were
about 50 ms slower on padded batches. The mode that compiled every batch
is removed.

With LAYA_WORKER_COMPILE=on and LAYA_WORKER_WEIGHTS=fp16 against the
worker with neither, in two paired runs: median latency ratio 0.62 for a
68-token one-question request, 0.80-0.83 at 198-484 tokens, 0.86 for
three questions and 0.82 for six; no input slower; answers within 0.0031;
worker memory 3.5 GB -> 2.8 GB.
@twu3202

twu3202 commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

The description lists other M-series chips as not verified, so I ran this on an Apple M5 (10-core GPU, 32 GB, macOS 26.5.2) at e10cf06.

Following the recipe in a fresh Python 3.12 venv (on other ports, since 8000 was taken) worked, including the first checkpoint download (55cf4c4, about 3 minutes). With LAYA_WORKER_COMPILE=on LAYA_WORKER_WEIGHTS=fp16, /health reported mps, fp16 weights and 3 graphs at ready, and the example request gave byte-identical responses directly and through the frontend. In that venv (Starlette 1.7.0, httpx 0.28.1) the 25 unit tests pass, and LAYA_CONTRACT=1 adds the 13 contract tests, which pass too.

Paired runs with paired.py, 300 pairs per input, median time ratio against the worker with neither option:

input on + fp16 on only fp16 only
W1, 68 tokens 0.514 0.830 0.769
W2, 198 tokens 0.334 0.921 0.476
W3, 484 tokens 0.303 0.947 0.426
W4, 3 questions 0.369 0.890 0.522
W5, 6 questions 0.595 0.844 0.725
W6, noul, 47 tokens 0.531 0.819 0.811
  • The criteria declared in the description hold here too: every one-question ratio is at most 0.531 with its 95% interval below 0.54, no input is slower, the largest |Δp| is 0.0054 with both options (0.0005 with on only, 0.0017 with fp16 only), and no graph was compiled after ready. No decision changed.
  • Against CPU fp32, the worker with both options and the worker with neither are within tolerance on all 22 questions each.
  • The worker was ready after 18.7 s with both options, against 3.4 s without, and its footprint went from 3.6 to 3.0 GB.
  • First one-question request after ready: in 23 fresh starts of the worker with both options it took 21 to 36 ms 21 times, and 327 and 409 ms the other two, which I could not tie to anything. laya-serve took 132 to 143 ms in 6 starts.

So the gains are larger on the M5 than on the M1 Pro, mostly from the fp16 weights: fp16 alone speeds up every input here, and W3 goes from 134 to 41 ms with both options.

The 1-minute load was between 2 and 4 during these runs, above the scripts' default gate of 2, so I used --max-load 8 and one run per configuration. The paired ratios are the numbers I would rely on.

Add the M5 validation from the PR discussion, where fp16 weights speed up
every input even without compile. Install httpx2 for the tests: recent
Starlette's TestClient needs it (httpx is only a deprecated fallback).
State that the first request after ready was occasionally slow on an M5
(2 of 23 fresh starts).
@cacheline999

Copy link
Copy Markdown
Author

@twu3202 Thanks for running this on an M5, and for checking it from a fresh venv including the download. I've taken three things from it into the PR (b00b285): the recipe lists the M5 and its numbers, the note that fp16 weights don't help without compile is now marked as M1 Pro only, and the occasional slow first request is written down as open.

About the two slow starts (327 and 409 ms): do you remember which input the first request was, and was it the same both times? The warmup covers one, three and six questions at a few lengths. If those two requests had a shape outside that, it points at the warmup. If it was the same short request that was fast in the other 21 starts, it's more likely something outside the worker.

@cacheline999

Copy link
Copy Markdown
Author

@hsliuustc0106 Thanks for the review. The test collection failure comes from Starlette's TestClient: recent versions need httpx2 and only fall back to httpx with a deprecation warning, and laya[serve] pulls in neither. The recipe now installs pytest httpx2 for the tests (b00b285); with that the 25 unit tests pass without the fallback. The 13 contract tests also need the checkpoint and run with LAYA_CONTRACT=1.

@twu3202

twu3202 commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

They were different requests. Both were the first request after /health turned ready, with LAYA_WORKER_COMPILE=on LAYA_WORKER_WEIGHTS=fp16:

  • 409 ms: one noul question ("Does the customer ask for a refund?") on "Please refund the duplicate charge.", 40 input tokens. The same request went first in 10 other starts and took 23 to 31 ms.
  • 327 ms: W1 from bench_http.py, which sends W1 first. That was the only start of that kind.

I ran both again at b00b285 (same worker code) with fresh starts: 30 with the noul request first, half of them right after a start without the options as in the 409 ms case, and 15 with W1 first. The first request took 24 to 38 ms for noul and 26 to 50 ms for W1, so neither slow start came back. recompiled_after_ready stayed false in all 45, and with TORCH_LOGS=recompiles each start logged one recompile, the autocast switch during warmup. The 1-minute load was 2.7 to 5.5.

So I can't say what caused the two, but it doesn't look tied to either request.

On CPU LAYA_WORKER_WEIGHTS=fp16 gives the same answers (max |dp| 0.0013)
but a 68-token request took 334 ms instead of 138 ms. The option's
description said "on MPS and CPU"; it now says it is meant for MPS, and
the worker logs a warning when a model with fp16 weights is on the CPU.
/health described the models as they were after warmup. laya moves a
model to the CPU when a request runs out of GPU memory and keeps
serving, so the device, dtypes and device_mismatch are now read on
every call.

Recipe, checked line by line against the measurements and a real worker:
- ready time with both options is 35-39 s on the M1 Pro, not 20-30 s
  (that was the earlier one-question-only compile);
- memory is 4.2 GB -> about 3 GB for a worker on its own (3.5 -> 2.8 GB
  was measured with two workers sharing the machine);
- first-request numbers are attributed to the configuration they were
  measured on;
- the uv route to Python 3.12 is `uv venv --python 3.12 --seed .venv`;
- LAYA_REQUIRE_DEVICE was exercised on a real worker with MPS reported
  unavailable: exit code 1 with the documented message;
- the first-pass report command globs *_feasibility.jsonl, and report.py
  skips result files of another kind instead of raising KeyError.
laya moves the model to the CPU when a request runs out of GPU memory and
keeps serving. Triggered on the real GPU by lowering PyTorch's MPS memory
limit, a worker with compile and fp16 weights then kept its fp16 weights
and recompiled its graphs for the CPU: the request that fell back took
336 s instead of 28 s, the next one 28 s, later ones 370-550 ms.

The worker's model wrapper now runs laya's own model in fp32 whenever the
inputs are on the CPU, restoring the weights once. After the same
fallback: 32-40 s for the request that fell back, then 140-270 ms, no
graph recompiled, /health showing cpu and float32. The two options are
no longer applied to a model that is on the CPU at startup.

The GPU path is unchanged: a paired check gave the same ratios as before
(0.62 for the 68-token request) and answers within 0.0031.
worker.py keeps startup, warmup and /health. optimize.py holds what
changes the model itself: fp16 weights, compile, the wrapper that runs
laya's own fp32 model on the CPU, and the rule that the options apply on
the GPU only. No behaviour change: same tests, and a paired check gave
the same ratios and answers as before.

README: the first-request numbers were from an early feasibility run
(230 ms / 73 ms); the measured ones are 0.7-1.1 s for laya-serve and
70-81 ms for the worker. It now lists what the worker adds the same way
as the PR does, says LAYA_REQUIRE_DEVICE is checked at startup, and
mentions the M5 validation.
@hsliuustc0106

Copy link
Copy Markdown
Contributor

please follow the structure of PR #19

- src/frontend/laya_mps.py: the HTTP worker, started with flags
  (PYTHONPATH=src python -m frontend.laya_mps --device mps --model english
  [--compile] [--weights fp16] [--require-device]) instead of LAYA_WORKER_*
  environment variables. laya-serve's own variables (LAYA_API_KEY, ...)
  still apply.
- src/models/laya/: model-side code only. engine.py holds the warmup,
  the per-model description /health returns and the revision record;
  optimize.py is unchanged.
- tests/laya/: the unit and contract tests, run with PYTHONPATH=src.
- recipe/laya/requirements-mps.txt pins the validated environment.
- recipe/laya/bench/: 10 scripts become 6. The answer comparison is a
  section of report.py, the frontend-overhead probe is paired.py's
  --a-url/--b-url mode, and the workload generator and token check are
  replaced by the committed workloads.jsonl plus its token counts in the
  README. The five result tables leave the repository and join the raw
  data as a release asset.
- README: LAYA's row in the models table points at the Apple Silicon
  worker.

No behaviour change: 26 unit and 13 contract tests pass, and a
feasibility pass of every script gives the same numbers as before
(68-token request 28 ms with both options, paired ratio 0.64, answers
within 0.0031, frontend overhead ratio 1.00-1.01).
@cacheline999

cacheline999 commented Oct 1, 2026 •

Copy link
Copy Markdown
Author

The PR description is updated to match. Let me know if the layout should go further.

Done in 005a874, and main is merged in (3d6cb57):

  • the HTTP worker is src/frontend/laya_mps.py, started like the Cua-S1 one: PYTHONPATH=src python -m frontend.laya_mps --device mps --model english [--compile] [--weights fp16], flags instead of the LAYA_WORKER_* variables;
  • src/models/laya/ keeps only the model-side code (engine.py, optimize.py);
  • tests are in tests/laya/, with recipe/laya/requirements-mps.txt pinning the environment;
  • the benchmark directory went from ten scripts to six, and the result tables left the repository and sit next to the raw data as release assets.

The PR description is updated to match. Let me know if the layout should go further.

…, honour --log-level

- Docstrings in optimize.py and profile_mps.py no longer carry one machine's numbers or a private spec id.
- /health gains compile.active, false once no model runs the compiled path (after a CPU fallback).
- --log-level applies to the worker's own log as well as uvicorn's.
- Startup failures print their traceback before exiting.
- Warn when laya's autocast row threshold is above what the warmup covers.
- compile_agent is a no-op on an already compiled agent; the wrapper is one class, found by isinstance.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed commit 3a7850751df2a3fc62b1e677c09b3971bc37a8cc.

Three actionable findings are attached inline: startup-only agent tracking misses later checkpoint loads, mixed benchmark/paired input crashes the documented report command, and bundled checkpoint revisions are missing from health.

Validation: 29 unit tests passed, 13 optional checkpoint contract tests skipped; ruff check, ruff format --check --line-length 120, git diff --check, and the strict MkDocs build passed. The health findings were reproduced through Laya 0.3.20's real Router and upstream HTTP routes with checkpoint loading/inference and snapshot download boundaries replaced. The report failure was reproduced with synthetic result files.

Execution limits: reused a CPU environment with torch 2.8.0+cpu and transformers 4.55.0, which differs from the Apple Silicon recipe's pins. No Apple Silicon execution, real-checkpoint contract tests, inference benchmarks, or actual GPU-to-CPU fallback runs were performed. CI and Docs runs for this head require approval.

Comment thread src/frontend/laya_mps.py Outdated
Comment thread recipe/laya/bench/report.py Outdated
Comment thread src/models/laya/engine.py Outdated
…ity on paired files and bundled revisions

- A checkpoint laya loads for a later request gets the same options, warmup and device check through the
  Router's on_load hook; an evicted one drops out of /health; one that cannot be prepared is unloaded.
  /health names a checkpoint under `preparing` while that runs.
- report.py: the parity section uses the records read() already filtered, so paired files no longer crash it.
- /health finds the revision of a bundled checkpoint under the download repository.
- Comments: drop a private spec id, headings repeated as comments, and a docstring restating its function.
# Conflicts:
#	README.md
#	src/models/laya/README.md
@cacheline999

Copy link
Copy Markdown
Author

All three are fixed in 584680d, and main is merged in (7dc6b68).

What I could run here that the review couldn't, on the M1 Pro:

  • a real multilingual load while serving, by explicit model and by a German state, with --compile --weights fp16: prepared in the first request, both checkpoints in /health, revision 55cf4c4 shown for the bundled one;
  • --require-device with the MPS limit lowered so the late checkpoint falls back to CPU: it is unloaded, that request gets 500, the preloaded one keeps serving.

…which laya-serve variables apply

- A late checkpoint that cannot be prepared is unloaded and the checkpoints laya evicted for it are loaded again.
- Only the variables laya-serve's app reads apply (LAYA_API_KEY); the worker warns about the launcher's ones.
- A second app built on the same router replaces the first app's hooks.
- paired.py: timed requests are not retried and failed pairs are counted; the summary shows each side's
  device at the end and flags a side that left its device or compiled path.
- bench: wait_ready keeps polling on a malformed response; cleanup terminates every process before waiting.
…nd; bench spawn fixes

- Recipe: the first request for a late checkpoint took about 70 s with --compile on the M1 Pro, past the
  frontend's 60 s backend timeout (504, then fine on retry); how to avoid it; memory with two checkpoints.
- /health: graphs compiled by a late checkpoint that is then unloaded are not reported as recompiles.
- bench_http.py, paired.py: a process that fails to start no longer leaves the other one running.
- bench_http.py: {device} and {model} in the spawn command, so --device/--model reach frontend.laya_mps.
- First-pass report compares against its own C2 run.

Copy link
Copy Markdown
Contributor

[P2] Preserve runs whose answer probes all fail

Reviewed 543acef. Location: recipe/laya/bench/report.py:233–240.

read_answers() creates a run entry only after a successful answers record, so this loop entirely omits configurations whose probes all produced answers_error. I reproduced this through bench_http.main() and report.main(), using the real W1/W4 workloads and mocked HTTP 200 responses without answers: C3 emitted two answer errors, but the report showed its latency and zero HTTP errors alongside a ‘4/4 questions within tolerance’ summary containing only healthy C2. Enumerate runs from environment/error records and represent failed probes explicitly rather than dropping the configuration.

Copy link
Copy Markdown
Contributor

[P3] Bind revisions to the checkpoint actually loaded

Reviewed 543acef. Location: src/models/laya/engine.py:78–81.

The bundled-repo lookup fixes missing revisions, but it reads a mutable per-repository value on every health call. If English loads snapshot A and a later multilingual load downloads snapshot B from the same repository, recording B overwrites the dictionary entry and /health now reports B for English despite its weights still being A. Reproduced with Laya 0.3.20’s real Router and HTTP routes, replacing only Agent loading/inference and snapshot-download boundaries: the multilingual request returned 200 and English’s reported revision changed A→B while its actual loaded revision stayed A. Bind revision identity to each successful checkpoint/agent load.

…nd what the warm numbers leave out

Worker and benchmark scripts
- /health keeps each checkpoint's revision from its own load: revisions are recorded per checkpoint (repo or
  repo/subfolder) and read once when the checkpoint has been prepared.
- report.py: the parity section lists every run with an env record; a run whose answer probes all failed, or
  an answer missing a question, shows as failures instead of disappearing or raising.
- paired.py --summarize: a run without an end record, a side without answers or a question missing on one
  side is reported instead of raising.

Tests
- Contract tests: answers match Laya run directly (choice, score, noul, combined, six questions), the port
  opens only after the warmup, the checkpoint stays loaded; LAYA_CONTRACT_DEVICE=mps and LAYA_CONTRACT_FLAGS
  run the same contract on the GPU, without and with the options.
- /health is checked against Laya's real Router after random sequences of loads, failed loads and evictions.
- Unit tests from the issue's requirements, and for behaviours a mutation run showed were executed but not
  asserted (late checkpoints get the GPU options before their warmup, logging, defaults, /health values).

Recipe
- What back-to-back latencies leave out, measured on the M1 Pro: a request after 2-5 s of idle takes
  105-115 ms against 25 ms; a new input length costs about 15 ms once; with both options the footprint grows
  about 5 MB per input length seen, to 5.3 GB after all of them.
- Checkpoint download: 97 s for 846 MB. Contract tests on the GPU.
@cacheline999

Copy link
Copy Markdown
Author

Both fixed in f2b7c6a.

  • report: the parity section lists every run with an env record; a run whose probes all failed, or an answer missing a question, shows as failures.
  • revision: recorded per checkpoint (repo or repo/subfolder) and read once when the checkpoint has been prepared, so a later download from the same repository no longer changes what a loaded one reports. Tested with Laya's real Router and two checkpoints of one repository at different commits.

…cOS; memory growth is from compile

- report.py: the parity section lists benchmark runs (they write a phase record before any answers) and runs
  with answers or failed probes; a profile_mps.py run, which never answers, no longer shows as failures.
- The run-header test asserts the macOS machine probes only on macOS, and checks that they degrade to None
  elsewhere.
- Recipe: the per-length memory growth comes from --compile, not from fp16 weights.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Help wanted] Accelerated Laya serving on Apple Silicon

3 participants