Laya on Apple Silicon: worker, benchmarks and recipe - #30
cacheline999 wants to merge 23 commits into
Conversation
Add a Laya worker (src/models/laya/worker.py) that runs laya-serve unchanged except for startup and /health: it binds only after a warmup, and /health reports the device, dtypes, checkpoint and revision the model actually uses. LAYA_REQUIRE_DEVICE=1 exits instead of serving on the wrong device, and LAYA_WORKER_COMPILE=single compiles the one-question path. Add benchmarks/laya: fixed inputs, in-process and HTTP benchmarks (laya-serve directly and behind the Rust frontend), output parity against CPU with tolerances fixed in advance, an MPS profile and a report built from raw JSONL. Add recipe/laya/apple-silicon.md with setup, frontend, tests and benchmark commands on a Mac.
Reports built from the measured runs on an M1 Pro (CPU and MPS in-process, laya-serve, laya-serve behind the Rust frontend, the worker with and without LAYA_WORKER_COMPILE=single): latency tables, output parity and the paired frontend overhead. The raw JSONL is published as a release asset of the fork; benchmarks/laya/README.md has the link and checksum. frontend_overhead.py sends each request directly and through the frontend back to back, so background load that shifts separate runs cancels out.
96747b3 to
326838a
Compare
Worker: warm up, compile and describe every model the router has loaded, not only LAYA_WORKER_MODEL. With LAYA_MODELS unset laya-serve preloads all checkpoints, and a request auto-routed to one of the others was served cold. A model that was not preloaded is no longer loaded just for warmup. /health lists each model under `models`; device_mismatch is set if any model is off the requested device. Worker: report the revision the weights were actually downloaded from, recorded from snapshot_download's path, instead of reading refs/main from the cache, which can move or not apply to a local checkpoint. bench_http: close the connection on any failed request so the next one reconnects (a timeout used to turn every later request on that thread into CannotSendRequest), record a failed answers request instead of aborting the run, and count only successful requests in req/s.
Keep the repository to the src/ and recipe/ layout: the scripts that reproduce the recipe's numbers now live next to it, like the Cua-S1 benchmarks under recipe/cua_s1/. The README describes only how to run them and where the results are.
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed commit d44d3fe31834b00f517f386e0081f3eea3cb84bc.
No actionable source findings. Reviewed warmup across loaded models, actual-device health reporting, compile wrapper, upstream route reuse, HTTP/CPU benchmarks and parity checks. Unit-test collection blocked: available Laya environments lack the httpx2 dependency required by their installed Starlette; no Apple Silicon execution or checkpoint contract run.
Keep the checkpoint's fp16 weights instead of laya's fp32 upcast on MPS (act_head stays fp32, since laya feeds it fp32 features). With LAYA_WORKER_COMPILE=single, two paired runs with both workers alive lowered warm p50 by about 14% for one-question requests (median ratio 0.855-0.857 at 68 tokens), 9-11% for longer and six-question requests, raised it 4% for three questions, and cut the worker's memory from 3.6 GB to 2.7 GB. Answers stayed within 0.0031 of the fp32 worker's. paired.py runs two worker configurations at once and sends each request to both back to back; separate runs could not resolve a 10% difference under background load. parity.py now applies the fp16 tolerance to every request of a worker whose weights are fp16.
LAYA_WORKER_COMPILE is now off or on. On compiles one-question batches end to end, as before, and for several questions compiles only the encoder: Laya's decision head is two nn.TransformerEncoderLayer with a padding mask, which lose PyTorch's fused path when compiled and were about 50 ms slower on padded batches. The mode that compiled every batch is removed. With LAYA_WORKER_COMPILE=on and LAYA_WORKER_WEIGHTS=fp16 against the worker with neither, in two paired runs: median latency ratio 0.62 for a 68-token one-question request, 0.80-0.83 at 198-484 tokens, 0.86 for three questions and 0.82 for six; no input slower; answers within 0.0031; worker memory 3.5 GB -> 2.8 GB.
|
The description lists other M-series chips as not verified, so I ran this on an Apple M5 (10-core GPU, 32 GB, macOS 26.5.2) at Following the recipe in a fresh Python 3.12 venv (on other ports, since 8000 was taken) worked, including the first checkpoint download ( Paired runs with
So the gains are larger on the M5 than on the M1 Pro, mostly from the fp16 weights: fp16 alone speeds up every input here, and W3 goes from 134 to 41 ms with both options. The 1-minute load was between 2 and 4 during these runs, above the scripts' default gate of 2, so I used |
Add the M5 validation from the PR discussion, where fp16 weights speed up every input even without compile. Install httpx2 for the tests: recent Starlette's TestClient needs it (httpx is only a deprecated fallback). State that the first request after ready was occasionally slow on an M5 (2 of 23 fresh starts).
|
@twu3202 Thanks for running this on an M5, and for checking it from a fresh venv including the download. I've taken three things from it into the PR (b00b285): the recipe lists the M5 and its numbers, the note that fp16 weights don't help without compile is now marked as M1 Pro only, and the occasional slow first request is written down as open. About the two slow starts (327 and 409 ms): do you remember which input the first request was, and was it the same both times? The warmup covers one, three and six questions at a few lengths. If those two requests had a shape outside that, it points at the warmup. If it was the same short request that was fast in the other 21 starts, it's more likely something outside the worker. |
|
@hsliuustc0106 Thanks for the review. The test collection failure comes from Starlette's |
|
They were different requests. Both were the first request after
I ran both again at b00b285 (same worker code) with fresh starts: 30 with the So I can't say what caused the two, but it doesn't look tied to either request. |
On CPU LAYA_WORKER_WEIGHTS=fp16 gives the same answers (max |dp| 0.0013) but a 68-token request took 334 ms instead of 138 ms. The option's description said "on MPS and CPU"; it now says it is meant for MPS, and the worker logs a warning when a model with fp16 weights is on the CPU.
/health described the models as they were after warmup. laya moves a model to the CPU when a request runs out of GPU memory and keeps serving, so the device, dtypes and device_mismatch are now read on every call. Recipe, checked line by line against the measurements and a real worker: - ready time with both options is 35-39 s on the M1 Pro, not 20-30 s (that was the earlier one-question-only compile); - memory is 4.2 GB -> about 3 GB for a worker on its own (3.5 -> 2.8 GB was measured with two workers sharing the machine); - first-request numbers are attributed to the configuration they were measured on; - the uv route to Python 3.12 is `uv venv --python 3.12 --seed .venv`; - LAYA_REQUIRE_DEVICE was exercised on a real worker with MPS reported unavailable: exit code 1 with the documented message; - the first-pass report command globs *_feasibility.jsonl, and report.py skips result files of another kind instead of raising KeyError.
laya moves the model to the CPU when a request runs out of GPU memory and keeps serving. Triggered on the real GPU by lowering PyTorch's MPS memory limit, a worker with compile and fp16 weights then kept its fp16 weights and recompiled its graphs for the CPU: the request that fell back took 336 s instead of 28 s, the next one 28 s, later ones 370-550 ms. The worker's model wrapper now runs laya's own model in fp32 whenever the inputs are on the CPU, restoring the weights once. After the same fallback: 32-40 s for the request that fell back, then 140-270 ms, no graph recompiled, /health showing cpu and float32. The two options are no longer applied to a model that is on the CPU at startup. The GPU path is unchanged: a paired check gave the same ratios as before (0.62 for the 68-token request) and answers within 0.0031.
worker.py keeps startup, warmup and /health. optimize.py holds what changes the model itself: fp16 weights, compile, the wrapper that runs laya's own fp32 model on the CPU, and the rule that the options apply on the GPU only. No behaviour change: same tests, and a paired check gave the same ratios and answers as before. README: the first-request numbers were from an early feasibility run (230 ms / 73 ms); the measured ones are 0.7-1.1 s for laya-serve and 70-81 ms for the worker. It now lists what the worker adds the same way as the PR does, says LAYA_REQUIRE_DEVICE is checked at startup, and mentions the M5 validation.
|
please follow the structure of PR #19 |
# Conflicts: # recipe/README.md
- src/frontend/laya_mps.py: the HTTP worker, started with flags (PYTHONPATH=src python -m frontend.laya_mps --device mps --model english [--compile] [--weights fp16] [--require-device]) instead of LAYA_WORKER_* environment variables. laya-serve's own variables (LAYA_API_KEY, ...) still apply. - src/models/laya/: model-side code only. engine.py holds the warmup, the per-model description /health returns and the revision record; optimize.py is unchanged. - tests/laya/: the unit and contract tests, run with PYTHONPATH=src. - recipe/laya/requirements-mps.txt pins the validated environment. - recipe/laya/bench/: 10 scripts become 6. The answer comparison is a section of report.py, the frontend-overhead probe is paired.py's --a-url/--b-url mode, and the workload generator and token check are replaced by the committed workloads.jsonl plus its token counts in the README. The five result tables leave the repository and join the raw data as a release asset. - README: LAYA's row in the models table points at the Apple Silicon worker. No behaviour change: 26 unit and 13 contract tests pass, and a feasibility pass of every script gives the same numbers as before (68-token request 28 ms with both options, paired ratio 0.64, answers within 0.0031, frontend overhead ratio 1.00-1.01).
Done in 005a874, and main is merged in (3d6cb57):
The PR description is updated to match. Let me know if the layout should go further. |
…, honour --log-level - Docstrings in optimize.py and profile_mps.py no longer carry one machine's numbers or a private spec id. - /health gains compile.active, false once no model runs the compiled path (after a CPU fallback). - --log-level applies to the worker's own log as well as uvicorn's. - Startup failures print their traceback before exiting. - Warn when laya's autocast row threshold is above what the warmup covers. - compile_agent is a no-op on an already compiled agent; the wrapper is one class, found by isinstance.
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed commit 3a7850751df2a3fc62b1e677c09b3971bc37a8cc.
Three actionable findings are attached inline: startup-only agent tracking misses later checkpoint loads, mixed benchmark/paired input crashes the documented report command, and bundled checkpoint revisions are missing from health.
Validation: 29 unit tests passed, 13 optional checkpoint contract tests skipped; ruff check, ruff format --check --line-length 120, git diff --check, and the strict MkDocs build passed. The health findings were reproduced through Laya 0.3.20's real Router and upstream HTTP routes with checkpoint loading/inference and snapshot download boundaries replaced. The report failure was reproduced with synthetic result files.
Execution limits: reused a CPU environment with torch 2.8.0+cpu and transformers 4.55.0, which differs from the Apple Silicon recipe's pins. No Apple Silicon execution, real-checkpoint contract tests, inference benchmarks, or actual GPU-to-CPU fallback runs were performed. CI and Docs runs for this head require approval.
…ity on paired files and bundled revisions - A checkpoint laya loads for a later request gets the same options, warmup and device check through the Router's on_load hook; an evicted one drops out of /health; one that cannot be prepared is unloaded. /health names a checkpoint under `preparing` while that runs. - report.py: the parity section uses the records read() already filtered, so paired files no longer crash it. - /health finds the revision of a bundled checkpoint under the download repository. - Comments: drop a private spec id, headings repeated as comments, and a docstring restating its function.
# Conflicts: # README.md # src/models/laya/README.md
|
All three are fixed in 584680d, and main is merged in (7dc6b68). What I could run here that the review couldn't, on the M1 Pro:
|
…which laya-serve variables apply - A late checkpoint that cannot be prepared is unloaded and the checkpoints laya evicted for it are loaded again. - Only the variables laya-serve's app reads apply (LAYA_API_KEY); the worker warns about the launcher's ones. - A second app built on the same router replaces the first app's hooks. - paired.py: timed requests are not retried and failed pairs are counted; the summary shows each side's device at the end and flags a side that left its device or compiled path. - bench: wait_ready keeps polling on a malformed response; cleanup terminates every process before waiting.
…nd; bench spawn fixes
- Recipe: the first request for a late checkpoint took about 70 s with --compile on the M1 Pro, past the
frontend's 60 s backend timeout (504, then fine on retry); how to avoid it; memory with two checkpoints.
- /health: graphs compiled by a late checkpoint that is then unloaded are not reported as recompiles.
- bench_http.py, paired.py: a process that fails to start no longer leaves the other one running.
- bench_http.py: {device} and {model} in the spawn command, so --device/--model reach frontend.laya_mps.
- First-pass report compares against its own C2 run.
|
[P2] Preserve runs whose answer probes all fail Reviewed 543acef. Location: recipe/laya/bench/report.py:233–240. read_answers() creates a run entry only after a successful answers record, so this loop entirely omits configurations whose probes all produced answers_error. I reproduced this through bench_http.main() and report.main(), using the real W1/W4 workloads and mocked HTTP 200 responses without answers: C3 emitted two answer errors, but the report showed its latency and zero HTTP errors alongside a ‘4/4 questions within tolerance’ summary containing only healthy C2. Enumerate runs from environment/error records and represent failed probes explicitly rather than dropping the configuration. |
|
[P3] Bind revisions to the checkpoint actually loaded Reviewed 543acef. Location: src/models/laya/engine.py:78–81. The bundled-repo lookup fixes missing revisions, but it reads a mutable per-repository value on every health call. If English loads snapshot A and a later multilingual load downloads snapshot B from the same repository, recording B overwrites the dictionary entry and /health now reports B for English despite its weights still being A. Reproduced with Laya 0.3.20’s real Router and HTTP routes, replacing only Agent loading/inference and snapshot-download boundaries: the multilingual request returned 200 and English’s reported revision changed A→B while its actual loaded revision stayed A. Bind revision identity to each successful checkpoint/agent load. |
…nd what the warm numbers leave out Worker and benchmark scripts - /health keeps each checkpoint's revision from its own load: revisions are recorded per checkpoint (repo or repo/subfolder) and read once when the checkpoint has been prepared. - report.py: the parity section lists every run with an env record; a run whose answer probes all failed, or an answer missing a question, shows as failures instead of disappearing or raising. - paired.py --summarize: a run without an end record, a side without answers or a question missing on one side is reported instead of raising. Tests - Contract tests: answers match Laya run directly (choice, score, noul, combined, six questions), the port opens only after the warmup, the checkpoint stays loaded; LAYA_CONTRACT_DEVICE=mps and LAYA_CONTRACT_FLAGS run the same contract on the GPU, without and with the options. - /health is checked against Laya's real Router after random sequences of loads, failed loads and evictions. - Unit tests from the issue's requirements, and for behaviours a mutation run showed were executed but not asserted (late checkpoints get the GPU options before their warmup, logging, defaults, /health values). Recipe - What back-to-back latencies leave out, measured on the M1 Pro: a request after 2-5 s of idle takes 105-115 ms against 25 ms; a new input length costs about 15 ms once; with both options the footprint grows about 5 MB per input length seen, to 5.3 GB after all of them. - Checkpoint download: 97 s for 846 MB. Contract tests on the GPU.
|
Both fixed in f2b7c6a.
|
…cOS; memory growth is from compile - report.py: the parity section lists benchmark runs (they write a phase record before any answers) and runs with answers or failed probes; a profile_mps.py run, which never answers, no longer shows as failures. - The run-header test asserts the macOS machine probes only on macOS, and checks that they degrade to None elsewhere. - Recipe: the per-length memory growth comes from --compile, not from fp16 weights.
Purpose
Closes #3. Serves Laya on Apple Silicon (PyTorch MPS) behind
/v1/systemoneand measures it.Laid out like the Cua-S1 worker (#19): the HTTP worker is
src/frontend/laya_mps.py, started asPYTHONPATH=src python -m frontend.laya_mps --device mps --model english; the model-side code issrc/models/laya/{engine,optimize}.py; tests are intests/laya/; the recipe pins its environment inrecipe/laya/requirements-mps.txt. The worker runs laya-serve (laya[serve]==0.3.20) with its requesthandling unchanged and adds three optimizations and one fix:
--compile): one-question requests run the whole model throughtorch.compile; requests with several questions compile only the encoder, because Laya's decisionhead is slower compiled on MPS. 68-token request on an M1 Pro: ~46 → ~33 ms (−29%).
--weights fp16): keep the checkpoint's fp16 weights instead of Laya'sfp32 upcast on MPS (
act_headstays fp32). A further −14% on top of compile, and4.2 → about 3 GB of memory for a worker on its own, on the six benchmark inputs; with
--compileitgrows with the input lengths seen (see "What the warm numbers leave out").
multi-question requests. laya-serve answers
/healthbefore any forward pass, so the first requestafter it took 0.7–1.1 s; through the worker it takes 70–81 ms (62–78 ms with both options), still more
than a warm request because the GPU has been idle since the warmup. A checkpoint
Laya loads later (a request that names another model, or routes to it) is prepared the same way inside
its first request, and
/healthlists it underpreparingmeanwhile. With--compilethat requesttook about 70 s on the M1 Pro, past the frontend's 60 s backend timeout: through the frontend it got 504
and the same request sent again took 35 ms. The recipe says how to avoid it.
/healthtells the truth. It reports, read on every call, the device the model is actually on,dtypes, checkpoint and the revision the weights were loaded from. laya-serve echoes
LAYA_DEVICEand keeps reportingmpsafter Laya falls back to CPU;--require-devicemakes the worker exit at startup instead, andafter a fallback while serving the worker runs Laya's plain fp32 model and says so (
compile.activeturnsfalse).Together, 1 and 2 bring a 68-token one-question request to 0.62 of its time on the M1 Pro and 0.51 on an
M5, with every input faster and answers within 0.0031 of the unoptimized worker's (0.0054 on the M5);
results below.
recipe/laya/apple-silicon.md: setup,/health, frontend, tests and benchmark commands on a Mac.The six scripts behind the numbers below are in
recipe/laya/bench/. The result tables and the rawJSONL are release assets on my fork
(checksums in
recipe/laya/bench/README.md), so no data files sit in the repository.Test Plan
System1-Omni Version / Commit: based on
3062243(main after #2). Each measurement ran on the working treecommitted right after it: baseline and one-question compile as
2a47358, fp16 weights asd3d0817, alloptimizations together as
e10cf06. Each result file records the commit, versions, power and load.PYTHONPATH=src python -m pytest tests/laya(26).LAYA_CONTRACT=1 PYTHONPATH=src python -m pytest tests/laya(13 more: all decision types, 400/401/413/422, first request after ready within 2× warm p50).
Plain laya-serve fails exactly the two tests covering the readiness and device issues.
55cf4c4), on AC power, in two designs:(three for the configs that failed the gate after two) of 300 requests per input after 20 discarded,
configs interleaved in the same rounds. A result counts only if all its measured runs agree on p50
within 10%; no run was dropped. These started at a 1-minute load average ≤ 5 (the plan said ≤ 2, which
this machine never reached), so absolute numbers may be a little high.
each back to back in alternating order, 300 pairs per input after 20 discarded, two runs; the
statistic is the median per-pair ratio with a 95% bootstrap interval. Separate runs could not
resolve a 10% difference on this machine, so these had no load gate: they started at load 19–31
(fp16) and 11–15 (all together), with intervals within ±3% of the median.
Criteria were written down before each measured run.
ruff check/ruff formaton the Python files (CI runs Rust checks only).Test Result
All optimizations together
The worker with
--compile --weights fp16against the worker with neither,both running at once, every request sent to each back to back (
paired.py, 300 pairs per input). Declaredbefore measuring: a one-question median ratio ≤ 0.70 with the 95% interval below 0.75, no input slower,
answers within 1e-2, no graph compiled after ready. All held in both runs.
In absolute terms the 68-token request went from 55–59 ms to 35–36 ms in these runs (both workers share the
GPU and the machine was loaded; a worker on its own measured ~46 ms before and ~28 ms after, the latter in
feasibility runs only). Answers within 0.0031 of the
baseline's, no decision changes. Worker memory 3.5 GB → 2.8 GB in these runs, where two workers shared the
machine; on its own a worker uses 4.2 GB without the options and about 3 GB with them on these six inputs. Ready after 35–39 s
instead of 8–10 s (two fresh starts each). Separately, the worker's warmup before readiness brings the first
request after
/healthfrom 0.7–1.1 s (plain laya-serve) to 70–81 ms, or 62–78 ms with both options. On an M5, 2 of 23 fresh starts stillgave a slow first request (327 and 409 ms, against 21–36 ms otherwise); not explained yet.
What the warm numbers leave out
Everything above is for requests sent back to back on six fixed inputs. Measured afterwards on the M1 Pro:
one-question request takes 25 ms back to back with the options and 40 ms without; after 0.2–1 s of idle
it took about 50 ms and 60–68 ms, after 2–5 s 105–115 ms and 114–127 ms. An agent that asks once every
few seconds sees those numbers, and there the options are worth about 10%, not 37%. This is also why the
first request after ready is slower than a warm one.
with the options, 6 ms without. Not a recompile.
--compile, PyTorch keeps host memory for every input lengththe compiled model has run, about 5 MB each (fp16 weights alone add little): 2.8 → 3.3 GB after 100 new
lengths, 5.3 GB after all 477, against 4.0 GB without the options.
torch.mps.empty_cache()releasesmost of it and those lengths are cold again.
times above are with it cached.
No regression from the review rounds
The worker was restructured several times during review, so the measured code (
b00b285) was comparedwith the reviewed one on the M1 Pro (AC power, load 7.6–8.1, 300 pairs per input, both workers alive):
Old and new gave identical answers. The last column reproduces the paired results above within 0.03.
bench_http.pyatf2b7c6a: ready after 10.6 s without the options and 39.3 s with them; first requestafter ready 86 and 72 ms; footprint after warmup 4.2 and 2.8 GB.
What each part contributes
profile_mps.py) showed short requests bound by afixed per-request cost from ~4,000 small ops; compiling cut the 68-token request from ~46 to ~33 ms
(−29%) in two measured runs against the uncompiled worker, with multi-question requests unchanged.
path, two paired runs gave a 0.855–0.857 median ratio for the 68-token request. On the M1 Pro it did
not help without compiling: fp16 GEMMs are ~25% faster there, but the fixed cost hid that. On an M5 it
speeds up every input on its own (below).
slower: Laya's decision head (two
nn.TransformerEncoderLayerwith a padding mask) loses PyTorch's fusedpath when compiled and took ~55 ms instead of 3–12 ms. Compiling only the encoder made them ~14%
faster in a same-process probe; the combined run above measures it together with the rest.
On an M5
@twu3202 ran the recipe and
paired.pyon an Apple M5 (10-core GPU, 32 GB, macOS 26.5.2) ate10cf06(comment): all tests pass,
and against the worker with neither option the median ratios were 0.51–0.53 for one-question requests at
47–68 tokens, 0.30–0.33 at 198–484 tokens, 0.37 for three questions and 0.60 for six. There fp16 weights
alone speed up every input (0.43–0.81). The declared criteria hold on that machine too.
Tried and dropped
slower, one-question slower than dynamic.
reference_compile: transformers 5.x no longer reads it (0 graphs, same latency).1e-2 tolerance). Keeping encoder layer 3 in fp32 fixed that input but not held-out ones.
through MPS (a 28-layer chain of Laya's four GEMMs alone takes 28.5 ms at 68 tokens, fp32, against
~33 ms for the whole request). MLX 0.32.3 was within 3% of PyTorch MPS or slower on 22 of 24 GEMM
shape/precision cells (faster by 5% and 10% on two fp16 cells) and on the chain (32.7 vs 28.5 ms fp32, 25.2 vs 21.0 ms fp16), measured in one
process at load ~9. A port would not buy speed on this hardware.
This is about speed for the 0.4B Laya, which fits in memory; a larger model that has to be quantized
to fit may decide differently ([RFC] Support emerging System 1 models, with CLM as the second model after LAYA #9).
torch.mps.compile_shader, 8×8 simdgroup matrices, fp16 weights, fp32accumulate, exact against fp32): timed alone it beat
mmon three of Laya's four GEMM shapes, but in a28-layer chain at 68 tokens the best variant came to 0.997 of the fp16
mmchain and the others to1.05–1.29, with or without repacking the weights per column tile. The bar set beforehand was 0.90.
Split-K through
bmmdid not help either (−4% at best). At 8 rows the chain still takes ~12 ms, of whichdispatch is ~2 ms; the rest scales with the weights walked and is the same for
mmand the customkernel. So what remains at short inputs is not MPSGraph overhead but the cost of walking the weights,
and
--weights fp16already runs them at the checkpoint's own precision. Measured at load 17–24, soratios only.
Baseline
Warm p50 in ms at concurrency 1, range over the measured runs.
*means the runs differ by more than10%, so that number is indicative only.
over in-process. CPU numbers are unstable on this machine (other applications were running).
the 68-token request and latency grows 4x.
background load rather than the frontend. Sending every request directly and through the frontend
back to back (now
paired.py --a-url/--b-url; 120 pairs per input): the frontend adds +0.1 to +0.5 msat the median, and 1 of 720 pairs differed by more than 10 ms.
after ready: 697–1135 ms for laya-serve, 75–81 ms for the worker.
2.0 GB on CPU; the frontend uses 3 MB.
Output parity
Every configuration and run against CPU fp32: 330/330 questions with the same decision and |Δp|
within 1e-3 (1e-2 where laya autocasts to fp16).
Follow-ups (not in this PR)
(20% of the GEMM rows in the 3-question input are padding) and runs the decision head eagerly. A
block-diagonal mask with per-row positions is exact and would make every request a single row on the
compiled path; estimated −12 to −17% for the 3-question input, nothing for one question.
0.5 s held a request after any pause at about 45 ms for 8% GPU load, and every 0.2 s at 31–41 ms for 18%;
a tiny GPU operation as heartbeat did nothing. Not in this PR: the worker reuses laya-serve's request
handling, whose inference thread it cannot reach, and driving MPS from a second thread aborted the process.
The slow first requests on the M5 (327 and 409 ms in 2 of 23 starts) are probably the same effect.
Not verified: M-series chips other than the M1 Pro and M5.
Self-review
Done at
005a874, after laying the files out like #19 (worker insrc/frontend/, model code insrc/models/laya/, tests intests/laya/, flags instead of environment variables, pinned requirements,six benchmark scripts instead of ten, no result files in the repository). The diff review found and fixed: warmup and
/healthcovering only one of several loadedmodels; the revision read from the cache ref instead of the loaded snapshot; three error-path bugs in
bench_http.py; fp16 weights described as suitable for CPU (same answers there, but 334 ms instead of138 ms; now documented and warned about); and
/healthdescribing the model as it was at startup, so alater fallback to the CPU stayed invisible (now read on every call).
The recipe was then checked line by line against the measurements and a real worker. Corrected: ready
time with both options (35–39 s, not the 20–30 s of the earlier one-question-only compile), the memory
baseline, which configuration the first-request numbers belong to, the uv route to Python 3.12, and the
first-pass report command, which failed once paired results were in the directory.
--require-devicewas exercised on a real worker with MPS reported unavailable: exit code 1 with the documented message.
The second review's findings (checkpoints loaded while serving, the report's parity section on paired
files, revisions of bundled checkpoints) are fixed in
584680d. Checked on the M1 Pro with the realmultilingual checkpoint: loaded by an explicit
modeland by a German state with--compile --weights fp16,both checkpoints described in
/healthwith revision55cf4c4; and with--require-deviceand the MPSlimit lowered so the late checkpoint falls back to CPU, it is unloaded, that request gets 500 and the
preloaded checkpoint keeps serving.
A further pass over the whole diff, each finding reproduced by a test before fixing (
87c920f): a latecheckpoint that fails preparation no longer costs the checkpoint Laya evicted for it; the docs said
laya-serve's environment variables still apply, but only
LAYA_API_KEYdoes, and the worker now warnsabout the others;
paired.pytimed requests with retries and its summary did not show a side that fellback to CPU; two cleanup and polling faults in the benchmark scripts. The evicted-checkpoint fix is covered
by a unit test on Laya's real Router with stub agents only: it needs three checkpoints to trigger.
A second pass, again reproduced before fixing (
1450884): the late-load cost above, which the recipe haddescribed as seconds; benchmark scripts leaving a worker running when the next process failed to start, and
not passing
--model/--deviceto a spawnedfrontend.laya_mps;/healthreporting recompiles foreverafter a late checkpoint failed its warmup; and the first-pass report having no parity reference.
The third review's findings (runs whose answer probes all failed vanishing from the report, and a
checkpoint's revision changing when another one is downloaded from the same repository) are fixed in
f2b7c6a, together with four more crashes of the result readers on damaged files that a self-audit found.The contract tests now also run against a worker on the GPU (
LAYA_CONTRACT_DEVICE=mps, with and withoutLAYA_CONTRACT_FLAGS="--compile --weights fp16") and compare its answers with Laya run directly; 20 passedin each of the three configurations.
A further self-review fixed in
668f331: a profile run's results no longer show as parity failures inreport.py, the run-header test no longer assumes macOS tools, and the memory growth above is attributedto
--compilerather than to the fp16 weights.Checks at
668f331: 72 unit tests, 92 withLAYA_CONTRACT=1,ruff checkandruff format --checkonthe Python files.
cargo fmt --all --check,cargo clippy --workspace --locked --all-targets -- -D warningsand
cargo test --workspace --lockedpassed at87c920f, after merging main; no Rust file is changed by this PR. Laya's fallback to CPUwas triggered on the real GPU by lowering PyTorch's MPS memory limit (
PYTORCH_MPS_HIGH_WATERMARK_RATIO)and sending a 64-question request. The request still returned 200 and
/healthswitched todevice: cpu,device_mismatch: true. The first run of this showed a worker with compile and fp16 weights degradingbadly afterwards (336 s for the request that fell back, a 28 s recompile, then 370–550 ms per request),
so the worker now runs Laya's fp32 model uncompiled whenever the inputs are on the CPU. Rerun with both
options, fp16 only and compile only: 32–40 s for the request that fell back, then 140–270 ms, no graph
recompiled, answers within 0.0002 of those before the fallback.