Skip to content

[Feat] Served Laya as a decision model (--model laya-served) - #35

Open
cacheline999 wants to merge 14 commits into
ThinkFlowLab:mainfrom
cacheline999:served-laya
Open

cacheline999 wants to merge 14 commits into
ThinkFlowLab:mainfrom
cacheline999:served-laya

Conversation

@cacheline999

Copy link
Copy Markdown
Contributor

Why

Closes #20. --model laya loads Laya in every process, so each CLI run and each MCP call pays the load and nothing shares a warm model. system1-omni can now keep Laya warm in a worker (ThinkFlowLab/system1-omni#30), and this PR lets agents use it.

How

s1a/decision_models/served.py is the new backend. Open it first.

  • ServedLayaClient posts to /v1/systemone and owns the connection handling:

    • one deadline per decision (LAYA_SERVED_TIMEOUT_S, default 5 s);
    • one retry on a dropped connection, a 502 or a 504;
    • one retry after Retry-After on a 503;
    • one X-Request-Id per decision, reused on the retry;
    • every other status mapped to MODEL_CALL_FAILED or MODEL_SERVICE_CONFIG_ERROR.
  • ServedLayaModel builds the same body as in-process Laya with laya_question(). It shares the full-window check with LayaModel, now check_window() in laya.py.

  • It also records who answered: Decision.model is <checkpoint>@<revision>. The details come from the response's served_by when a server sends one. Otherwise they come from the worker's /health, re-read after 30 s, since the worker reports there when Laya falls back to the CPU. Plain laya-serve reports less, so the record keeps routing.repo for it.

  • The contract is written down in docs/served-laya.md, as design and trade-offs. There are two OpenAPI 3.1 specs in docs/api/:

    • laya-systemone.current.openapi.yaml is today's servers, checked against recorded traffic;
    • laya-systemone.openapi.yaml is the target, with problem+json errors, served_by, Server-Timing and /readyz.

    The gap between the two belongs to system1-omni; section 4 of the doc lists where each part would go.

What

  • --model laya-served works on every agent, on decide, on the rails and in MCP decide. tests/test_served_laya_registration.py fails if a new list offers laya without it.
  • Environment variables:
    • LAYA_SERVED_URL is required: the worker, omni-jev or laya-serve.
    • Optional: LAYA_SERVED_MODEL, LAYA_SERVED_API_KEY, LAYA_SERVED_TIMEOUT_S, LAYA_SERVED_MAX_LEN.
    • No cloud key is read.
  • Tool-front ticks now keep model for every decision model. A served model's ticks also keep served_by, url, request_id and server_timing.
  • Docs: decision-models.md, configuration.md, .env.example, the CHANGELOG, and a run section in served-laya.md.
  • The dev extra gains jsonschema, pyyaml and referencing for the spec check (pyproject.toml, uv.lock, CONTRIBUTING). They were already in the environment through mcp and openjiuwen.
  • The branch also moves the laya extra to 0.3.20, the version the worker runs.

The worker features used here (warm before listening, a live /health with checkpoint, revision and device) come from ThinkFlowLab/system1-omni#30, which is still open. Against plain laya-serve everything works, with the checkpoint as the only identity. The run section links the recipe on the PR branch until it merges.

Results

ticket_router on an M1 Pro, 3 seeds × 30 tickets per configuration. Laya is checkpoint 55cf4c4 with laya 0.3.20. The worker is system1-omni 3d6cb57 (the head of ThinkFlowLab/system1-omni#30), started with --compile --weights fp16. Details, commands and every decision are in evals/ticket_router/SERVED_LAYA.md.

correct p50 ms p95 ms
in process (--model laya, MPS, fp32) 63/90 88 101
worker, direct 63/90 72 88
worker through omni-jev 63/90 72 76

All three routed every one of the 90 decisions the same. In-process and worker probabilities differ by at most 0.001. In-process Laya is slower here because it runs fp32 and uncompiled, not because of HTTP.

Open questions

  • The name laya-served, chosen so records tell served runs from in-process ones.
  • The target interface in docs/api/laya-systemone.openapi.yaml. If it looks right, the worker-side parts belong in system1-omni.

Verification

  • uv run ruff format --check . && uv run ruff check . && uv run ty check: all pass.

  • uv run pytest -q and scripts/smoke.sh: 586 passed and 43 skipped with the laya extra installed; smoke: ok. With torch, laya and transformers blocked from import, the new tests pass (113 passed, 2 skipped), as on CI's core install.

  • CHANGELOG.md and the docs say what the code does now.

  • New tests run over httpx.MockTransport:

    • the shared DecisionModelContract;
    • each row of the error table;
    • the deadline across a retry and the window check;
    • identity, including the 30 s refresh with a fake clock.
  • tests/data/served_laya/ holds 27 responses recorded from the worker, from omni-jev (with the worker up and down) and from laya-serve. They validate against the as-is spec.

  • Live on the same M1 Pro and worker:

    check result
    worker ready after 37 s
    decide via the CLI and over MCP (stdio) answered
    injection_guard rail precision 0.9, recall 0.9
    ticket_router, direct and through omni-jev 21/30
    worker stopped behind omni-jev fails with the 502 after one retry
    plain laya-serve answers, with identity reduced to the checkpoint
    CPU worker (--device cpu) with LAYA_API_KEY set 401 without LAYA_SERVED_API_KEY; with it, 21/30 and ticks show cpu, float32, not compiled

laya 0.3.9 builds the encoder under transformers' no_init_weights inside
laya.load and loads the checkpoint with strict=True, so the
without_weight_init() wrapper from ThinkFlowLab#22 and its test stubs are no longer
needed. The lock moves laya from 0.3.5 to 0.3.9, the first release with
the skip; later releases change MPS precision and are left for their
own bump.

The system_one config-error test pins the version lookup, so it passes
with the laya extra installed as well as on the core install.

Closes ThinkFlowLab#28
system1-omni's Laya worker pins laya[serve]==0.3.20. Comparing in-process
Laya with the served one (ThinkFlowLab#20) needs the same library on both sides, so the
in-process extra moves from 0.3.9 to 0.3.20. 0.3.20 autocasts to fp16 on
MPS for requests with five or more questions; CPU runs and smaller requests
are unchanged.
docs/served-laya.md designs the served Laya decision model: requirements,
the target interface, what today's server lacks and where each part gets
built, the client (--model laya-served), error handling and trade-offs.

docs/api/laya-systemone.openapi.yaml is the target interface as OpenAPI
3.1 (RFC 9457 errors with stable codes, request ids, Server-Timing,
/livez and /readyz, 503 with Retry-After, served_by per response); each
item is marked implemented or planned. laya-systemone.current.openapi.yaml
specifies today's server (laya-serve 0.3.20 behind the system1-omni
worker and frontend), checked against traffic from a running worker.
ServedLayaModel asks a Laya served over HTTP (system1-omni's worker, its
omni-jev frontend, or plain laya-serve) through POST /v1/systemone, with
the body the in-process model builds (laya_question). The client holds one
deadline per decision, retries once on a dropped connection, 502 or 504,
and after Retry-After on 503, sends one X-Request-Id per decision, and maps
every other status to MODEL_CALL_FAILED or MODEL_SERVICE_CONFIG_ERROR.

Decision.model is the served checkpoint and revision, read from the
response's served_by when a server sends it, else from /health (refreshed
after 30 s, since the worker reports a CPU fallback there), else from
routing.repo against plain laya-serve. The window check is shared with
LayaModel (check_window).

Tests run over httpx.MockTransport: the shared contract, the error and
retry mapping, identity, and responses recorded from a real worker, the
frontend and laya-serve, which are checked against the as-implemented
OpenAPI spec (jsonschema, pyyaml and referencing join the dev extra).
--model laya-served on every agent, on decide, on the rails and in MCP
decide. A test finds any name list, match or Literal that has laya without
laya-served, so a new front cannot miss it.
…nkFlowLab#20)

A tool-front tick keeps Decision.model as model and, from a served model,
served_by, url, request_id and server_timing (Decision.provenance), so a
run's artifacts show the checkpoint, revision and device of each step.
decision-models.md and configuration.md list laya-served and its
variables; served-laya.md gains a run section (start the worker once, use
it from the CLI and MCP) and matches system1-omni#30 on /health freshness
and GPU-only worker options.
A refused connection after the retry now says the worker listens only once
warm and points to system1-omni's recipe; a first /health read that fails
no longer logs about a previous reading that does not exist. The --model
help of every front names laya-served.
90 decisions per configuration on an M1 Pro: in-process Laya, the
system1-omni worker (compile + fp16) and the same worker behind omni-jev
routed every (seed, ticket) pair the same and scored 63/90 each; p50 101,
77 and 82 ms. compare_served.py builds the table from the job dirs and
keeps every decision in served_laya_records.json.
The worker now starts as `python -m frontend.laya_mps --compile --weights
fp16` and its /health reports `compile.enabled` instead of `compile.mode`.
served_by records `compiled` (true/false); both specs, the run section
and the not-up message follow. Worker fixtures were recorded again at
3d6cb57; the other responses came out byte for byte the same.

The ticket-router comparison was rerun at 3d6cb57: 63/90 in every
configuration, all 90 decisions routed the same, p50 88 ms in process and
72 ms served, direct and through omni-jev.
…FlowLab#20)

The registration check read s1a's sources with the platform encoding,
cp1252 on Windows, and reported paths with backslashes there. Every new
text read and write now names utf-8, and paths are reported in POSIX form.
…stopped episodes (ThinkFlowLab#20)

The identity refresh every 30 s ran before the decision's deadline began,
so a stalled /health could stretch one decision to LAYA_SERVED_TIMEOUT_S
plus 2 s. The refresh now shares the decision's deadline and takes at most
half of what is left.

compare_served.py paired the batch's ticket ids with the ticks, so a job
with an episode that stopped early raised in zip(). It now pairs the
processed routes with the ticks and checks that each pair agrees.
…rors (ThinkFlowLab#20)

The not-up error quoted the M1 Pro ready time with --compile and fp16,
which misleads on a CPU worker or another Mac; it now says the worker may
still be starting and points to the recipe. The timeout comment gives its
reason instead of a measured latency range, and the browser comment names
served Laya as the HTTP one. The run section quotes the real message.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Help wanted] Collaborate on system1-omni Laya integration and Mac evaluation

1 participant