Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 26 additions & 2 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
[workspace]
members = ["src/frontend", "src/models/cua_s1/native", "src/models/laya"]
members = ["src/frontend", "src/models/cua_s1/native", "src/models/qwen3_5/native", "src/models/open_jev/native", "src/models/laya"]
resolver = "3"
10 changes: 6 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Documentation: <https://thinkflowlab.github.io/system1-omni/>

A community-maintained inference engine for prefill-only System1-Omni models, designed around a Rust frontend, model-owned execution, and high-performance CUDA and Metal backends.

The Rust frontend forwards requests to a separately running model worker. The Cua-S1 4B 0.2 `text` adapter has a native worker with CUDA kernels in this repository; other in-repository model engines and GPU backends are not implemented yet.
The Rust frontend forwards requests to a separately running model worker. The Cua-S1 4B 0.2 `text` adapter and Open-Jev-27B-v1.1 have native workers using shared CUDA kernels in this repository.

## Run the frontend

Expand Down Expand Up @@ -49,7 +49,7 @@ Implementation code lives under `src/`; recipes and documentation stay at the re
| [`recipe/`](recipe/) | Model setup instructions, launch commands, configuration examples, and example requests. |
| [`docs/`](docs/) | Project documentation and architecture assets. |

The frontend, Cua-S1 native worker and Laya checkpoint reader are Cargo workspace members. The other model and backend directories currently document planned work; they do not prescribe process boundaries.
The frontend, both native workers, their shared Qwen3.5/3.8 prefill implementation and the Laya checkpoint reader are Cargo workspace members. The other model and backend directories currently document planned work; they do not prescribe process boundaries.

## Supported models

Expand All @@ -59,14 +59,16 @@ LAYA can run as an external Python worker for text requests; its in-repository m
| --- | --- |
| LAYA | [External worker](recipe/laya/README.md); [CPU checkpoint reader](src/models/laya/README.md); model execution planned |
| Cua-S1 4B 0.2 (`text` adapter) | [Python worker](recipe/cua_s1/text.md); [native worker](recipe/cua_s1/native.md), CUDA, run on sm_89 |
| Open-Jev-27B-v1.1 | [Native Rust/CUDA worker](recipe/open_jev/native.md); eager independent text candidates; [L20X validation](recipe/open_jev/validation.md) |

CUDA and Metal coverage will be documented per model as implementations are added and validated.

## Benchmarks

See the [GPU serving benchmark](benchmarks/README.md) for request replay,
output-fidelity checks, and the CUDA comparison protocol. GPU performance
measurements are pending.
output-fidelity checks, and the CUDA comparison protocol. The
[Open-Jev L20X results](recipe/open_jev/validation.md) cover 74 single-candidate
requests and a matched comparison with OpenJev-Fast.

## Stay Tuned with Us

Expand Down
2 changes: 2 additions & 0 deletions recipe/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,8 @@
the worker and connect the Rust frontend.
- [Cua-S1 4B 0.2 native text worker](cua_s1/native.md): build the CUDA library and
the Rust worker, export the merged weights and start the worker.
- [Open-Jev-27B-v1.1 native text worker](open_jev/native.md): export the merged
text backbone and trained decision head, then serve with Rust and CUDA.

Recipes contain setup, launch commands and examples. Reusable implementation code
belongs under `src/`.
4 changes: 2 additions & 2 deletions recipe/cua_s1/native.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ each exact prompt length warms the GEMM plans and captures the forward pass;
later requests replay it with freshly uploaded token ids. At most eight lengths
are cached. Growing the scratch allocation clears the captures before freeing
their buffers. Capture adds first-use latency; leave the variable unset to use
the eager control. Rebuild both the worker and CUDA library together (ABI 3).
the eager control. Rebuild both the worker and CUDA library together (ABI 4).
If capture fails, the worker returns the completed eager result and disables
Graph capture/replay for its remaining lifetime, logging the failure to stderr.

Expand All @@ -39,5 +39,5 @@ The request tests need no GPU; the kernel tests compare attention and the chunke
```sh
cargo test -p omni-cua-s1-native
CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \
cargo test --release -p omni-cua-s1-native --test kernels -- --ignored
cargo test --release -p omni-qwen3-5-native --test kernels -- --ignored
```
27 changes: 27 additions & 0 deletions recipe/open_jev/example-request.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
{
"state": "A customer says: I was charged twice for one order. The service is working normally. There is no sign of unauthorized access.",
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this issue?",
"criteria": {
"billing": "Problems with charges, invoices, refunds or payments.",
"security": "Unauthorized access or account compromise.",
"technical": "Service unavailable or a software malfunction."
}
},
"refund_review": {
"type": "noul",
"instructions": "Does this message describe a duplicate charge?"
},
"urgency": {
"type": "score",
"instructions": "Assess the urgency using only the given evidence.",
"criteria": [
"Routine: no service disruption or active security compromise is reported.",
"Urgent: an ongoing service disruption is reported.",
"Critical: active unauthorized access is reported."
]
}
}
}
70 changes: 70 additions & 0 deletions recipe/open_jev/export_merged.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
"""Export the pinned Open-Jev-27B-v1.1 text backbone and scalar head on CPU.

Use the reference environment documented in native.md. This preparation step
needs about 110 GB of host RAM and 52 GB of output storage, without a GPU.
"""

import argparse
import json
from pathlib import Path

import torch
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoTokenizer

BASE_REVISION = "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0"
CHECKPOINT_REVISION = "28cf73067d5b337860bbef3c85b8b82ba8730956"


def main():
parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
parser.add_argument("--base", required=True, type=Path)
parser.add_argument("--checkpoint", required=True, type=Path)
parser.add_argument("--out", required=True, type=Path)
parser.add_argument("--max-length", type=int, default=4096)
args = parser.parse_args()
config = json.loads((args.checkpoint / "model.json").read_text())
if config["model_id"] != "Qwen/Qwen3.8-27B" or config["revision"] != BASE_REVISION:
raise ValueError("expected Open-Jev-27B-v1.1's pinned base")
if args.out.exists():
raise ValueError("output already exists; choose a new export directory")
if not 1 <= args.max_length <= 16384:
raise ValueError("max length must be within 1..=16384")
temperature = json.loads((args.checkpoint / "temperature.json").read_text())["temperature"]
tokenizer = AutoTokenizer.from_pretrained(args.base, local_files_only=True)
marker = "\x00OMNI_OPEN_JEV\x00"
chat = tokenizer.apply_chat_template(
[{"role": "user", "content": marker}], tokenize=False,
add_generation_prompt=True, enable_thinking=False,
)
if chat.count(marker) != 1:
raise ValueError("expected a single-user text chat template")
prefix, suffix = chat.split(marker)
head = torch.load(args.checkpoint / "head.pt", map_location="cpu", weights_only=True)
if head["weight"].shape != (1, 5120) or head["bias"].shape != (1,):
raise ValueError("expected a 5120-wide trained scalar head")
if not all(torch.isfinite(v).all() for v in head.values()):
raise ValueError("non-finite scalar head")
full = AutoModelForImageTextToText.from_pretrained(
args.base, torch_dtype=torch.bfloat16, device_map={"": "cpu"},
attn_implementation="sdpa", local_files_only=True,
)
backbone = full.model.language_model
del full
backbone = PeftModel.from_pretrained(backbone, args.checkpoint / "adapter")
backbone = backbone.merge_and_unload(safe_merge=True)
backbone.save_pretrained(args.out, max_shard_size="5GB")
tokenizer.save_pretrained(args.out)
# Written last: the native worker refuses incomplete exports or plain base weights.
(args.out / "open_jev_export.json").write_text(json.dumps({
"format": "open-jev-text-merged/1",
"model_id": config["model_id"], "base_revision": BASE_REVISION,
"checkpoint_revision": CHECKPOINT_REVISION, "temperature": temperature,
"max_length": args.max_length, "chat_prefix": prefix, "chat_suffix": suffix,
"head_weight": head["weight"].float().reshape(-1).tolist(),
"head_bias": head["bias"].float().item(),
}, allow_nan=False) + "\n")


if __name__ == "__main__":
main()
125 changes: 125 additions & 0 deletions recipe/open_jev/native.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,125 @@
# Open-Jev-27B-v1.1 native text worker

The worker owns request compilation, tokenization, candidate scoring and typed
responses in Rust. It uses the native CUDA prefill implementation introduced in
[PR #19](https://github.com/ThinkFlowLab/system1-omni/pull/19), shared with Cua-S1
under [`src/models/qwen3_5/native/`](../../src/models/qwen3_5/native/).
Python is required only to prepare the merged checkpoint.

It supports `choice` (1–255 candidates), `score` (2–10 levels), and `noul`
(yes/no). Each candidate has an independent prompt; the last hidden state goes
through Open-Jev's trained FP32 scalar head. Noul uses logits `[0, score]`.
The saved calibration temperature is applied before normalizing each complete
question. There is no autoregressive generation. Structured state and descriptions
use Open-Jev's sorted JSON rendering; question and candidate order is preserved.

## Prepare the checkpoint

Use the reference dependencies from
[Open-Jev @ 3308a15](https://github.com/Zefan-Cai/Open-Jev/tree/3308a15ccd7eea1df7a37d6ddc39b023b801ba16):
PyTorch 2.8 or newer, Transformers 5.10.2, PEFT 0.19.1, Accelerate 1.13.0,
and safetensors. An optional `kernels` installation must be compatible with that
Transformers release. Run these commands from the repository root:

```sh
hf download Qwen/Qwen3.8-27B \
--revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 \
--local-dir weights/Qwen3.8-27B
hf download ZefanCai/Open-Jev-27B-v1.1 \
--revision 28cf73067d5b337860bbef3c85b8b82ba8730956 \
--include 'package/checkpoint/*' --local-dir weights/Open-Jev-27B-v1.1
CUDA_VISIBLE_DEVICES='' python recipe/open_jev/export_merged.py \
--base weights/Qwen3.8-27B \
--checkpoint weights/Open-Jev-27B-v1.1/package/checkpoint \
--out weights/open-jev-27b-merged
```

CPU export needs roughly 110 GB of RAM and 52 GB of output storage. It merges
LoRA in BF16 and saves the trained head, temperature, and single-user chat
template in `open_jev_export.json`. The worker refuses a plain base checkpoint
or an incomplete export. The saved limit defaults to 4096 tokens per candidate;
`--max-length` may raise it to 16384. Oversize prompts fail before inference.

## Build and serve

The CUDA kernels require compute capability 8.0 or newer. The current build
target below is Ada (`89`); pass your GPU's compute capability explicitly.
The CUDA shared library and both Rust workers must be rebuilt together because
the gated-attention entry point updates the library ABI to version 4 alongside
the shared CUDA Graph entry points.

```sh
src/backends/cuda/qwen3_5/build.sh target/release 89
cargo build --release --locked -p omni-open-jev-native -p omni-jev
OPEN_JEV_MODEL=weights/open-jev-27b-merged \
target/release/omni-open-jev-native
```

`OPEN_JEV_HOST` and `OPEN_JEV_PORT` default to `127.0.0.1` and `8000`.
`OPEN_JEV_CUDA_LIB` overrides the default library next to the executable.
The worker loads all text weights onto visible CUDA device 0, performs a real
warmup inference, then exposes `/health` and `/v1/systemone`.
Use a reservation before any GPU command on hosts with a GPU scheduler.

In another terminal, start the existing Rust frontend:

```sh
OMNI_JEV_BIND=127.0.0.1:8080 OMNI_JEV_BACKEND_URL=http://127.0.0.1:8000 \
target/release/omni-jev
curl http://127.0.0.1:8080/v1/systemone \
-H 'Content-Type: application/json' --data-binary @recipe/open_jev/example-request.json
```

The worker accepts the model's base name `Qwen/Qwen3.8-27B`, `open-jev`,
`jev-latest`, and `open-jev-27b-v1.1`; the response model is the base name,
matching Open-Jev. Error wording and metadata differ from the reference service.
Requests are bounded to 4 MiB, 4096 questions, and 65536 candidate sequences.

## Validation and optimization scope

Tests and fixtures live in the repository-level `tests/` tree: Open-Jev's typed
contract and tokenizer cases are in
[`tests/open_jev/`](../../tests/open_jev/), and shared Qwen JSON, configuration
and CUDA reference tests are in [`tests/qwen3_5/`](../../tests/qwen3_5/).
The default suites below run on CPU without downloading model weights:

```sh
cargo test --locked -p omni-open-jev-native -p omni-qwen3-5-native
cargo test --locked -p omni-jev --test frontend
```

The frontend mock-worker API coverage is tracked in
[issue #46](https://github.com/ThinkFlowLab/system1-omni/issues/46) and
[PR #58](https://github.com/ThinkFlowLab/system1-omni/pull/58). Checkpoint tokenizer
and CUDA kernel tests are opt-in; the latter require a GPU reservation:

```sh
# Inside a GPU reservation, after building the library:
CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \
cargo test --release --locked -p omni-qwen3-5-native --test kernels -- --ignored
```

CPU golden fixtures come from Open-Jev's request compiler and response formatter
at the revision above. Kernel tests compare attention and Gated DeltaNet with
float64 references and require exact BF16 equality between fused attention gating
and a separate gate pass. Residual RMSNorm and packed SiLU are checked against
rounded references, including odd widths and unaligned pointers. The shared
kernel tests retain PR #19's
`CUA_S1_CUDA_LIB` environment variable.

The worker reuses PR #19's fused norm, activation, QK/RoPE and chunked Gated
DeltaNet operations. Attention's sigmoid gate is fused into its output epilogue,
preserving both BF16 rounding points and removing one launch and one output
read/write pass per full-attention layer (16 layers for this model). Residual
RMSNorm keeps thread values in registers at widths 2560/5120. MLP SiLU uses
16-byte BF16 loads/stores when width, stride and pointers permit it, retaining
both BF16 rounding points; other layouts use the scalar path.

This recipe leaves `CUA_S1_GRAPH` unset and runs one eager forward pass per
candidate. The shared backend retains Cua-S1's opt-in CUDA Graph path, but
Open-Jev graph replay remains unvalidated. Prefix sharing, GEMM autotuning,
quantization and multimodal inference are not implemented. The
[L20X validation](validation.md) reports full-checkpoint results for 74
single-candidate requests, including probability differences and timing
variability. It does not establish general accuracy parity or a speedup over
OpenJev-Fast; the author's B300 results use different hardware and workloads.
Loading
Loading