Documentation: https://thinkflowlab.github.io/system1-omni/
A community-maintained inference engine for prefill-only System1-Omni models, designed around a Rust frontend, model-owned execution, and high-performance CUDA and Metal backends.
The Rust frontend forwards requests to a separately running model worker. The Cua-S1 4B 0.2 text adapter has a native worker with CUDA kernels in this repository; other in-repository model engines and GPU backends are not implemented yet.
From the repository root, with stable Rust installed:
cargo build --release --locked
OMNI_JEV_BIND=127.0.0.1:8080 \
OMNI_JEV_BACKEND_URL=http://127.0.0.1:8000 \
./target/release/omni-jevStart the worker separately. See the frontend documentation for the HTTP interface and configuration, or the Laya recipe for a CPU text worker and response checks.
Share serving infrastructure; let each model own its execution.
| Layer | Responsibility |
|---|---|
| Rust frontend | API, request lifecycle, and response delivery through a small engine interface. |
| System1-Omni models | Model-specific preprocessing and postprocessing, batching, state, execution, and kernel selection. |
| CUDA backend | High-performance GPU operations for NVIDIA GPUs. |
| Metal backend | High-performance GPU operations for Apple GPUs. |
Each model owns its complete request-to-result path. Shared utilities stay minimal and are extracted when implementations need the same functionality. Backends can optimize for their hardware without requiring identical internal implementations.
Implementation code lives under src/; recipes and documentation stay at the repository root.
| Directory | Responsibility |
|---|---|
src/frontend/ |
Rust serving code, Python worker adapters, and the small engine interface. |
src/models/ |
Model implementations, one directory per model: preprocessing, batching, state, execution, and output processing. |
src/backends/cuda/ |
NVIDIA GPU operations and kernel integration. |
src/backends/metal/ |
Apple GPU operations and kernel integration. |
recipe/ |
Model setup instructions, launch commands, configuration examples, and example requests. |
docs/ |
Project documentation and architecture assets. |
The frontend and the Cua-S1 native worker are Cargo workspace members. The other model and backend directories currently document planned work; they do not prescribe process boundaries.
LAYA can run as an external Python worker for text requests; its in-repository model engine is still planned. The Cua-S1 4B 0.2 text adapter runs as a Python worker or as a native worker on CUDA:
| Model | Status |
|---|---|
| LAYA | External worker; model engine planned |
Cua-S1 4B 0.2 (text adapter) |
Python worker; native worker, CUDA, run on sm_89 |
CUDA and Metal coverage will be documented per model as implementations are added and validated.
See the GPU serving benchmark for request replay, output-fidelity checks, and the CUDA comparison protocol. GPU performance measurements are pending.
If you find system1-omni useful, give us a star on GitHub to support the project and help others discover it!
