Skip to content

Latest commit

 

History

96 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

System1-Omni

Documentation: https://thinkflowlab.github.io/system1-omni/

A community-maintained inference engine for prefill-only System1-Omni models, designed around a Rust frontend, model-owned execution, and high-performance CUDA and Metal backends.

The Rust frontend forwards requests to a separately running model worker. The Cua-S1 4B 0.2 text adapter has a native worker with CUDA kernels in this repository; other in-repository model engines and GPU backends are not implemented yet.

Run the frontend

From the repository root, with stable Rust installed:

cargo build --release --locked
OMNI_JEV_BIND=127.0.0.1:8080 \
OMNI_JEV_BACKEND_URL=http://127.0.0.1:8000 \
  ./target/release/omni-jev

Start the worker separately. See the frontend documentation for the HTTP interface and configuration, or the Laya recipe for a CPU text worker and response checks.

Architecture

Share serving infrastructure; let each model own its execution.

System1-Omni architecture: Rust frontend, model-owned execution, and CUDA and Metal backends

Layer Responsibility
Rust frontend API, request lifecycle, and response delivery through a small engine interface.
System1-Omni models Model-specific preprocessing and postprocessing, batching, state, execution, and kernel selection.
CUDA backend High-performance GPU operations for NVIDIA GPUs.
Metal backend High-performance GPU operations for Apple GPUs.

Each model owns its complete request-to-result path. Shared utilities stay minimal and are extracted when implementations need the same functionality. Backends can optimize for their hardware without requiring identical internal implementations.

Repository layout

Implementation code lives under src/; recipes and documentation stay at the repository root.

Directory Responsibility
src/frontend/ Rust serving code, Python worker adapters, and the small engine interface.
src/models/ Model implementations, one directory per model: preprocessing, batching, state, execution, and output processing.
src/backends/cuda/ NVIDIA GPU operations and kernel integration.
src/backends/metal/ Apple GPU operations and kernel integration.
recipe/ Model setup instructions, launch commands, configuration examples, and example requests.
docs/ Project documentation and architecture assets.

The frontend and the Cua-S1 native worker are Cargo workspace members. The other model and backend directories currently document planned work; they do not prescribe process boundaries.

Supported models

LAYA can run as an external Python worker for text requests; its in-repository model engine is still planned. The Cua-S1 4B 0.2 text adapter runs as a Python worker or as a native worker on CUDA:

Model Status
LAYA External worker; model engine planned
Cua-S1 4B 0.2 (text adapter) Python worker; native worker, CUDA, run on sm_89

CUDA and Metal coverage will be documented per model as implementations are added and validated.

Benchmarks

See the GPU serving benchmark for request replay, output-fidelity checks, and the CUDA comparison protocol. GPU performance measurements are pending.

Stay Tuned with Us

If you find system1-omni useful, give us a star on GitHub to support the project and help others discover it!

GitHub repository screenshot demonstrating a click on Star, turning the star yellow and showing Starred

About

community maintained vllm/vllm-omni style inference engine for system1 models in Rust for extreme efficient performance

Resources

Contributing

Stars

15 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages