The model to consider
Open-Jev-27B-v1.1, with the Open-Jev reference implementation and its Qwen3.8-27B text backbone. OpenJev-Fast provides optimization motivation; its hardware measurements are not performance evidence for System1-Omni.
The closest planned or implemented model
Cua-S1's native Rust/CUDA worker from merged PR #19 already implements the related Qwen hybrid-attention prefill path. Open-Jev can share that implementation while owning its request compiler, tokenizer, trained scalar head, calibration, and response contract.
What's your difficulty of supporting the model you want?
Implement text-only candidate scoring through the existing Rust frontend and /v1/systemone API:
- Compile independent candidate prompts for
choice, ordinal score, and noul questions, preserving reference formatting and calibration.
- Export the pinned LoRA adapter into a merged BF16 backbone and retain the trained FP32 scalar head and temperature.
- Share the existing native Qwen prefill implementation with Cua-S1.
- Fuse the attention sigmoid gate into the CUDA attention epilogue, preserving the separate implementation's BF16 rounding points.
- Document preparation, serving, numerical limitations, and reproducible validation.
This initial scope uses eager inference. Prefix sharing, CUDA Graphs, GEMM tuning, quantization, and multimodal support are outside this issue.
Use case and motivation
Support typed text decisions with Rust request processing and native CUDA execution, extending the infrastructure introduced in #19 to Open-Jev rather than requiring its Python service at runtime. Python remains a checkpoint preparation dependency.
Acceptance and current evidence
GPU execution is pending: the previous scheduler wait timed out without running the validation job. No native full-model accuracy or latency improvement is claimed.
Before submitting a new issue
Implementation
Draft PR: #55. Based on merged #19; CPU checks pass and GPU validation remains pending.
The model to consider
Open-Jev-27B-v1.1, with the Open-Jev reference implementation and its Qwen3.8-27B text backbone. OpenJev-Fast provides optimization motivation; its hardware measurements are not performance evidence for System1-Omni.
The closest planned or implemented model
Cua-S1's native Rust/CUDA worker from merged PR #19 already implements the related Qwen hybrid-attention prefill path. Open-Jev can share that implementation while owning its request compiler, tokenizer, trained scalar head, calibration, and response contract.
What's your difficulty of supporting the model you want?
Implement text-only candidate scoring through the existing Rust frontend and
/v1/systemoneAPI:choice, ordinalscore, andnoulquestions, preserving reference formatting and calibration.This initial scope uses eager inference. Prefix sharing, CUDA Graphs, GEMM tuning, quantization, and multimodal support are outside this issue.
Use case and motivation
Support typed text decisions with Rust request processing and native CUDA execution, extending the infrastructure introduced in #19 to Open-Jev rather than requiring its Python service at runtime. Python remains a checkpoint preparation dependency.
Acceptance and current evidence
GPU execution is pending: the previous scheduler wait timed out without running the validation job. No native full-model accuracy or latency improvement is claimed.
Before submitting a new issue
Implementation
Draft PR: #55. Based on merged #19; CPU checks pass and GPU validation remains pending.