Skip to content

[New Model]: Native Rust/CUDA support for Open-Jev-27B-v1.1 #54

Description

@hsliuustc0106

The model to consider

Open-Jev-27B-v1.1, with the Open-Jev reference implementation and its Qwen3.8-27B text backbone. OpenJev-Fast provides optimization motivation; its hardware measurements are not performance evidence for System1-Omni.

The closest planned or implemented model

Cua-S1's native Rust/CUDA worker from merged PR #19 already implements the related Qwen hybrid-attention prefill path. Open-Jev can share that implementation while owning its request compiler, tokenizer, trained scalar head, calibration, and response contract.

What's your difficulty of supporting the model you want?

Implement text-only candidate scoring through the existing Rust frontend and /v1/systemone API:

  • Compile independent candidate prompts for choice, ordinal score, and noul questions, preserving reference formatting and calibration.
  • Export the pinned LoRA adapter into a merged BF16 backbone and retain the trained FP32 scalar head and temperature.
  • Share the existing native Qwen prefill implementation with Cua-S1.
  • Fuse the attention sigmoid gate into the CUDA attention epilogue, preserving the separate implementation's BF16 rounding points.
  • Document preparation, serving, numerical limitations, and reproducible validation.

This initial scope uses eager inference. Prefix sharing, CUDA Graphs, GEMM tuning, quantization, and multimodal support are outside this issue.

Use case and motivation

Support typed text decisions with Rust request processing and native CUDA execution, extending the infrastructure introduced in #19 to Open-Jev rather than requiring its Python service at runtime. Python remains a checkpoint preparation dependency.

Acceptance and current evidence

  • Rust request/response golden fixtures for all three question types.
  • Exact pinned-tokenizer parity and prompt length-limit checks.
  • CPU checkpoint export and reference request smoke tests.
  • Workspace formatting, clippy, CPU tests, and release builds; CUDA compilation for sm_89.
  • Execute fused-versus-separate attention gating equality tests on a reserved GPU, including the retained attention and Gated DeltaNet reference checks.
  • Validate the complete native checkpoint against the reference model and exercise direct and frontend-proxied inference after real readiness.

GPU execution is pending: the previous scheduler wait timed out without running the validation job. No native full-model accuracy or latency improvement is claimed.

Before submitting a new issue

  • Searched the issue tracker and checked the documentation. Existing broader model discussions do not track this native Open-Jev implementation.

Implementation

Draft PR: #55. Based on merged #19; CPU checks pass and GPU validation remains pending.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    new modelRequests to support a new model

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions