Skip to content

open_jev: add native Rust/CUDA text worker - #55

Draft
hsliuustc0106 wants to merge 7 commits into
mainfrom
codex-open-jev-native
Draft

hsliuustc0106 wants to merge 7 commits into
mainfrom
codex-open-jev-native

Conversation

@hsliuustc0106

@hsliuustc0106 hsliuustc0106 commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Closes #54

Serve Open-Jev-27B-v1.1 through the existing Rust frontend with a native Rust/CUDA text worker. It compiles and tokenizes independent candidate prompts, applies the trained scalar decision head and saved temperature, and returns typed choice, score and noul answers.

  • Extract the Qwen prefill implementation shared by Cua-S1 and Open-Jev, preserving Cua-S1's public imports.
  • Add a CPU checkpoint export recipe and real inference before worker readiness.
  • Fuse attention gating, cache residual RMSNorm values and pack BF16 MLP SiLU loads/stores while preserving rounding and reduction order. Other SiLU layouts use the scalar fallback.
  • Add reference contract, tokenization and kernel coverage, plus setup and L20X results.

The CUDA ABI is now 4; rebuild the shared library and both native workers together.

Test Plan

System1-Omni Version / Commit: base 1be7d41eb74d5b49fd25194042041319950e5dce; head 40062bc12ac4b0d6364f61f0db963a49a1fde4f9.

Open-Jev tests and fixtures live under the repository-level tests/open_jev/; shared Qwen JSON/configuration and CUDA tests live under tests/qwen3_5/. Cargo explicitly registers the integration tests, and private unit tests load their files from that same top-level tree. Frontend CPU mock-worker coverage is tracked separately in issue #46 and PR #58.

Check prompts, typed answers and token IDs against pinned Open-Jev fixtures. Run the workspace checks and reserved-GPU kernel/worker validation. The recorded performance comparison fixes cached RMSNorm in both native variants and measures scalar versus packed SiLU on 74 single-candidate JevBench requests: one excluded feasibility pass and two measured passes per configuration. Hardware, revisions, controls and acceptance gates are in the linked results.

Test Result

Passed locally after cleanup:

cargo fmt --all --check
cargo clippy --workspace --locked --all-targets -- -D warnings
cargo test --workspace --locked
cargo build --workspace --release --locked
  • After moving the tests: all workspace Rust checks passed, including 29 CPU tests; the pinned-checkpoint tokenizer test and strict MkDocs build passed. All 38 registered test names are preserved, including ignored GPU/checkpoint cases. The relocated six-test GPU suite compiled in release mode without device execution.
  • Six CUDA ABI 4 tests and live worker/frontend smoke checks previously passed on reserved L20X GPU 2 at 61b83b3. The CUDA source and GPU test contents are unchanged; their location and registration changed, and device checks were not repeated. SM89 compilation passed previously; SM89 device execution and full Cua-S1 checkpoint inference were not checked.
  • Recorded packed-SiLU improvement: warm HTTP mean 49.527→48.297 ms (2.483%), with all 74 native probabilities unchanged. Long-request SiLU time 21.217→6.380 ms (69.93%), closing 94.19% of that measured kernel-family gap to OpenJev-Fast.

The comparison covers single-candidate, eager text inference. Fast's preparation variability prevents claiming a stable aggregate winner, and the measured outputs do not establish native/Fast accuracy parity. Multi-candidate prefix sharing and Open-Jev CUDA Graph replay remain unvalidated. Nsight Systems supplied timelines because Nsight Compute counters were denied.

Self-review

Contributor checklist from CONTRIBUTING.md, left for completion before requesting review:

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[New Model]: Native Rust/CUDA support for Open-Jev-27B-v1.1

1 participant