Skip to content

[RFC]: Native Rust/CUDA inference for Mapika/decider-2b #57

Description

@hsliuustc0106

Motivation

Add native inference for Mapika/decider-2b, starting with its released 2B v11 text checkpoint on CUDA. It answers choice, score, and noul questions by scoring allowed answers directly, without generating text. A native worker would let System1-Omni serve this checkpoint with Rust request processing and the project's CUDA execution path.

At the inspected main revision, the frontend proxies requests to a model worker and Cua-S1 has a native Qwen3.5 worker. Decider uses a related Qwen3.5 backbone, but its prompt format, candidate vocabulary, temperatures, Score execution, and answer formatting belong to Decider. Support is not currently implemented.

Proposed Change

Target and reference

Freeze these sources for the first implementation and its fixtures:

Artifact Pinned revision
System1-Omni architecture and current native worker fd0d87e7ebe2d76ea5fe93d125bcf3323d1ee601
Decider-2B v11 weights, tokenizer, and configuration 533964dae8be954c5b5e19fa4948e48408094c1e
Author's reference runtime, decider-ai 1.8.1 50d0be0d7cb43d2066965ce5fa7f3fe4e489a60f

The checkpoint's model configuration specifies 24 layers: 18 Gated DeltaNet layers and 6 full-attention layers; hidden size 2048, MLP size 6144, 8 attention heads, 2 KV heads, and tied input/output embeddings. Its BF16 checkpoint already includes the merged fine-tuning weights. There is no separate adapter merge or trained pointer/scalar head to add at runtime.

Ownership and integration

  • Add the model-owned implementation under src/models/decider/, with a native Rust worker and setup/validation instructions under recipe/decider/. Load the pinned local safetensors checkpoint, tokenizer, config.json, and decider_config.json directly; record artifact hashes during preparation.
  • Serve POST /v1/systemone and GET /health behind the existing Rust frontend. Keep the frontend's current byte-forwarding boundary; Decider owns JSON interpretation, tokenization, execution, and response construction. Serving requires no Python, PyTorch, or external decider.serve process. Python remains useful for preparing reference fixtures and comparisons.
  • Reuse src/backends/cuda/qwen3_5/. Coordinate with open_jev: add native Rust/CUDA text worker #55, which proposes extracting the native Qwen layer loop/loader into src/models/qwen3_5/native/; that extraction is pending, not current main. If Decider lands first, make the smallest shared extraction needed by the existing Cua-S1 worker and Decider. Preserve the Cua-S1 API and regression coverage. Coordinate the same boundary with Kev support in [New Model]: Support Kev (0.5B–9B) — second Qwen3.5 decision model, reusing the qwen3_5 CUDA backend #27.
  • Start with the released defaults: text, BF16, plain state-first prompts, independent questions, and isolated Score levels. Execute complete scoring rows eagerly through one serialized GPU executor. Tokenize the common state once on CPU; initially the GPU may recompute it for each row. Bound expanded rows and total prompt tokens before submitting device work, and document admission limits.
  • Readiness follows checkpoint/configuration validation, device initialization, and a real warmup inference. Report the loaded model revision and effective temperatures in worker metadata. The frontend continues to use the worker's real health result.

Decider contract to preserve

Use the pinned runtime's request rendering and output assembly, HTTP preparation, and token construction as the compatibility reference.

Part Required behavior
State and questions Preserve JSON order, Unicode, author serialization, descriptions, and long-array index annotation. Question identifiers stay outside model text. Retain the released state truncation rule: the tokenized Context: prefix plus state is capped at 32,768 tokens; question suffixes are separate. This is not a 32,768-token cap on the complete scoring row.
Prompt Use the checkpoint's plain Context / Question / Options / Answer layout and its tokenization boundaries. No answer token is inserted. For the initial independent-row mode, read the final answer-slot hidden state.
Choice Support 2–255 options in request order. Up to 10 options use the narrow rendering; larger sets use the reference's explicit token assembly. Build the 255-entry label table from the tokenizer: A–Z, then eligible single-token two-letter labels. A fixed 26-letter readout is insufficient.
Noul Score no and yes, including optional true/false descriptions and the reference fallback instruction when descriptions supply the question. Return the calibrated probability of yes.
Score Support 2–10 ordered levels. With the released isolated_levels=true, each level becomes its own yes/no row after stripping a leading numeric level label. Normalize the per-level P(yes) values, retain the reference zero-mass handling, and compute the expected level. This is not a single softmax over level labels.
Calibration Load the released configuration: Choice 1.164, Noul 1.624, Score 1.124; fallback temperature 1.145. Each isolated Score row uses the Score temperature. Preserve neutralize_none=false and schema_first=false. Validate configuration before readiness.
Readout Project only the allowed label rows of the tied LM output weight. Match the reference BF16 projection/cast boundaries before FP32 temperature scaling and softmax; qualify remaining backend numerical differences through validation.
Responses Preserve decider-2b-v11 identity, typed answers, probability/score rounding, tie-breaking, and the author's Choice/Score confidence formulas. Preserve certainty, x_p_max, legend, level_fit, and fit_mass where emitted. Keep the reference unique-prefix token accounting and usage.output_tokens=0; report actual processed rows separately. Empty questions return empty answers without device execution.

Do not inherit Cua-S1's chat prompt, 26-option limit, or entropy-based confidence: these would change Decider's behavior. Unsupported execution modes must be rejected explicitly rather than silently interpreted as the defaults.

Delivery and alternatives

  1. Contract and checkpoint support: native configuration/loading, Rust request/token fixtures, and output assembly checked against the frozen author code.
  2. Native CUDA worker: reuse the Qwen prefill implementation, add the Decider readout, and validate direct and frontend-proxied requests against the pinned reference.
  3. Measured optimization follow-ups: request-local shared-prefix execution, row batching, and CUDA Graphs only after the eager path passes parity checks. Shared-prefix reuse must preserve full-attention KV, Gated DeltaNet recurrent state, convolution history, and suffix positions; reusing only attention KV is incomplete.

An external decider.serve worker is a useful reference/baseline but does not deliver native inference. A separate Qwen implementation would duplicate the existing layer loop and kernels. Switching this checkpoint to chat/schema-first prompts, packed questions, or a different Score readout would change the reference computation.

Initial scope excludes other Decider sizes, vision, Metal/CPU inference, GGUF/quantization, schema caches, packed independent=false execution, and cross-request scheduling changes. The published tensor metadata contains 1,881,825,088 BF16 parameters, about 3.76 GB of tensor data; device memory also includes workspaces and activations. No native memory or performance improvement has been measured.

Validation and acceptance

  • CPU golden fixtures match the author's rendered inputs, exact token IDs, row plans, candidate IDs, and output assembly. Cover Choice counts 2/10/11/26/255, Score counts 2/3/10, Noul descriptions, mixed requests, Unicode/structured state, array-index annotation, empty questions, truncation, invalid criteria, and unsupported modes.
  • The loader validates the pinned configuration, BF16 tensor shapes, tied output weights, and the complete 255-token label table. Wrong artifacts fail before readiness.
  • Reserved-GPU checks exercise the 2B attention/Gated DeltaNet shapes, complete checkpoint inference, and direct/front-end transport. Existing Cua-S1 checks continue to pass.
  • Freeze the comparison manifest and numerical gates before GPU execution. Suggested initial gates: maximum absolute answer-probability drift 0.02, Score expectation drift 0.1, and matching discrete decisions when the reference top-two margin is at least 0.05. Report raw per-level fit and assembled probabilities, worst cases, and low-margin disagreements. These are proposed gates for discussion, not established parity results.
  • Reuse the HTTP benchmark harness with Decider-aware validation. Its current rounding limitation must be addressed for the author's four-decimal distributions, especially with 255 options; preserve the public response rather than renormalizing it to pass a validator.
  • Any speed claim compares the pinned author's CUDA runtime and native worker on the same exact reserved GPU, BF16 checkpoint, requests, tokenizer, temperatures, affinity, readiness behavior, and cache policy. Declare runtime implementation as the changed variable and record necessary batching/kernel differences. Start with a CPU preflight and one feasibility run per configuration, then two measured runs per configuration at each predeclared concurrency. Separate preparation and process-to-readiness from warmed p50/p95, requests/s, decisions/s, errors, and peak memory. Stop on a correctness failure or the run limit; preserve raw results and variability.
  • Document preparation, build/run commands, supported hardware, limits, and validation status. Apply the repository's contribution/self-review checks to implementation PRs.

Feedback Period

Initial feedback through 2026-10-08. The main questions are whether to coordinate the shared Qwen module extraction with #55, whether the proposed eager/default-mode scope is appropriate, and which Decider-specific numerical gates should be frozen before GPU validation.

Anything else

This issue proposes implementation; it does not claim delivered support or native speedups. Preparation for this RFC inspected the pinned sources and ran eight CPU-only reference prompt/schema probes, including the 10/11-option rendering boundary, 255 distinct label tokens, Score expansion, mixed structured requests, and empty questions. Four invalid-criteria probes were rejected. These used the tokenizer and pure preprocessing code only; no model weights were loaded and no GPU inference or benchmarks were run.

Related: #9 (model/backend ownership), #27 (Kev/Qwen reuse), #55 (pending shared native Qwen implementation), #19 (merged native Cua-S1 worker), and #39 (serving comparison protocol).

Before submitting a new issue

  • Searched the existing issues and PRs and read the current architecture, frontend/model documentation, contribution guide, and RFC template. No existing issue or PR specifically tracks native Decider-2B support.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    RFCRequest for comments on design changes

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions