You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Add native inference for Mapika/decider-2b, starting with its released 2B v11 text checkpoint on CUDA. It answers choice, score, and noul questions by scoring allowed answers directly, without generating text. A native worker would let System1-Omni serve this checkpoint with Rust request processing and the project's CUDA execution path.
At the inspected main revision, the frontend proxies requests to a model worker and Cua-S1 has a native Qwen3.5 worker. Decider uses a related Qwen3.5 backbone, but its prompt format, candidate vocabulary, temperatures, Score execution, and answer formatting belong to Decider. Support is not currently implemented.
Proposed Change
Target and reference
Freeze these sources for the first implementation and its fixtures:
Artifact
Pinned revision
System1-Omni architecture and current native worker
The checkpoint's model configuration specifies 24 layers: 18 Gated DeltaNet layers and 6 full-attention layers; hidden size 2048, MLP size 6144, 8 attention heads, 2 KV heads, and tied input/output embeddings. Its BF16 checkpoint already includes the merged fine-tuning weights. There is no separate adapter merge or trained pointer/scalar head to add at runtime.
Ownership and integration
Add the model-owned implementation under src/models/decider/, with a native Rust worker and setup/validation instructions under recipe/decider/. Load the pinned local safetensors checkpoint, tokenizer, config.json, and decider_config.json directly; record artifact hashes during preparation.
Serve POST /v1/systemone and GET /health behind the existing Rust frontend. Keep the frontend's current byte-forwarding boundary; Decider owns JSON interpretation, tokenization, execution, and response construction. Serving requires no Python, PyTorch, or external decider.serve process. Python remains useful for preparing reference fixtures and comparisons.
Start with the released defaults: text, BF16, plain state-first prompts, independent questions, and isolated Score levels. Execute complete scoring rows eagerly through one serialized GPU executor. Tokenize the common state once on CPU; initially the GPU may recompute it for each row. Bound expanded rows and total prompt tokens before submitting device work, and document admission limits.
Readiness follows checkpoint/configuration validation, device initialization, and a real warmup inference. Report the loaded model revision and effective temperatures in worker metadata. The frontend continues to use the worker's real health result.
Preserve JSON order, Unicode, author serialization, descriptions, and long-array index annotation. Question identifiers stay outside model text. Retain the released state truncation rule: the tokenized Context: prefix plus state is capped at 32,768 tokens; question suffixes are separate. This is not a 32,768-token cap on the complete scoring row.
Prompt
Use the checkpoint's plain Context / Question / Options / Answer layout and its tokenization boundaries. No answer token is inserted. For the initial independent-row mode, read the final answer-slot hidden state.
Choice
Support 2–255 options in request order. Up to 10 options use the narrow rendering; larger sets use the reference's explicit token assembly. Build the 255-entry label table from the tokenizer: A–Z, then eligible single-token two-letter labels. A fixed 26-letter readout is insufficient.
Noul
Score no and yes, including optional true/false descriptions and the reference fallback instruction when descriptions supply the question. Return the calibrated probability of yes.
Score
Support 2–10 ordered levels. With the released isolated_levels=true, each level becomes its own yes/no row after stripping a leading numeric level label. Normalize the per-level P(yes) values, retain the reference zero-mass handling, and compute the expected level. This is not a single softmax over level labels.
Calibration
Load the released configuration: Choice 1.164, Noul 1.624, Score 1.124; fallback temperature 1.145. Each isolated Score row uses the Score temperature. Preserve neutralize_none=false and schema_first=false. Validate configuration before readiness.
Readout
Project only the allowed label rows of the tied LM output weight. Match the reference BF16 projection/cast boundaries before FP32 temperature scaling and softmax; qualify remaining backend numerical differences through validation.
Responses
Preserve decider-2b-v11 identity, typed answers, probability/score rounding, tie-breaking, and the author's Choice/Score confidence formulas. Preserve certainty, x_p_max, legend, level_fit, and fit_mass where emitted. Keep the reference unique-prefix token accounting and usage.output_tokens=0; report actual processed rows separately. Empty questions return empty answers without device execution.
Do not inherit Cua-S1's chat prompt, 26-option limit, or entropy-based confidence: these would change Decider's behavior. Unsupported execution modes must be rejected explicitly rather than silently interpreted as the defaults.
Delivery and alternatives
Contract and checkpoint support: native configuration/loading, Rust request/token fixtures, and output assembly checked against the frozen author code.
Native CUDA worker: reuse the Qwen prefill implementation, add the Decider readout, and validate direct and frontend-proxied requests against the pinned reference.
Measured optimization follow-ups: request-local shared-prefix execution, row batching, and CUDA Graphs only after the eager path passes parity checks. Shared-prefix reuse must preserve full-attention KV, Gated DeltaNet recurrent state, convolution history, and suffix positions; reusing only attention KV is incomplete.
An external decider.serve worker is a useful reference/baseline but does not deliver native inference. A separate Qwen implementation would duplicate the existing layer loop and kernels. Switching this checkpoint to chat/schema-first prompts, packed questions, or a different Score readout would change the reference computation.
Initial scope excludes other Decider sizes, vision, Metal/CPU inference, GGUF/quantization, schema caches, packed independent=false execution, and cross-request scheduling changes. The published tensor metadata contains 1,881,825,088 BF16 parameters, about 3.76 GB of tensor data; device memory also includes workspaces and activations. No native memory or performance improvement has been measured.
Validation and acceptance
CPU golden fixtures match the author's rendered inputs, exact token IDs, row plans, candidate IDs, and output assembly. Cover Choice counts 2/10/11/26/255, Score counts 2/3/10, Noul descriptions, mixed requests, Unicode/structured state, array-index annotation, empty questions, truncation, invalid criteria, and unsupported modes.
The loader validates the pinned configuration, BF16 tensor shapes, tied output weights, and the complete 255-token label table. Wrong artifacts fail before readiness.
Reserved-GPU checks exercise the 2B attention/Gated DeltaNet shapes, complete checkpoint inference, and direct/front-end transport. Existing Cua-S1 checks continue to pass.
Freeze the comparison manifest and numerical gates before GPU execution. Suggested initial gates: maximum absolute answer-probability drift 0.02, Score expectation drift 0.1, and matching discrete decisions when the reference top-two margin is at least 0.05. Report raw per-level fit and assembled probabilities, worst cases, and low-margin disagreements. These are proposed gates for discussion, not established parity results.
Reuse the HTTP benchmark harness with Decider-aware validation. Its current rounding limitation must be addressed for the author's four-decimal distributions, especially with 255 options; preserve the public response rather than renormalizing it to pass a validator.
Any speed claim compares the pinned author's CUDA runtime and native worker on the same exact reserved GPU, BF16 checkpoint, requests, tokenizer, temperatures, affinity, readiness behavior, and cache policy. Declare runtime implementation as the changed variable and record necessary batching/kernel differences. Start with a CPU preflight and one feasibility run per configuration, then two measured runs per configuration at each predeclared concurrency. Separate preparation and process-to-readiness from warmed p50/p95, requests/s, decisions/s, errors, and peak memory. Stop on a correctness failure or the run limit; preserve raw results and variability.
Document preparation, build/run commands, supported hardware, limits, and validation status. Apply the repository's contribution/self-review checks to implementation PRs.
Feedback Period
Initial feedback through 2026-10-08. The main questions are whether to coordinate the shared Qwen module extraction with #55, whether the proposed eager/default-mode scope is appropriate, and which Decider-specific numerical gates should be frozen before GPU validation.
Anything else
This issue proposes implementation; it does not claim delivered support or native speedups. Preparation for this RFC inspected the pinned sources and ran eight CPU-only reference prompt/schema probes, including the 10/11-option rendering boundary, 255 distinct label tokens, Score expansion, mixed structured requests, and empty questions. Four invalid-criteria probes were rejected. These used the tokenizer and pure preprocessing code only; no model weights were loaded and no GPU inference or benchmarks were run.
Searched the existing issues and PRs and read the current architecture, frontend/model documentation, contribution guide, and RFC template. No existing issue or PR specifically tracks native Decider-2B support.
Motivation
Add native inference for Mapika/decider-2b, starting with its released 2B v11 text checkpoint on CUDA. It answers
choice,score, andnoulquestions by scoring allowed answers directly, without generating text. A native worker would let System1-Omni serve this checkpoint with Rust request processing and the project's CUDA execution path.At the inspected main revision, the frontend proxies requests to a model worker and Cua-S1 has a native Qwen3.5 worker. Decider uses a related Qwen3.5 backbone, but its prompt format, candidate vocabulary, temperatures, Score execution, and answer formatting belong to Decider. Support is not currently implemented.
Proposed Change
Target and reference
Freeze these sources for the first implementation and its fixtures:
fd0d87e7ebe2d76ea5fe93d125bcf3323d1ee601533964dae8be954c5b5e19fa4948e48408094c1edecider-ai1.8.150d0be0d7cb43d2066965ce5fa7f3fe4e489a60fThe checkpoint's model configuration specifies 24 layers: 18 Gated DeltaNet layers and 6 full-attention layers; hidden size 2048, MLP size 6144, 8 attention heads, 2 KV heads, and tied input/output embeddings. Its BF16 checkpoint already includes the merged fine-tuning weights. There is no separate adapter merge or trained pointer/scalar head to add at runtime.
Ownership and integration
src/models/decider/, with a native Rust worker and setup/validation instructions underrecipe/decider/. Load the pinned local safetensors checkpoint, tokenizer,config.json, anddecider_config.jsondirectly; record artifact hashes during preparation.POST /v1/systemoneandGET /healthbehind the existing Rust frontend. Keep the frontend's current byte-forwarding boundary; Decider owns JSON interpretation, tokenization, execution, and response construction. Serving requires no Python, PyTorch, or externaldecider.serveprocess. Python remains useful for preparing reference fixtures and comparisons.src/backends/cuda/qwen3_5/. Coordinate with open_jev: add native Rust/CUDA text worker #55, which proposes extracting the native Qwen layer loop/loader intosrc/models/qwen3_5/native/; that extraction is pending, not current main. If Decider lands first, make the smallest shared extraction needed by the existing Cua-S1 worker and Decider. Preserve the Cua-S1 API and regression coverage. Coordinate the same boundary with Kev support in [New Model]: Support Kev (0.5B–9B) — second Qwen3.5 decision model, reusing the qwen3_5 CUDA backend #27.Decider contract to preserve
Use the pinned runtime's request rendering and output assembly, HTTP preparation, and token construction as the compatibility reference.
Context:prefix plus state is capped at 32,768 tokens; question suffixes are separate. This is not a 32,768-token cap on the complete scoring row.Context / Question / Options / Answerlayout and its tokenization boundaries. No answer token is inserted. For the initial independent-row mode, read the final answer-slot hidden state.A–Z, then eligible single-token two-letter labels. A fixed 26-letter readout is insufficient.noandyes, including optional true/false descriptions and the reference fallback instruction when descriptions supply the question. Return the calibrated probability ofyes.isolated_levels=true, each level becomes its own yes/no row after stripping a leading numeric level label. Normalize the per-levelP(yes)values, retain the reference zero-mass handling, and compute the expected level. This is not a single softmax over level labels.1.164, Noul1.624, Score1.124; fallback temperature1.145. Each isolated Score row uses the Score temperature. Preserveneutralize_none=falseandschema_first=false. Validate configuration before readiness.decider-2b-v11identity, typed answers, probability/score rounding, tie-breaking, and the author's Choice/Score confidence formulas. Preservecertainty,x_p_max,legend,level_fit, andfit_masswhere emitted. Keep the reference unique-prefix token accounting andusage.output_tokens=0; report actual processed rows separately. Empty questions return empty answers without device execution.Do not inherit Cua-S1's chat prompt, 26-option limit, or entropy-based
confidence: these would change Decider's behavior. Unsupported execution modes must be rejected explicitly rather than silently interpreted as the defaults.Delivery and alternatives
An external
decider.serveworker is a useful reference/baseline but does not deliver native inference. A separate Qwen implementation would duplicate the existing layer loop and kernels. Switching this checkpoint to chat/schema-first prompts, packed questions, or a different Score readout would change the reference computation.Initial scope excludes other Decider sizes, vision, Metal/CPU inference, GGUF/quantization, schema caches, packed
independent=falseexecution, and cross-request scheduling changes. The published tensor metadata contains 1,881,825,088 BF16 parameters, about 3.76 GB of tensor data; device memory also includes workspaces and activations. No native memory or performance improvement has been measured.Validation and acceptance
Feedback Period
Initial feedback through 2026-10-08. The main questions are whether to coordinate the shared Qwen module extraction with #55, whether the proposed eager/default-mode scope is appropriate, and which Decider-specific numerical gates should be frozen before GPU validation.
Anything else
This issue proposes implementation; it does not claim delivered support or native speedups. Preparation for this RFC inspected the pinned sources and ran eight CPU-only reference prompt/schema probes, including the 10/11-option rendering boundary, 255 distinct label tokens, Score expansion, mixed structured requests, and empty questions. Four invalid-criteria probes were rejected. These used the tokenizer and pure preprocessing code only; no model weights were loaded and no GPU inference or benchmarks were run.
Related: #9 (model/backend ownership), #27 (Kev/Qwen reuse), #55 (pending shared native Qwen implementation), #19 (merged native Cua-S1 worker), and #39 (serving comparison protocol).
Before submitting a new issue