Skip to content

Cua-S1 multimodal stack: validation results and cross-PR integration checklist #36

Description

@Levius-Fubuki

Cua-S1 multimodal stack: validation results and cross-PR integration checklist

The multimodal stack #12 → #15 → #17 → #18 → #20 → #22 → #24 → #33 is ready for maintainer review with the following validation evidence and explicit integration limits. No new multimodal runtime regression was observed in the tested combination. This is not a claim that all open PRs can be merged unchanged or that every integration check is green.

Versions and scope

Validation results

Check Result
Standalone fixed #33 178 tests pass with Torch; actual depth-1 CPU checkout passes 177 with one optional Torch skip
Combined Cua text + multimodal suite 222 passed, including local pinned-tokenizer checks
Combined Rust workspace Format and release build pass; 30 tests pass, 10 asset-dependent tests ignored
Laya worker, separate pytest process 22 passed, 13 real-model contract tests skipped
LFM, its own pinned environment 26 passed, including the tiny-model CPU engine tests
CUDA checker unit tests 53 passed; repository contract check separately fails as described below
Shared response-comparison recipe Five positive/negative cases pass
Multimodal-owned Python files Lint and format pass in the integration checkout
Real multimodal worker through combined frontend 12 direct/proxied request pairs have identical status, Content-Type and response bytes

The HTTP check used the release frontend built from the combined Cargo workspace on macOS, connected over an SSH tunnel to the combined checkout's worker on an RTX 4090. It covered health, cold/capture/replay, changed images/text, invalid aspect ratios, malformed/invalid requests, and recovery after rejection. All 12 valid model predictions across the two routes exactly matched their eager references. Graph counters recorded 1 capture and 21 replays, with zero capture OOMs/errors or numerical mismatches.

The earlier full GPU matrix and reproduction material remain in the #33 experiment report. The full performance matrix was not repeated for this combination because multimodal execution and its pinned dependencies did not change.

Integration items that still need coordination

  1. Merge resolutions: Bound Cua-S1 Graph capture work with request admission and cooldown #33 has README conflicts with docs: add Cua-S1 4B 0.2 inference contract #11/cua_s1: add a text worker that loads Qwen3.5-4B through Transformers #13/cua_s1: add a native CUDA text worker #19/Add CPU checkpoint loading for Laya #21/Add LFM2.5-350M choice worker with candidate batching #32. Combining the other branches also requires reconciling Cargo.toml/Cargo.lock, recipe/README.md and src/models/laya/README.md. The temporary workspace retains all four Rust members: frontend, Cua native, Laya and CLM.
  2. cua_s1: add a native CUDA text worker #19 and the multimodal workflow's shared recipe scope: the existing Ruff command now sees five E741/E731 errors in recipe/cua_s1/check_native.py and diff_workers.py; those two files and diff_corpus.py fail formatting. The individual multimodal files remain clean.
  3. Existing cua_s1: add a native CUDA text worker #19 Clippy failure: on Rust 1.98, engine.rs:172 and model.rs:925 trigger chunks_exact_to_as_chunks under -D warnings. This also reproduces on cua_s1: add a native CUDA text worker #19 alone, so it is not introduced by multimodal integration.
  4. cua_s1: add a native CUDA text worker #19 + cuda: add the backend contract and its checker #25 backend contract: qwen3_5 contains CUDA sources but has no backend manifest required by the new checker. The actual repository check exits 1 even though the checker's own unit tests pass.
  5. Laya on Apple Silicon: worker, benchmarks and recipe #30 + Add LFM2.5-350M choice worker with candidate batching #32 test isolation: both use tests/test_worker.py and a bare import worker. Default combined collection fails with an import-file mismatch; importlib mode still loads the wrong worker module. Separate pytest processes pass. Package/module isolation or explicitly separate CI invocations are needed.
  6. Python environment separation: LFM pins huggingface-hub==1.31.0, while Cua pins ==1.32.0. Installing both requirements into one environment is unsatisfiable. The tests used separate environments without loosening pins.

These results do not validate other models' full checkpoints, native CUDA kernel execution, Apple Silicon/MPS, all merge orders, or concurrent multi-model GPU memory use. Each frontend instance still targets one configured upstream; it does not automatically route different model aliases. Multiple workers also need distinct listening ports.

Maintainer review checklist

  • Review the cumulative multimodal stack at Bound Cua-S1 Graph capture work with request admission and cooldown #33, including the two validation fixes.
  • Approve/run the fork workflows: multimodal CPU checks and CI. Both currently report action_required; the local evidence is not a substitute for green GitHub checks.
  • Decide whether to merge the multimodal stack independently or coordinate a larger integration.
  • For a larger integration, resolve the listed shared checks, manifests, test isolation and environment requirements without discarding either modality's documentation.

Integration snapshot

Independent tip SHA
#33 be44f57ed8a34c4122a6a03fb0adf33bf3000c51
#19 (includes #11/#13) 72a5e0f47203e208bd23033a1dab1b21f4a4ca5d
#21 c8cebc4db70bea6855fac1c2913f335f4df23f95
#29 (includes #28) fd841e1967e0a58b0a1af7b8425e83579f1c576c
#23 d61c0d139621fd0f8941360009859c38c45b5e17
#25 391cf45e4013b833473e81b71508d547217f1605
#30 d44d3fe31834b00f517f386e0081f3eea3cb84bc
#32 a402606cd9ee6a8c129be3df51b1c923ec7bf324

The unpublished integration fixture is commit 5096ea22687ec000314f2a1e825083c9bfeb15a5. Logs, a reproducible Git bundle and the direct/proxy JSON results were retained locally. The GPU server was shut down after retrieving and checksum-verifying the results.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions