You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Cua-S1 multimodal stack: validation results and cross-PR integration checklist
The multimodal stack #12 → #15 → #17 → #18 → #20 → #22 → #24 → #33 is ready for maintainer review with the following validation evidence and explicit integration limits. No new multimodal runtime regression was observed in the tested combination. This is not a claim that all open PRs can be merged unchanged or that every integration check is green.
Versions and scope
Base main: 30622438bbe824563a066a0525631ae6e1beaabb.
Multimodal head: be44f57, including the shallow-checkout test fix and pre-inference image aspect-ratio validation.
Only README/index documentation, Cargo workspace membership and the lockfile were reconciled. Production and test source files were retained from their final constituent PR heads. Multimodal execution, frontend source and multimodal dependency pins are unchanged from Bound Cua-S1 Graph capture work with request admission and cooldown #33.
This temporary integration is a review fixture, not a proposed merge of every PR into main.
178 tests pass with Torch; actual depth-1 CPU checkout passes 177 with one optional Torch skip
Combined Cua text + multimodal suite
222 passed, including local pinned-tokenizer checks
Combined Rust workspace
Format and release build pass; 30 tests pass, 10 asset-dependent tests ignored
Laya worker, separate pytest process
22 passed, 13 real-model contract tests skipped
LFM, its own pinned environment
26 passed, including the tiny-model CPU engine tests
CUDA checker unit tests
53 passed; repository contract check separately fails as described below
Shared response-comparison recipe
Five positive/negative cases pass
Multimodal-owned Python files
Lint and format pass in the integration checkout
Real multimodal worker through combined frontend
12 direct/proxied request pairs have identical status, Content-Type and response bytes
The HTTP check used the release frontend built from the combined Cargo workspace on macOS, connected over an SSH tunnel to the combined checkout's worker on an RTX 4090. It covered health, cold/capture/replay, changed images/text, invalid aspect ratios, malformed/invalid requests, and recovery after rejection. All 12 valid model predictions across the two routes exactly matched their eager references. Graph counters recorded 1 capture and 21 replays, with zero capture OOMs/errors or numerical mismatches.
The earlier full GPU matrix and reproduction material remain in the #33 experiment report. The full performance matrix was not repeated for this combination because multimodal execution and its pinned dependencies did not change.
cua_s1: add a native CUDA text worker #19 and the multimodal workflow's shared recipe scope: the existing Ruff command now sees five E741/E731 errors in recipe/cua_s1/check_native.py and diff_workers.py; those two files and diff_corpus.py fail formatting. The individual multimodal files remain clean.
Python environment separation: LFM pins huggingface-hub==1.31.0, while Cua pins ==1.32.0. Installing both requirements into one environment is unsatisfiable. The tests used separate environments without loosening pins.
These results do not validate other models' full checkpoints, native CUDA kernel execution, Apple Silicon/MPS, all merge orders, or concurrent multi-model GPU memory use. Each frontend instance still targets one configured upstream; it does not automatically route different model aliases. Multiple workers also need distinct listening ports.
Approve/run the fork workflows: multimodal CPU checks and CI. Both currently report action_required; the local evidence is not a substitute for green GitHub checks.
Decide whether to merge the multimodal stack independently or coordinate a larger integration.
For a larger integration, resolve the listed shared checks, manifests, test isolation and environment requirements without discarding either modality's documentation.
The unpublished integration fixture is commit 5096ea22687ec000314f2a1e825083c9bfeb15a5. Logs, a reproducible Git bundle and the direct/proxy JSON results were retained locally. The GPU server was shut down after retrieving and checksum-verifying the results.
Cua-S1 multimodal stack: validation results and cross-PR integration checklist
The multimodal stack #12 → #15 → #17 → #18 → #20 → #22 → #24 → #33 is ready for maintainer review with the following validation evidence and explicit integration limits. No new multimodal runtime regression was observed in the tested combination. This is not a claim that all open PRs can be merged unchanged or that every integration check is green.
Versions and scope
30622438bbe824563a066a0525631ae6e1beaabb.be44f57, including the shallow-checkout test fix and pre-inference image aspect-ratio validation.Validation results
The HTTP check used the release frontend built from the combined Cargo workspace on macOS, connected over an SSH tunnel to the combined checkout's worker on an RTX 4090. It covered health, cold/capture/replay, changed images/text, invalid aspect ratios, malformed/invalid requests, and recovery after rejection. All 12 valid model predictions across the two routes exactly matched their eager references. Graph counters recorded 1 capture and 21 replays, with zero capture OOMs/errors or numerical mismatches.
The earlier full GPU matrix and reproduction material remain in the #33 experiment report. The full performance matrix was not repeated for this combination because multimodal execution and its pinned dependencies did not change.
Integration items that still need coordination
recipe/cua_s1/check_native.pyanddiff_workers.py; those two files anddiff_corpus.pyfail formatting. The individual multimodal files remain clean.engine.rs:172andmodel.rs:925triggerchunks_exact_to_as_chunksunder-D warnings. This also reproduces on cua_s1: add a native CUDA text worker #19 alone, so it is not introduced by multimodal integration.qwen3_5contains CUDA sources but has no backend manifest required by the new checker. The actual repository check exits 1 even though the checker's own unit tests pass.tests/test_worker.pyand a bareimport worker. Default combined collection fails with an import-file mismatch; importlib mode still loads the wrongworkermodule. Separate pytest processes pass. Package/module isolation or explicitly separate CI invocations are needed.huggingface-hub==1.31.0, while Cua pins==1.32.0. Installing both requirements into one environment is unsatisfiable. The tests used separate environments without loosening pins.These results do not validate other models' full checkpoints, native CUDA kernel execution, Apple Silicon/MPS, all merge orders, or concurrent multi-model GPU memory use. Each frontend instance still targets one configured upstream; it does not automatically route different model aliases. Multiple workers also need distinct listening ports.
Maintainer review checklist
action_required; the local evidence is not a substitute for green GitHub checks.Integration snapshot
be44f57ed8a34c4122a6a03fb0adf33bf3000c5172a5e0f47203e208bd23033a1dab1b21f4a4ca5dc8cebc4db70bea6855fac1c2913f335f4df23f95fd841e1967e0a58b0a1af7b8425e83579f1c576cd61c0d139621fd0f8941360009859c38c45b5e17391cf45e4013b833473e81b71508d547217f1605d44d3fe31834b00f517f386e0081f3eea3cb84bca402606cd9ee6a8c129be3df51b1c923ec7bf324The unpublished integration fixture is commit
5096ea22687ec000314f2a1e825083c9bfeb15a5. Logs, a reproducible Git bundle and the direct/proxy JSON results were retained locally. The GPU server was shut down after retrieving and checksum-verifying the results.