Xingtong Ge1,2, Yutong Wang3, Lunjie Zhu1, Haitao Lin4, Fangyu Lin1, Yushi Huang1, Xin Zhang2, Yi Zhang2, Yu Liu2, Jun Zhang1
1The Hong Kong University of Science and Technology, 2Vivix Group Limited, 3The University of Sydney, 4Westlake University
Preprint, 2026
A successor to Salt, extending few-step distribution matching from video distillation to causal joint audio–video generation.
Few-step streaming audio–video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD. CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher. This self-supervised signal encourages the student to extract semantic information and improves cross-modal alignment. Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation. At 480p, Salt++ improves visual and motion quality by 57% and 45% over OmniForcing on JavisBench under the same 4-step causal setting. A separate scale-wise post-training stage extends Salt++ to 4-step 1664×960 generation, outperforming bidirectional LTX-2 on six of seven reported metrics.
Four-step streaming audio–video generation at 480p. Salt++ retains scene and facial detail that the causal baseline loses at the same step budget.
🔊 Watch the generated clips with audio on the project page. Audio and video are generated jointly, so audio–visual synchrony can be judged directly.
- Paper on arXiv
- Project page with audio–video samples
- Inference scripts
- Open-source model weights
- Training scripts
Stay tuned 🚀
Prior recipes switch objectives between stages and score the causal generator with bidirectional models. Salt++ keeps a single AR DMD objective throughout post-training and only shifts its conditioning from clean context to generated rollout:
- Causal Self-Flow (CSF) varies the history while holding the noisy target fixed, turning contextual information asymmetry into a representation-learning signal for the autoregressive teacher.
- Context-aligned AR DMD shares one causal mask and prefix across generator sampling, fake-score training, and real-score evaluation, so distribution matching alone suffices — no separate consistency-distillation stage.
- Scale-wise post-training extends the recipe to 1664×960 with two low-resolution and two high-resolution generator calls per block, still four generator calls in total.
Two roles of causal context in post-training. (a) CSF exploits information asymmetry in causal histories for self-supervised representation learning. (b) Mismatched (top) and aligned (bottom) contexts across the generator, real score, and fake score.
Bold = best, underline = second best among the 4-step rows.
| Model | Causal | Steps | VQ ↑ | MQ ↑ | AQ ↑ | CLIP ↑ | IB-AV ↑ | Javis ↑ | DeSync ↓ |
|---|---|---|---|---|---|---|---|---|---|
| LTX-2 Base | ✗ | 40 | 1.884 | 0.566 | 4.986 | 0.311 | 0.239 | 0.200 | 0.608 |
| AR Teacher (CSF) | ✓ | 40 | 2.304 | 0.857 | 4.565 | 0.316 | 0.202 | 0.162 | 0.746 |
| OmniForcing | ✓ | 4 | 1.807 | 0.699 | 4.718 | 0.303 | 0.163 | 0.124 | 0.745 |
| Salt++ (TF-dCM route) | ✓ | 4 | 2.013 | 0.826 | 4.976 | 0.313 | 0.229 | 0.185 | 0.710 |
| Salt++ | ✓ | 4 | 2.838 | 1.010 | 4.991 | 0.316 | 0.184 | 0.146 | 0.759 |
| Model | Causal | Steps | Aesthetic ↑ | Imaging ↑ | Subject Consistency ↑ | Background Consistency ↑ |
|---|---|---|---|---|---|---|
| LTX-2 Base | ✗ | 40 | 53.89 | 67.06 | 95.94 | 95.63 |
| OmniForcing | ✓ | 4 | 56.74 | 68.41 | 96.62 | 95.32 |
| Salt++ (TF-dCM route) | ✓ | 4 | 53.65 | 62.87 | 96.58 | 96.05 |
| Salt++ | ✓ | 4 | 55.43 | 69.98 | 97.01 | 96.24 |
Bold = best across all methods.
| Model | Causal | Steps | VQ ↑ | MQ ↑ | AQ ↑ | CLIP ↑ | IB-AV ↑ | Javis ↑ | DeSync ↓ |
|---|---|---|---|---|---|---|---|---|---|
| LTX-2 Base | ✗ | 40 + 3 | 2.231 | 0.607 | 4.867 | 0.311 | 0.169 | 0.145 | 0.658 |
| OmniForcing | ✓ | 4 | 2.298 | 0.832 | 4.721 | 0.293 | 0.165 | 0.128 | 0.755 |
| Salt++ | ✓ | 4 | 2.730 | 0.957 | 5.113 | 0.318 | 0.193 | 0.157 | 0.768 |
Scale-wise post-training at 1664×960, using two low-resolution and two high-resolution generator calls per block.
If you find this work useful, please cite:
@misc{ge2026saltpp,
title={Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation},
author={Ge, Xingtong and Wang, Yutong and Zhu, Lunjie and Lin, Haitao and Lin, Fangyu and Huang, Yushi and Zhang, Xin and Zhang, Yi and Liu, Yu and Zhang, Jun},
year={2026},
eprint={2609.36995},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.36995}
}This work builds on LTX-2 as the bidirectional audio–video foundation model, follows the causal streaming setup of OmniForcing, and evaluates with JavisBench and VBench. Please also cite the original projects when using their components.



