Skip to content
#

block-diffusion

Here are 6 public repositories matching this topic...

Qwen3.8-27B (Unsloth UD-Q4_K_XL GGUF) served at 37.4 tok/s decode over a 226,048-token context on a single NVIDIA L4 24 GB: llama.cpp with DFlash 2 block-diffusion speculative decoding, a one-line CUDA kernel-routing patch worth 16%, one-command GCP provisioning, quality verification, and a measured account of where the ceiling is.

  • Updated Aug 20, 2026
  • Shell

Production-grade fine-tuning & LoRA toolkit for Chatterbox-Flash zero-shot TTS models. Combines parallel block diffusion and FlashInfer acceleration with smart placeholder vocabulary extension supporting languages. Features offline feature preprocessing, Silero VAD silence trimming, zero-padding leakage prevention, and fast voice cloning for custom

  • Updated Aug 15, 2026
  • Python

Improve this page

Add a description, image, and links to the block-diffusion topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the block-diffusion topic, visit your repo's landing page and select "manage topics."

Learn more