Skip to content
#

dflash

Here are 29 public repositories matching this topic...

Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-context draft acceptance on DGX Spark. 6 HF variants (BF16/NVFP4/MTP/MTP-XS), docker-compose, and QuickStart.

  • Updated Jul 3, 2026
  • Python
ChaosEngineAI

Local AI workstation — discover, run, chat, benchmark, and generate images from open-weight models. DFlash/DDTree speculative decoding, TurboQuant & TriAttention cache compression strategies, MLX + llama.cpp + vLLM + MTPLX backends.

  • Updated Aug 1, 2026
  • Python

Qwen3.8-27B (Unsloth UD-Q4_K_XL GGUF) served at 37.4 tok/s decode over a 226,048-token context on a single NVIDIA L4 24 GB: llama.cpp with DFlash 2 block-diffusion speculative decoding, a one-line CUDA kernel-routing patch worth 16%, one-command GCP provisioning, quality verification, and a measured account of where the ceiling is.

  • Updated Aug 20, 2026
  • Shell

Improve this page

Add a description, image, and links to the dflash topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the dflash topic, visit your repo's landing page and select "manage topics."

Learn more