Optimized SGLang runtime for Qwen3.8-27B FP8 with DFlash2 and Qwen3.8 Flash-Next NVFP4 with FR-Spec on one NVIDIA RTX PRO 6000 Blackwell 96 GB GPU (SM120): 524K context, HiCache and NIXL.
-
Updated
Sep 19, 2026 - Python
Optimized SGLang runtime for Qwen3.8-27B FP8 with DFlash2 and Qwen3.8 Flash-Next NVFP4 with FR-Spec on one NVIDIA RTX PRO 6000 Blackwell 96 GB GPU (SM120): 524K context, HiCache and NIXL.
A policy layer above transport for KV movement, workload-aware admissibility, and explainable routing in disaggregated inference.
Open-source Python control plane for LLM inference: launch, route, scale, and observe disaggregated serving on infrastructure you control.
Measuring the KV-transfer tax of disaggregated prefill/decode LLM serving (vLLM + NIXL, 4x A10G): on PCIe-only hardware, disagg loses to plain data-parallel replication — measured, committed, reproducible.
To associate your repository with the nixl topic, visit your repo's landing page and select "manage topics."