Prerequisites
Feature Description
Please implement adaptive-kv-streaming
Motivation
I can run higher quants for both model and KV and have a much higher context with minimal drop in performance.
Possible Implementation
Add a flag to enable or disable it.
Prerequisites
Feature Description
Please implement adaptive-kv-streaming
Motivation
I can run higher quants for both model and KV and have a much higher context with minimal drop in performance.
Possible Implementation
Add a flag to enable or disable it.
I had a slight misunderstanding of fork's mechanism, it can be tested as is with
--kv-stream-stage-mib.Result: it's clear now why there are no KLD benchmarks in the fork repo... it's quite lossy.
Same story every time. In this case, with horrific performance as well. You're literally better off dropping the KV cache to a lower quant.
KLD at -c 131072 / 128k, wikitext-2 test, Qwen3.5-4B-BF16