-
Notifications
You must be signed in to change notification settings - Fork 1.1k
Pull requests: FlashML-org/FreeToken
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
qwen4_exp: serve the block-FP8 dense projections natively (+25% decode)
#392
opened Sep 5, 2026 by
gberasmus87
Loading…
fix(models): better support for mixed-precision compressed-tensors NVFP4
#390
opened Sep 5, 2026 by
Sam-Izdat
Loading…
fix(scheduler): reserve paged KV at allocation granularity
#367
opened Sep 3, 2026 by
taking-lying-flat
Contributor
Loading…
feat(kvcache): store the KV cache as fp8 e4m3 codes (--kv-cache-dtype…
#354
opened Sep 2, 2026 by
ArqAlice
Loading…
fix(server): stop gemma4 tool-call markers leaking into content
#346
opened Sep 2, 2026 by
vianbas
Loading…
fix(kernel): exact single-launch triton top-k/top-p sampling
#345
opened Sep 2, 2026 by
jason-fxz
Collaborator
Loading…
qwen3_5_moe: run lm_head on sampled rows only (fixes 32 GB first-prefill OOM)
#342
opened Sep 2, 2026 by
chrisqianz
Loading…
feat(bench): add reproducible serving performance harness
#341
opened Sep 2, 2026 by
tuxevil
Loading…
fix(engine): reserve the pool's dummy page in the --moe-cache-auto KV floor
#340
opened Sep 2, 2026 by
dejay2
Loading…
fix(kernels): drop the D2H sync from the varlen GDN/KDA prefill conv
#339
opened Sep 2, 2026 by
dejay2
Loading…
perf(ple): fuse the n-gram row-id hash into one Triton kernel
#338
opened Sep 2, 2026 by
dejay2
Loading…
fix(memory): honor cgroup memory limits when sizing expert banks
#334
opened Sep 1, 2026 by
Avicennasis
Loading…
qwen4_exp: load modelopt MIXED_PRECISION (NVFP4 experts + block-FP8 dense) checkpoints
#320
opened Sep 1, 2026 by
gberasmus87
Loading…
Previous Next
ProTip!
Follow long discussions with comments:>50.