Skip to content

Log the recurrent-state slot allocation at load; document max_batch_size for recurrent models - #467

Open
matthematics1137 wants to merge 2 commits into
theroyallab:mainfrom
matthematics1137:recurrent-batch-slot-cost
Open

matthematics1137 wants to merge 2 commits into
theroyallab:mainfrom
matthematics1137:recurrent-batch-slot-cost

Conversation

@matthematics1137

@matthematics1137 matthematics1137 commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Is your pull request related to a problem? Please describe.

Recurrent state slots can be a substantial part of a model's memory budget, but
the load log does not show their cost. This came out of our investigation of
drafted inference on a 12 GB card: reducing max_batch_size was an important
fitting control, and its effect was not obvious.

Why should this feature be added?

Add one INFO line after creating the main cache, using the recurrent layers'
storage_size() values. It reports planned total storage, storage per batch
slot, and max_history, with a single-user hint when there is more than one
slot. It is silent without recurrent layers or history. Allocation behavior
and defaults are unchanged.

The config/schema/sample docs explain the per-slot cost and correct the
transformer batch-size default from 32 to the 128 used by the backend.

Examples

For 48 GDN layers with 48 value heads of dimension 128, four slots and four
history tokens, the shape-only test produces:

Recurrent state storage: 4 slots, 2910 MiB total (728 MiB per slot), max_history: 4. Single-user setups can set max_batch_size: 1.

This is calculated state storage, not a measurement of total process VRAM.
It excludes weights, paged KV cache and other buffers. History layouts differ
by architecture, and some state tensors are CPU-side. The revised wording
deliberately does not describe all recurrent models as keeping full copies
of each checkpoint in VRAM.

Additional context

Updated against main at 816c321, retaining the checkpoint controls from #468
and resolving the overlapping documentation hunk.

Validation:

  • Six new CPU tests call the logging method with real GDN, sliding-attention,
    and PLE state objects on the meta device, covering multi/single slots,
    no-history/no-layer cases, and architecture-specific storage accounting.
  • python -m pytest -q tests/test_*.py: 197 passed, using exllamav3 Python
    sources at 12414d0 with CUDA devices hidden.
  • Ruff 0.11.10: full-project checks pass; all 106 Python files are formatted.
  • Generated max_batch_size sample comments match the schema; its default
    remains None. No live server/GPU load test was run.

AI assistance: the original patch was prepared with Claude Code; this update,
tests, and revised description were prepared with Codex on my behalf.

For recurrent models with a draft model, every batch slot holds
draft_num_tokens + 1 copies of the recurrent state, allocated on the GPU
at load (728 MiB per slot for a 27B hybrid at draft length 4, 2.9 GiB at
the default max_batch_size of 4). Log the size once the cache is built and
say in the max_batch_size docs that single-user recurrent setups should
set 1. The default is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Correct the transformer batch-size default in the docs. Report layer storage without assuming full checkpoint copies or GPU residency for every architecture. Add shape-based CPU tests and retain the merged checkpoint controls.

Assisted-by: Codex

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant