Skip to content

Preserve backbone precision and extend smoke coverage - #4

Open
AngeLouCN wants to merge 2 commits into
mainfrom
codex/infra-execution-smoke-pr
Open

AngeLouCN wants to merge 2 commits into
mainfrom
codex/infra-execution-smoke-pr

Conversation

@AngeLouCN

@AngeLouCN AngeLouCN commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Ordinary decision scoring retains unused generation caches and casts the entire sequence of hidden states to FP32, although the pointer head reads only selected vectors. Route text, media and prefix execution through BackboneAdapter.forward, retain the backbone's native hidden-state dtype, and cast only the vectors needed by the pointer head. Ordinary scoring defaults to use_cache=False; explicit prefix reuse continues to request a cache.

Extend the backbone smoke CLI with separate frozen-weight precision and decision-mode options. Handle tokenizers that reuse existing delimiter tokens without requiring jevany_token_schema metadata, reject incompatible direct-token training flags, and use a tiny GLM fixture with dimensions previously verified for BF16 SDPA on A100.

Validation on the restored two-commit branch:

  • CPU regression: python -m pytest tests -m 'not server' -q -ra --strict-config --strict-markers — 432 passed, 37 skipped, 7 deselected.
  • python -m build --no-isolation — source distribution and wheel both built successfully.
  • Optional CUDA checks are skipped in the CPU suite; no new GPU run was performed for this revision.

Base: main at 625aed9.

Candidate-token latency optimization and its accuracy/latency results are reviewed separately in #9.

@AngeLouCN
AngeLouCN force-pushed the codex/infra-execution-smoke-pr branch from cd4e17b to 3d48583 Compare September 30, 2026 20:38
@ZhiningLiu1998
ZhiningLiu1998 marked this pull request as ready for review October 1, 2026 01:02
@AngeLouCN
AngeLouCN force-pushed the codex/infra-execution-smoke-pr branch from 3d48583 to 77db932 Compare October 1, 2026 06:41
@AngeLouCN AngeLouCN changed the title Preserve backbone precision and extend smoke coverage Optimize candidate-token readout and preserve backbone execution Oct 1, 2026
Route text, media and prefix forwards through the existing adapter contract. Keep sequence hidden states in their native dtype and cast only the pointer readout vectors. Validate cache behavior and exact BF16 output/gradient parity; the CPU regression suite passes 303 tests.
Handle existing delimiter token schemas, expose frozen-weight precision and direct-token smoke, and reject incompatible direct-mode flags. Use tiny GLM dimensions validated for BF16 SDPA training and checkpoint reload on A100.
@AngeLouCN AngeLouCN changed the title Optimize candidate-token readout and preserve backbone execution Preserve backbone precision and extend smoke coverage Oct 1, 2026
@AngeLouCN
AngeLouCN force-pushed the codex/infra-execution-smoke-pr branch from 77db932 to 95139d2 Compare October 1, 2026 07:00

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant