Skip to content

Keep half-precision Q-value normalization finite - #126

Open
sylvesterkaczmarek wants to merge 1 commit into
google-deepmind:mainfrom
sylvesterkaczmarek:fix/half-precision-q-normalization-20261005
Open

sylvesterkaczmarek wants to merge 1 commit into
google-deepmind:mainfrom
sylvesterkaczmarek:fix/half-precision-q-normalization-20261005

Conversation

@sylvesterkaczmarek

Copy link
Copy Markdown

Summary

Fixes #125.

Promote Q-range calculations and supplied bounds to at least float32, then restore the established result dtype. This avoids a float16 zero-epsilon division and overflowing differences between finite bounds. The public half-precision MuZero regression now selects the same action as float32.

Unvisited-action semantics, explicit wider dtypes and ordinary float32 results are preserved. The mixed-value calculation from PR #120 is untouched. Already-overflowed Q-values and explicit zero denominators are outside the correction.

Validation

python -m pytest -q --pyargs mctx (from a temporary child of the repository)

39 tests passed, covering the complete MCTX test suite. Fifteen new tests check equal/wide ranges, float16/bfloat16/float32 inputs, explicit half and wider bounds, JIT, unvisited actions, and a real 16-simulation MuZero search. Existing JSON tree fixtures also pass.

The full suite ran from an isolated temporary child directory of the checkout using --pyargs mctx, as required by the repository's relative tree-fixture paths. No GPU/TPU or trained-model evaluation was run. No changes to general search, visit accounting or action masking are included.

Tested locally on macOS CPU using real module imports. New test formatting, scoped static checks, syntax checks and git diff --check pass. No dependency or workflow changes.

Signed-off-by: Sylvester Kaczmarek <16242628+sylvesterkaczmarek@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Half-precision Q normalization produces NaNs and selects the wrong MuZero action

1 participant