Repository navigation
Support num_simulations=0 in all policies. - #123
Open
carlosgmartin wants to merge 1 commit into
Open
carlosgmartin wants to merge 1 commit into
carlosgmartin wants to merge 1 commit into
Conversation
With no simulations, the search loop is skipped, and the MuZero and Stochastic MuZero policies act on the root prior, which excludes invalid actions, instead of the uniform visit probabilities of an unvisited root.
carlosgmartin
force-pushed
the
num_simulations_zero
branch
from
October 10, 2026 03:06
2738a82 to
a34ddc0
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
With no simulations, a policy should act based on the prior. Currently, it doesn't:
gumbel_muzero_policyfails withIndexError: index is out of bounds for axis 1 with size 0.muzero_policyandstochastic_muzero_policyact uniformly at random over all actions, including invalid ones.The cause and fix for each is described below.
gumbel_muzero_policysearchtraces the simulation loop body even when it runs zero times. Tracing reachesgumbel_muzero_root_action_selection, which indexes a table with one column per simulation (action_selection.py:147), so it fails with:Fix: skip the loop in
searchwhennum_simulations == 0. The existing post-search code then acts on the prior: the action isargmax(gumbel + logits)over valid actions, andaction_weightsis the prior restricted to valid actions.The behavior for
num_simulations >= 1is unchanged.muzero_policyandstochastic_muzero_policyaction_weightscomes fromTree.summary().visit_probs, which falls back to1 / num_actionswhen the root has no visits (tree.py:110). This ignores both the prior andinvalid_actions.Fix: when
num_simulations == 0, use the softmax of the root's prior logits, which already include Dirichlet noise and exclude invalid actions.The behavior for
num_simulations >= 1is unchanged.Tests
Added
test_gumbel_muzero_policy_without_simulationsandtest_muzero_policies_without_simulations. Both fail without this change and pass with it. All existing tests pass.