Skip to content

Add Ability to Reproduce Failed Unit Tests #707

Description

@seanlaw

Currently, when there are precision-related issues, it is not obvious what random seed produced those failed unit/coverage tests. It would be nice to figure out a way to log this information in PyTest so that we can easily reproduce the failure. It'll require some trial-and-error and may not be easy/possible.

Activity

  1. seanlaw commented on May 23, 2023

    @seanlaw
    ContributorAuthor

    Based on some brief research, I've learned that, we should stop setting the np.random.seed since (according to this)

    The problem comes in larger projects or projects with imports that could also set the seed. Using np.random.seed(number) sets what NumPy calls the global random seed, which affects all uses to the np.random.* module. Some imported packages or other scripts could reset the global random seed to another random seed with np.random.seed(another_number), which may lead to undesirable changes to your output and your results becoming unreproducible. For the most part, you will only need to ensure you use the same random numbers for specific parts of your code (like tests or functions).

    So, if we have multi-threading when running our tests, each thread could potentially be setting the random global random seed at the same time and therefore it can change the arrays that are produced. Instead, within each test file, we should

    1. Generate a random integer, test_state (essentially a seed)
    2. Create a (local) pesudo random number generator, rng, based on the test_state
    3. Use the random number generator, rng, everywhere within this test file where a random number or array is needed
    4. Pass the test_state into every test function so that it will be written/recorded when a test fails

    Untested Example

    # In some test file
    import numpy as np
    test_state = np.random.randint(1_000_000)
    rng = np.random.default_rng(test_state)
    
    test_data = [
        (
            np.array([9, 8100, -60, 7], dtype=np.float64),
            np.array([584, -11, 23, 79, 1001, 0, -19], dtype=np.float64),
        ),
        (
            rng.random.uniform(-1000, 1000, [8]).astype(np.float64),
            rng.random.uniform(-1000, 1000, [64]).astype(np.float64),
        ),
    ]
    
    
    @pytest.mark.parametrize("state", test_state)
    @pytest.mark.parametrize("Q, T", test_data)
    def test_compute_mean_std_multidimensional(state, Q, T):
        m = Q.shape[0]
    
        Q = np.array([Q, rng.random.uniform(-1000, 1000, [Q.shape[0]])])
        T = np.array([T, T, rng.random.uniform(-1000, 1000, [T.shape[0]])])
    
        ref_μ_Q, ref_σ_Q = naive_compute_mean_std_multidimensional(Q, m)
        ref_M_T, ref_Σ_T = naive_compute_mean_std_multidimensional(T, m)
        comp_μ_Q, comp_σ_Q = core.compute_mean_std(Q, m)
        comp_M_T, comp_Σ_T = core.compute_mean_std(T, m)
    
        npt.assert_almost_equal(ref_μ_Q, comp_μ_Q)
        npt.assert_almost_equal(ref_σ_Q, comp_σ_Q)
        npt.assert_almost_equal(ref_M_T, comp_M_T)
        npt.assert_almost_equal(ref_Σ_T, comp_Σ_T)
    
    @pytest.mark.parametrize("state", test_state)
    @pytest.mark.parametrize("Q, T", test_data)
    def test_njit_sliding_dot_product(state, Q, T):
        ref_mp = naive_rolling_window_dot_product(Q, T)
        comp_mp = core._sliding_dot_product(Q, T)
        npt.assert_almost_equal(ref_mp, comp_mp)
    

    Maybe something like this? Note that the random state is set once at the beginning of the file and then it is iterated upon to generate the necessary data but the initial state is never changed. I believe/hypothesize that one would need to explicitly set the random state to the test_state in order to get reproduce the failed test.

  2. seanlaw commented on Jan 1, 2026

    @seanlaw
    ContributorAuthor

    Perhaps, instead of:

    test_state = np.random.randint(1_000_000)
    

    We can use the built in $RANDOM integer value from the shell, which returns a value between [0, 32767].

    So, inside test.sh

    export STUMPY_SEED=$RANDOM
    

    and then this can be referenced inside of each testfile via:

    test_state= os.getenv('STUMPY_SEED')
    rng = np.random.default_rng(test_state)
    

    and then see above

  3. seanlaw commented on Jan 2, 2026

    @seanlaw
    ContributorAuthor

    You can count the number of tests being collected for each testfile via:

    pytest --collect-only -rsx -W ignore::RuntimeWarning -W ignore::DeprecationWarning -W ignore::UserWarning $testfile | grep "<Function" -c
    
  4. seanlaw commented on Jan 2, 2026

    @seanlaw
    ContributorAuthor

    This code demonstrates what the state is and how it can be restored:

    import numpy as np
    
    rng = np.random.default_rng(12345)
    print(rng.bit_generator.state)
    print(rng.integers(0, 300, size=3))
    print(rng.integers(0, 300, size=3))
    print(rng.integers(0, 300, size=3))
    
    
    rng = np.random.default_rng()
    rng.bit_generator.state = {'bit_generator': 'PCG64', 'state': {'state': 33261208707367790463622745601869196757, 'inc': 268209174141567072605526753992732310247}, 'has_uint32': 0, 'uinteger': 0}
    print(rng.integers(0, 300, size=3))
    print(rng.integers(0, 300, size=3))
    print(rng.integers(0, 300, size=3))
    

    So print(rng.bit_generator.state) will print the exact dictionary:

    {'bit_generator': 'PCG64', 'state': {'state': 33261208707367790463622745601869196757, 'inc': 268209174141567072605526753992732310247}, 'has_uint32': 0, 'uinteger': 0}
    

    This dict can be set as the (bit generator) state as shown above. So, with this, you can:

    1. Create an rng
    2. Call the rng three times
    3. And note/save the state of the generator right before it executes another "draw" from rng
    import numpy as np
    
    rng = np.random.default_rng(12345)
    rng.integers(0, 300, size=3)
    rng.integers(0, 300, size=3)
    rng.integers(0, 300, size=3)
    
    print(rng.bit_generator.state)
    print(rng.integers(0, 300, size=3))
    

    Produces:

    {'bit_generator': 'PCG64', 'state': {'state': 124332563986525153014824637231659249044, 'inc': 268209174141567072605526753992732310247}, 'has_uint32': 1, 'uinteger': 1679802728}
    [117 251  99]
    

    and then you can set this state and draw:

    rng = np.random.default_rng()
    rng.bit_generator.state = {'bit_generator': 'PCG64', 'state': {'state': 124332563986525153014824637231659249044, 'inc': 268209174141567072605526753992732310247}, 'has_uint32': 1, 'uinteger': 1679802728}
    print(rng.integers(0, 300, size=3))
    

    and this would produce the same array:

    [117 251  99]
    
  5. seanlaw commented on Jan 2, 2026

    @seanlaw
    ContributorAuthor

    Obviously, the same thing can be done to generate an array of floats:

    rng = np.random.default_rng(12345)
    rng.uniform(0, 300, size=3)
    rng.uniform(0, 300, size=3)
    rng.uniform(0, 300, size=3)
    
    print(rng.bit_generator.state)
    print(rng.uniform(0, 300, size=3))
    

    and then restore the state:

    # Restore state
    
    rng = np.random.default_rng()
    rng.bit_generator.state = {'bit_generator': 'PCG64', 'state': {'state': 151817738446295939502568876194759307672, 'inc': 268209174141567072605526753992732310247}, 'has_uint32': 0, 'uinteger': 0}
    print(rng.uniform(0, 300, size=3))
    
    
  6. self-assigned this
    on Feb 14, 2026
  7. Vansh-Sharmaa commented on Jul 18, 2026

    @Vansh-Sharmaa

    Hi @seanlaw,

    I've implemented a solution for this issue in PR #1155. Here's a summary of the approach taken:

    Problem: When a pytest test fails due to a random seed-related flakiness, it is impossible to reproduce the failure because no seed is recorded.

    Solution:

    1. Added a STUMPY_SEED environment variable that is automatically set at the start of every test session (via conftest.py). If not already set, a random integer seed is generated.
      1. The seed is printed at pytest startup via a pytest_configure hook so it always appears in the test output.
      1. The global config.RNG (NumPy Generator) is seeded with STUMPY_SEED at session start, making all random draws deterministic and reproducible.
      1. To reproduce a failing test run, users simply re-run with the same seed: STUMPY_SEED=<seed> pytest tests/
      1. Added tests/test_seed.py to verify the environment variable is correctly set and produces deterministic random sequences.
        This approach is minimal, non-breaking, and does not require any changes to existing test logic.
  8. seanlaw commented on Jul 18, 2026

    @seanlaw
    ContributorAuthor

    This approach is minimal, non-breaking, and does not require any changes to existing test logic.

    @Vansh-Sharmaa Unfortunately, this is not true and will break when you use the same seed/state across different versions of NumPy

  9. seanlaw commented on Jul 21, 2026

    @seanlaw
    ContributorAuthor

    Fixed in #1154

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions