Repository navigation
Add Ability to Reproduce Failed Unit Tests #707
Description
Activity
- addedhelp wantedExtra attention is neededExtra attention is needed
on Nov 6, 2022 Based on some brief research, I've learned that, we should stop setting the
np.random.seedsince (according to this)The problem comes in larger projects or projects with imports that could also set the seed. Using np.random.seed(number) sets what NumPy calls the global random seed, which affects all uses to the np.random.* module. Some imported packages or other scripts could reset the global random seed to another random seed with np.random.seed(another_number), which may lead to undesirable changes to your output and your results becoming unreproducible. For the most part, you will only need to ensure you use the same random numbers for specific parts of your code (like tests or functions).
So, if we have multi-threading when running our tests, each thread could potentially be setting the random global random seed at the same time and therefore it can change the arrays that are produced. Instead, within each test file, we should
- Generate a random integer,
test_state(essentially a seed) - Create a (local) pesudo random number generator,
rng, based on thetest_state - Use the random number generator,
rng, everywhere within this test file where a random number or array is needed - Pass the
test_stateinto every test function so that it will be written/recorded when a test fails
Untested Example
# In some test file import numpy as np test_state = np.random.randint(1_000_000) rng = np.random.default_rng(test_state) test_data = [ ( np.array([9, 8100, -60, 7], dtype=np.float64), np.array([584, -11, 23, 79, 1001, 0, -19], dtype=np.float64), ), ( rng.random.uniform(-1000, 1000, [8]).astype(np.float64), rng.random.uniform(-1000, 1000, [64]).astype(np.float64), ), ] @pytest.mark.parametrize("state", test_state) @pytest.mark.parametrize("Q, T", test_data) def test_compute_mean_std_multidimensional(state, Q, T): m = Q.shape[0] Q = np.array([Q, rng.random.uniform(-1000, 1000, [Q.shape[0]])]) T = np.array([T, T, rng.random.uniform(-1000, 1000, [T.shape[0]])]) ref_μ_Q, ref_σ_Q = naive_compute_mean_std_multidimensional(Q, m) ref_M_T, ref_Σ_T = naive_compute_mean_std_multidimensional(T, m) comp_μ_Q, comp_σ_Q = core.compute_mean_std(Q, m) comp_M_T, comp_Σ_T = core.compute_mean_std(T, m) npt.assert_almost_equal(ref_μ_Q, comp_μ_Q) npt.assert_almost_equal(ref_σ_Q, comp_σ_Q) npt.assert_almost_equal(ref_M_T, comp_M_T) npt.assert_almost_equal(ref_Σ_T, comp_Σ_T) @pytest.mark.parametrize("state", test_state) @pytest.mark.parametrize("Q, T", test_data) def test_njit_sliding_dot_product(state, Q, T): ref_mp = naive_rolling_window_dot_product(Q, T) comp_mp = core._sliding_dot_product(Q, T) npt.assert_almost_equal(ref_mp, comp_mp)Maybe something like this? Note that the random state is set once at the beginning of the file and then it is iterated upon to generate the necessary data but the initial state is never changed. I believe/hypothesize that one would need to explicitly set the random state to the
test_statein order to get reproduce the failed test.- Generate a random integer,
Perhaps, instead of:
test_state = np.random.randint(1_000_000)We can use the built in
$RANDOMinteger value from the shell, which returns a value between [0, 32767].So, inside
test.shexport STUMPY_SEED=$RANDOMand then this can be referenced inside of each testfile via:
test_state= os.getenv('STUMPY_SEED') rng = np.random.default_rng(test_state)and then see above
You can count the number of tests being collected for each testfile via:
pytest --collect-only -rsx -W ignore::RuntimeWarning -W ignore::DeprecationWarning -W ignore::UserWarning $testfile | grep "<Function" -cThis code demonstrates what the state is and how it can be restored:
import numpy as np rng = np.random.default_rng(12345) print(rng.bit_generator.state) print(rng.integers(0, 300, size=3)) print(rng.integers(0, 300, size=3)) print(rng.integers(0, 300, size=3)) rng = np.random.default_rng() rng.bit_generator.state = {'bit_generator': 'PCG64', 'state': {'state': 33261208707367790463622745601869196757, 'inc': 268209174141567072605526753992732310247}, 'has_uint32': 0, 'uinteger': 0} print(rng.integers(0, 300, size=3)) print(rng.integers(0, 300, size=3)) print(rng.integers(0, 300, size=3))So
print(rng.bit_generator.state)will print the exact dictionary:{'bit_generator': 'PCG64', 'state': {'state': 33261208707367790463622745601869196757, 'inc': 268209174141567072605526753992732310247}, 'has_uint32': 0, 'uinteger': 0}This
dictcan be set as the (bit generator) state as shown above. So, with this, you can:- Create an
rng - Call the
rngthree times - And note/save the
stateof the generator right before it executes another "draw" fromrng
import numpy as np rng = np.random.default_rng(12345) rng.integers(0, 300, size=3) rng.integers(0, 300, size=3) rng.integers(0, 300, size=3) print(rng.bit_generator.state) print(rng.integers(0, 300, size=3))Produces:
{'bit_generator': 'PCG64', 'state': {'state': 124332563986525153014824637231659249044, 'inc': 268209174141567072605526753992732310247}, 'has_uint32': 1, 'uinteger': 1679802728} [117 251 99]and then you can set this state and draw:
rng = np.random.default_rng() rng.bit_generator.state = {'bit_generator': 'PCG64', 'state': {'state': 124332563986525153014824637231659249044, 'inc': 268209174141567072605526753992732310247}, 'has_uint32': 1, 'uinteger': 1679802728} print(rng.integers(0, 300, size=3))and this would produce the same array:
[117 251 99]- Create an
Obviously, the same thing can be done to generate an array of floats:
rng = np.random.default_rng(12345) rng.uniform(0, 300, size=3) rng.uniform(0, 300, size=3) rng.uniform(0, 300, size=3) print(rng.bit_generator.state) print(rng.uniform(0, 300, size=3))and then restore the state:
# Restore state rng = np.random.default_rng() rng.bit_generator.state = {'bit_generator': 'PCG64', 'state': {'state': 151817738446295939502568876194759307672, 'inc': 268209174141567072605526753992732310247}, 'has_uint32': 0, 'uinteger': 0} print(rng.uniform(0, 300, size=3))- modified the milestones: Python 1.14.0 Release, Python 1.15.0 Release, Python 1.16.0 Release
on Feb 1, 2026 Hi @seanlaw,
I've implemented a solution for this issue in PR #1155. Here's a summary of the approach taken:
Problem: When a pytest test fails due to a random seed-related flakiness, it is impossible to reproduce the failure because no seed is recorded.
Solution:
- Added a
STUMPY_SEEDenvironment variable that is automatically set at the start of every test session (viaconftest.py). If not already set, a random integer seed is generated. -
- The seed is printed at pytest startup via a
pytest_configurehook so it always appears in the test output.
- The seed is printed at pytest startup via a
-
- The global
config.RNG(NumPy Generator) is seeded withSTUMPY_SEEDat session start, making all random draws deterministic and reproducible.
- The global
-
- To reproduce a failing test run, users simply re-run with the same seed:
STUMPY_SEED=<seed> pytest tests/
- To reproduce a failing test run, users simply re-run with the same seed:
-
- Added
tests/test_seed.pyto verify the environment variable is correctly set and produces deterministic random sequences.
This approach is minimal, non-breaking, and does not require any changes to existing test logic.
- Added
- Added a
This approach is minimal, non-breaking, and does not require any changes to existing test logic.
@Vansh-Sharmaa Unfortunately, this is not true and will break when you use the same seed/state across different versions of NumPy
Fixed in #1154
Currently, when there are precision-related issues, it is not obvious what random seed produced those failed unit/coverage tests. It would be nice to figure out a way to log this information in PyTest so that we can easily reproduce the failure. It'll require some trial-and-error and may not be easy/possible.