Skip to content

feat(repo): test xstate's guard evaluation as a conformance spec and give checkStateIn a data-last form - #81

Open
systemfsoftware-maker wants to merge 17 commits into
mainfrom
xs/xstate-guards
Open

systemfsoftware-maker wants to merge 17 commits into
mainfrom
xs/xstate-guards

Conversation

@systemfsoftware-maker

@systemfsoftware-maker systemfsoftware-maker commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Capability 6 (guards) of @systemfsoftware/xstate: guard evaluation gets one conformance spec over a hand-written admission model, and its decisions move into src/transitionGuards.ts, which joins the mutate set.

What changes

  • src/transitionGuards.ts (new, mutated, linted, tsgo-enrolled). It holds the pure decisions:
    • checkStateIn, moved from utils.ts, now with a data-last form checkStateIn(stateValue)(snapshot) dispatched by argument count;
    • candidate admission, in order: event pattern, session matcher, internal guard, then the transition function;
    • the selection rule for a transition function's outcome (a returned value admits unless it is undefined; an enqueued effect admits and is not reusable);
    • first-admitted-candidate in document order.
  • The shell keeps the user-code calls. stateUtils.ts still runs the transition function and catches the effect signal, then hands the tagged outcome to the core. StateNode.next and the eventless walk use the core's decisions. The evaluateCandidate pass-through, utils.ts's matchesEvent and checkStateIn, and stateUtils.ts's isStateId re-export are gone.
  • Tests. tests/guards.conformance.test.ts (13 cases) replaces test/guards.test.ts and test/stateIn.test.ts; both leave the two unguarded lists.

The spec

A Conformance.sequential check over generated machines on one fixed parallel topology (regions L and R; L holds compound p with c1 and c2, plus q; R holds r1 and r2; context { n }; events GO{k} and TICK). Each node's on slot for each event is generated:

  • Absent, a string target (GO: '#q'; setup's types forbid the bare string, so its machines use { target }), an event pattern { target, matches: { k } }, or a transition function whose predicate is Always, Never, a context threshold, a payload threshold, a guards.atLeast source, or matchesState(path, value), and whose outcome is a move, a targetless stay (optionally bumping the context), or an emitted effect;
  • eventless slots on c2 and r2;
  • three factories (createMachine, setup, .provide, where provide's guard wins) and two drivers (actor with can and on('effect'); pure initialTransition/transition).

The model predicts, per step, the state value, context.n, can(event), the emitted effects and the active ids read through checkStateIn in both call forms. A liveness ledger requires each slot kind on every side its semantics allow (Target admitted, Absent rejected, Pattern and Fn both admitted and rejected), each fallback depth (p, the region, the root), a cross-region interim eventless step, every factory and driver, and checkStateIn true and false in both forms. Two planted subjects (a rejection written as {}, an effect outcome that enqueues nothing) must be judged model-diverged.

Cases:

  • A custom guard with parameters admits only when the context and the event together clear the threshold
  • A declared target admits the event and moves the actor to it
  • A guard in one region reads the interim value another region reached in the same macrostep
  • A guard source receives exactly the arguments the transition function passes
  • A guard source the machine does not implement leaves the actor in error status with a TypeError naming it
  • A rejection that returns an empty object consumes the event and is caught as a model divergence
  • A transition function that throws leaves the actor in error status with that error
  • A zero-parameter guard source is called with no arguments
  • An effect outcome that enqueues nothing falls through and is caught as a model divergence
  • Every machine and event sequence the model draws from seed 1 admits exactly what the published guards admit
  • Every machine and event sequence the model draws from seed 2 admits exactly what the published guards admit
  • Every machine and event sequence the model draws from seed 3 admits exactly what the published guards admit
  • checkStateIn answers path strings and state values like snapshot.matches, in both forms

Behaviour-to-test map (CONST-T9)

Replaces packages/xstate/test/guards.test.ts and packages/xstate/test/stateIn.test.ts.
Covering scenarios are the Gherkin titles in packages/xstate/tests/guards.conformance.test.ts:

  • OUTLINE — Every machine and event sequence the model draws from seed <seed> admits exactly what the published guards admit (scenario outline, seeds 1/2/3).
  • STATEIN — checkStateIn answers path strings and state values like snapshot.matches, in both forms.
  • TARGET — A declared target admits the event and moves the actor to it.
  • CUSTOM — A custom guard with parameters admits only when the context and the event together clear the threshold.
  • INTERIM — A guard in one region reads the interim value another region reached in the same macrostep.
  • MISSING — A guard source the machine does not implement leaves the actor in error status with a TypeError naming it.
  • THROW — A transition function that throws leaves the actor in error status with that error.
  • ARGS — A guard source receives exactly the arguments the transition function passes.
  • ZERO — A zero-parameter guard source is called with no arguments.
  • DIVERGE_EMPTY — A rejection that returns an empty object consumes the event and is caught as a model divergence.
  • DIVERGE_EFFECT — An effect outcome that enqueues nothing falls through and is caught as a model divergence.

guards.test.ts (24 it, 2 skipped)

guard conditions › should transition only if condition is met → OUTLINE (+ the model's Fn slot with a ContextAtLeast evaluator admits when the pre-event context clears the threshold, over both drivers and all three factories).
guard conditions › should transition if condition based on event is met → OUTLINE (+ the model's Fn slot with a PayloadAtLeast evaluator on GO admits; Named/Payload admission).
guard conditions › should not transition if condition based on event is not met → OUTLINE (+ the same Fn slot rejects, so the event falls through to the ancestor; ledger rejected.Fn).
guard conditions › should not transition if no condition is met → OUTLINE (+ a Fn slot returning undefined admits nothing; the model predicts the unchanged value and the ledger records the rejection, matching the old test's "no actions fired" assertion).
guard conditions › should work with defined string transitions → OUTLINE (+ Target slots: a plain string target always admits).
guard conditions › should work with guard objects → OUTLINE (+ Fn slots; the old test's guard-object syntax had already been collapsed to inline transition functions, so the behaviour is the generated Fn slots).
guard conditions › should work with defined string transitions (condition not met) → OUTLINE (+ the Fn/Target slots whose condition fails).
guard conditions › should guard against transition → OUTLINE (+ Fn slots whose when is Never/a false predicate never admit — the "guard against" case).
guard conditions › should allow a matching transition → OUTLINE (+ the model's In evaluator, matchesState(path, value) on the pre-event value, admits).
guard conditions › should check guards with interim states → INTERIM (+ a region-R always reads the interim value of region L after L moved in the same macrostep; the old test's A3→A4→A5 eventless chain is the same one-microstep interim step).

[function] guard conditions › should transition only if condition is met → OUTLINE.
[function] guard conditions › should transition if condition based on event is met → OUTLINE.
[function] guard conditions › should not transition if condition based on event is not met → OUTLINE.
[function] guard conditions › should not transition if no condition is met → OUTLINE.
[function] guard conditions › should work with defined string transitions → OUTLINE.
[function] guard conditions › should work with guard objects → OUTLINE.
[function] guard conditions › should work with defined string transitions (condition not met) → OUTLINE.
[function] guard conditions › it.skip should allow a matching transition → OUTLINE (+ same machine as the live guard conditions › should allow a matching transition; skipped in the old suite, the matching behaviour is the model's In evaluator).
[function] guard conditions › it.skip should check guards with interim states → INTERIM (+ same machine as the live guard conditions › should check guards with interim states; skipped in the old suite).

custom guards › should evaluate custom guards → CUSTOM (+ exactly the old machine: guards-style source with count + event.value > 3; the scenario asserts event value 4 reaches active and value 3 stays inactive).
guards - other › should allow for a fallback target to be a simple string → OUTLINE (+ the model's Fn slot with a Go outcome admits and moves to its target, which is the old test's function returning { target: 'c' }; TARGET pins the declared-target form).
guards - unknown references › should throw on a guard reference that is not implemented → MISSING (+ same guards.isRedy call: actor reaches status: 'error' with a TypeError whose message matches /guards[^\s]*isRedy[^\s]* is not a function/).
guards - plain function sources › passes only the caller-supplied params to the source → ARGS (+ same probe: guards.isAbove(context.count, 3) records [5, 3] and admits).
guards - plain function sources › supports zero-param guards → ZERO (+ a zero-parameter source is callable and admits).

stateIn.test.ts (9 it, 1 skipped)

transition "in" check › should transition if string state path matches current state value → STATEIN (+ the hand-written parallel machine answers a full state value { a: 'a1', b: 'b2' } true and a partial { a: 'a2' } false; the outline's In evaluator covers matchesState on the pre-event value).
transition "in" check › should transition if state node ID matches current state value → STATEIN (+ checkStateIn(snapshot, '#b_b2') true, and data-last checkStateIn('#b_b2')(snapshot) true).
transition "in" check › should not transition if string state path does not match current state value → STATEIN (+ checkStateIn(snapshot, '#b_b1') and data-last false for an inactive id).
transition "in" check › should not transition if state value matches current state value → STATEIN (+ the scenario's { a: 'a2' } false case; the old title is misnamed, the old test does transition when the b-region value matches).
transition "in" check › matching should be relative to grandparent (match) → STATEIN (+ checkStateIn(self.getSnapshot(), '#bar1') is an id-active lookup; the scenario answers ids true in both forms).
transition "in" check › matching should be relative to grandparent (no match) → STATEIN (+ the same id lookup false).
transition "in" check › should work to forbid events → OUTLINE (+ the parent node's In guard on { red: 'stop' } and the region fallback are the model's In slot with a fallback to the region/root).
transition "in" check › should be possible to use a referenced stateIn guard → OUTLINE (+ a Named guard source inside a transition function; the model's Named evaluator calls guards.atLeast, and the ledger records a named source asked under each factory).
transition "in" check › it.skip should be possible to check an ID with a path → uncovered: the test was already skipped in the old suite (its stateIn('#b.B1') guard was commented out and replaced by an inline matchesState('#b.B1', value) call), and no live case exercises the compound #id.child path form; the fixed scenario covers ids and plain paths, which is the form the old suite's live tests use.

Fixed scenarios added for behaviour the outline does not reach

  • INTERIM — the old should check guards with interim states multi-region, one-microstep eventless step (the outline's In slots are absolute paths only, so a dedicated scenario pins the interim-value read).
  • CUSTOM — the old custom guards › should evaluate custom guards parameterised source.
  • TARGET — a declared target moving the actor, pinning admission distinct from the generated Target slots.
  • MISSING / THROW / ARGS / ZERO / STATEIN — the error and guard-argument behaviours, each with a hand-written expectation.

Target form note

The generated Target slot emits the published bare-string form (GO: '#q') for the createMachine and provide factories; the setup factory builds the same topology with the object form ({ target: '#q' }) because setup's TransitionConfigOrTarget has no bare-string member and validates on keys against the declared event schemas. Both forms feed the engine's same admission, so the model predicts one behaviour.

API report

etc/xstate.api.md gains one overload; the other diffs in etc/ are chunk-hash renames in ae-forgotten-export comments.

export function checkStateIn(stateValue: StateValue): (snapshot: AnyMachineSnapshot) => boolean;

Changeset .changeset/guards-capability.md (minor) declares it.

Mutated set

packages/xstate/stryker.config.ts adds src/transitionGuards.ts. This PR's own release-gate run is to be the mutation proof: 0 Survived and 0 NoCoverage in that file. No run on this head has finished yet: runs 37956675935 and 37962882414 lost their runner mid-shard (exit 143), and run 37959701505's shard ended with exit 3 during the checker stage.

Enrolment (XS1)

src/transitionGuards.ts and the spec join the lint script, tsconfig.tsgo.json and XS1's enrolled list in packages/AGENTS.md. No disable comments; no casts or any in the spec or its fixtures.

Gates (local, on the head)

  • bin/dprint check, tsc -b, tsc -p tsconfig.test.json --noEmit, pnpm run lint, pnpm run lint:tsgo (14 src files checked, 0 diagnostics), pnpm run build (API check included): all exit 0.
  • Spec, seeds 1, 2 and 3: 13 passed each, identical case names.
  • Full xstate suite, seeds 1, 2 and 3: 137 files, 2334 passed, 23 skipped, 3 todo on each. Before the deletion it was 139 files, 2364 passed, 26 skipped; the two deleted files held 30 passing and 3 skipped cases.
  • No local mutation testing of any kind.

…ts (WIP, parked)

Capability 6 (guards) work in progress, parked by conductor ruling 217
for the release-gate work. src/transitionGuards.ts owns checkStateIn
(now with a data-last form), candidate admission and the
transition-function selection rule over a tagged outcome; stateUtils.ts
and StateNode.ts keep calling user code and delegate the decisions.
The module is enrolled in mutate, the lint script, tsconfig.tsgo.json
and XS1. The guards conformance spec is not finished and the old guard
tests are still in place.
…rked)

Adds src/transitionGuards.ts and the guards conformance spec path to
the lint script, src/transitionGuards.ts to tsconfig.tsgo.json and to
XS1's enrolled list, and regenerates the root API report for
checkStateIn's data-last form. Parked by conductor ruling 217.
…uards.ts

StateMachine.ts imports isStateId from transitionGuards.ts, so the
re-export in stateUtils.ts is gone. The eventless walk calls
admitTransitionCandidate through isAdmitted directly, so the
evaluateCandidate pass-through is gone too. A changeset declares
checkStateIn's data-last form, and the API reports are regenerated.
tests/guards.conformance.test.ts checks guard evaluation against a
hand-written admission model. The model covers transition functions,
string targets, event patterns, guards sources from createMachine, setup
and provide, inStates via matchesState, effect-only outcomes, the walk up
ancestors, parallel regions with a root fallback, and eventless steps
that read the value one region reached earlier in the same macrostep.
Each step also checks checkStateIn, data-first and data-last, against
the active ids.

Two planted subjects show the check rejects a rejection written as {}
and an effect outcome that enqueues nothing. The fixed scenarios cover a
missing guard source, a throwing transition function, the exact guard
arguments, checkStateIn's truth table and the old suite's custom-guard
and interim-state machines.

test/guards.test.ts and test/stateIn.test.ts are deleted and leave both
unguarded lists.
@systemfsoftware-maker systemfsoftware-maker changed the title xs/xstate guards feat(repo): test xstate's guard evaluation as a conformance spec and give checkStateIn a data-last form Oct 9, 2026
Two of three release-gate runs on #81 lost the whole ubuntu-latest VM
(16 GB) mid-shard. Node's default old-space ceiling is about 4 GiB per
isolate, and each test-runner child holds a main isolate plus a claimed
and a spare vitest worker thread, so two runners running a mutant that
allocates without bound can outgrow the VM before the 45 s mutant
timeout settles them.

testRunnerNodeArgs passes --max-old-space-size=2048 to every test-runner
child. The flag is process-wide in V8, so each worker thread gets the
same ceiling. A worker that reaches it is ended by Node; Stryker retries
the mutant twice and then records it as a RuntimeError, so the run
carries on. The full xstate suite passes on one thread with a 512 MiB
ceiling, so 2 GiB leaves room for the coverage instrumentation.
@systemfsoftware-maker
systemfsoftware-maker changed the base branch from main to xs/release-gate-runner-heap October 9, 2026 18:01
@systemfsoftware-maker
systemfsoftware-maker added this pull request to stack #83 October 9, 2026 18:01
On #81 the shard ends with exit 3 about 26 s into mutation testing, and
the job log keeps only the first 1024 characters of the child's stderr,
which are startup lines. The diagnostics re-run then ran for 16 minutes
until the VM was lost, so its stderr.log never reached an artifact.

A new step prints the last 32 KiB of each package's stryker.log as soon
as the shard fails, before anything else runs. The re-run is bounded by
timeout 600 (kill after 30 s more) and --concurrency 1, which gives one
checker and one test runner, and prints the last 32 KiB of its stderr,
stdout and stryker.log into the job log before the upload step.
At Stryker's default concurrency on the 4 vCPU hosted runner, a shard
starts 2 checkers, each driving a native tsgo process, and 2 test
runners. Run 37975117762 on #82 (main's code) and every gate run on #81
ended with exit 3 about two minutes in, or lost the 16 GB VM. The
bounded diagnostics re-run of the same project at concurrency 1 passed
its dry run and tested 276 mutants in five minutes without a failure.

concurrency: 2 splits into one checker and one test runner. The 2 GiB
test-runner heap cap stays as a second guard.
…r.log

Run 37978503748 printed "(no such file)" for packages/*/stryker.log.
This Stryker build accepts fileLogLevel but nothing writes the file:
the option appears only in the schema, the CLI table and the
fingerprint key list. The setting and every stryker.log path go.

The print step now tails the shard wrapper's own progress stream
(reports/mutation-stream.jsonl, the default for a run without
--progressStreamFile) and each project child's stream under
reports/shards. Every run writes its framed events to that file, one
synced line at a time. The diagnostics artifact keeps the wrapper's
stream, and the re-run prints its stderr and stdout.
`stryker run --shard` spawns one child per planned project with stdout
ignored, keeps 4096 characters of its stderr and prints 1024 of them
(stryker-js dist/main.mjs 113412-113483, 114204-114206). Every exit-3
run on #81 and #82 lost the error that way.

The shard step now runs the same children itself:
- stryker-plan-gate.ts gains a `shard` mode. It decodes the plan with the
  gate's ShardPlan schema and finds the shard by its index/count label,
  as the wrapper does (112975-112980, 113345-113362), refusing an unknown
  label. It writes "<index>\t<project>" lines in plan order.
- .github/scripts/run-shard.sh runs each project in its own directory
  with the wrapper's exact args (113415-113429) and seeding
  (113431-113437), writing stdout.log and stderr.log beside the stream
  under reports/shards/<index>/<project>/. Exit 1 logs the wrapper's
  below-threshold line and carries on (113478); any other exit, or a
  missing progress stream, prints both log tails and stops the shard,
  as the wrapper's sequential forEach does (113474-113483).

The diagnostics re-run and its artifact go: the shard artifact now
carries both logs. The verdict job, merge and gate are unchanged.
The guards capability no longer waits on a pull-request mutation gate.
Mutation runs on main only, so this branch goes back onto main without
#82's gate changes. release-gate.yml, the shard scripts,
stryker-plan-gate.ts with its tests, and stryker.shared.ts return to
main's content. They reach main through #82 itself.
kiro-systemf Bot pushed a commit that referenced this pull request Oct 9, 2026
…hards (#82)

* ci(repo): cap each mutation test runner isolate at a 2 GiB heap

Two of three release-gate runs on #81 lost the whole ubuntu-latest VM
(16 GB) mid-shard. Node's default old-space ceiling is about 4 GiB per
isolate, and each test-runner child holds a main isolate plus a claimed
and a spare vitest worker thread, so two runners running a mutant that
allocates without bound can outgrow the VM before the 45 s mutant
timeout settles them.

testRunnerNodeArgs passes --max-old-space-size=2048 to every test-runner
child. The flag is process-wide in V8, so each worker thread gets the
same ceiling. A worker that reaches it is ended by Node; Stryker retries
the mutant twice and then records it as a RuntimeError, so the run
carries on. The full xstate suite passes on one thread with a 512 MiB
ceiling, so 2 GiB leaves room for the coverage instrumentation.

* ci(repo): print a failed shard's logs before its diagnostics re-run

On #81 the shard ends with exit 3 about 26 s into mutation testing, and
the job log keeps only the first 1024 characters of the child's stderr,
which are startup lines. The diagnostics re-run then ran for 16 minutes
until the VM was lost, so its stderr.log never reached an artifact.

A new step prints the last 32 KiB of each package's stryker.log as soon
as the shard fails, before anything else runs. The re-run is bounded by
timeout 600 (kill after 30 s more) and --concurrency 1, which gives one
checker and one test runner, and prints the last 32 KiB of its stderr,
stdout and stryker.log into the job log before the upload step.

* ci(repo): run mutation shards with one checker and one test runner

At Stryker's default concurrency on the 4 vCPU hosted runner, a shard
starts 2 checkers, each driving a native tsgo process, and 2 test
runners. Run 37975117762 on #82 (main's code) and every gate run on #81
ended with exit 3 about two minutes in, or lost the 16 GB VM. The
bounded diagnostics re-run of the same project at concurrency 1 passed
its dry run and tested 276 mutants in five minutes without a failure.

concurrency: 2 splits into one checker and one test runner. The 2 GiB
test-runner heap cap stays as a second guard.

* ci(repo): print the failed shard's progress streams instead of stryker.log

Run 37978503748 printed "(no such file)" for packages/*/stryker.log.
This Stryker build accepts fileLogLevel but nothing writes the file:
the option appears only in the schema, the CLI table and the
fingerprint key list. The setting and every stryker.log path go.

The print step now tails the shard wrapper's own progress stream
(reports/mutation-stream.jsonl, the default for a run without
--progressStreamFile) and each project child's stream under
reports/shards. Every run writes its framed events to that file, one
synced line at a time. The diagnostics artifact keeps the wrapper's
stream, and the re-run prints its stderr and stdout.

* ci(repo): run each shard project's stryker child with its output kept

`stryker run --shard` spawns one child per planned project with stdout
ignored, keeps 4096 characters of its stderr and prints 1024 of them
(stryker-js dist/main.mjs 113412-113483, 114204-114206). Every exit-3
run on #81 and #82 lost the error that way.

The shard step now runs the same children itself:
- stryker-plan-gate.ts gains a `shard` mode. It decodes the plan with the
  gate's ShardPlan schema and finds the shard by its index/count label,
  as the wrapper does (112975-112980, 113345-113362), refusing an unknown
  label. It writes "<index>\t<project>" lines in plan order.
- .github/scripts/run-shard.sh runs each project in its own directory
  with the wrapper's exact args (113415-113429) and seeding
  (113431-113437), writing stdout.log and stderr.log beside the stream
  under reports/shards/<index>/<project>/. Exit 1 logs the wrapper's
  below-threshold line and carries on (113478); any other exit, or a
  missing progress stream, prints both log tails and stops the shard,
  as the wrapper's sequential forEach does (113474-113483).

The diagnostics re-run and its artifact go: the shard artifact now
carries both logs. The verdict job, merge and gate are unchanged.

* ci(repo): run the release gate on main only

Mutation testing runs on main, never on a pull request. The release
gate loses its pull_request trigger and everything that existed only
for it:
- the plan job's changed-files step and fetch-depth 2;
- the scope output, MUTATION_SCOPE and stryker.shared.ts's scopedMutate,
  so every enrolled file is mutated;
- stryker-plan-gate.ts's change scoping (GATE_FILES, scopeOf,
  resolveScope, readChanged, the out-of-scope report and the
  MalformedStrykerConfig refusal), its --changed and --scope-out flags,
  and their tests;
- the pull-request-only cache skip and cancel-in-progress.

The verdict job, the shard layout, run-shard.sh, concurrency 2, the
2 GiB heap cap and the log tails on failure stay.

The release-gate sandbox proof now finds the mutation step by
run-shard.sh instead of `stryker run`. It runs the shard selection and
shard steps with a stub stryker, and checks that the progress stream
lands where the shard artifact uploads it. README and CONTRIBUTING
describe the gate as main-only again.
@systemfsoftware-maker
systemfsoftware-maker removed this pull request from stack #83 October 9, 2026 21:14
@systemfsoftware-maker
systemfsoftware-maker changed the base branch from xs/release-gate-runner-heap to main October 9, 2026 21:14
@systemfsoftware-maker
systemfsoftware-maker added this pull request to stack #85 October 10, 2026 00:49
systemfsoftware-maker added a commit that referenced this pull request Oct 10, 2026
Merging #81 into capability 7 lengthened the XS1 row, and the merge
resolution left the table padded for the shorter row.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant