Skip to content

test(cpython): budget the crash-containment case for its real cost - #2666

Merged
chaliy merged 1 commit into
mainfrom
eve/clever-mccarthy-fasrpn
Oct 10, 2026
Merged

chaliy merged 1 commit into
mainfrom
eve/clever-mccarthy-fasrpn

Conversation

@chaliy

@chaliy chaliy commented Oct 10, 2026

Copy link
Copy Markdown
Contributor

What changed

deep_c_recursion_is_contained runs with a budget sized for the work it
actually does (300 s, via a CONTAINMENT_BUDGET helper that sets both the
shell's ExecutionLimits::timeout and CPythonLimits::max_duration) instead of
the 30 s default. The nesting depth is unchanged; the comments now record why.

No production code changes — this is the test's budget only.

Why

The test asserts the TM-PY-CPY-007 property: a guest C-stack overflow is
contained — the trap ends only that call (exit 1, Display-only message) and
the shell survives. Reaching that trap is inherently slow. Exhausting the
guest's 4 MiB wasm stack (MAX_WASM_STACK) through CPython's C recursion in
repr costs ~16 s of interpreted guest execution, and several times that under
AddressSanitizer.

Against a 30 s budget the shell's own wall clock won the race, so the call ended
124 and the trap was never reached. That is nightly run 250 (ASAN) going
red, and it is not ASAN-specific: the same test fails inside a contended
parallel suite run while passing in 16.28 s in isolation. ci.yml runs this
test (-- cpython), so it was a latent timing dependency in required CI that
main passes by margin, not by design.

Before / After

Before (default 30 s budget)
  isolated, pristine main ........ ok, 16.28 s          <- passes by ~14 s of slack
  inside full parallel suite ..... FAILED, "after 124"  <- trap never reached
  nightly 250 (ASAN) ............. FAILED, "after 124"

After (300 s containment budget)
  inside full parallel suite ..... ok

Why the depth is not the thing to cut

My first instinct was that 200,000 was gratuitous. Measured, it is not — the
cost is flat above the trap threshold, because the time goes into exhausting the
stack rather than building the list:

depth time outcome
1,000 0.11 s repr succeeds, exit 0 — no trap
10,000 3.69 s repr succeeds, exit 0 — no trap
20,000 12.93 s repr succeeds, exit 0 — no trap
22,000 12.65 s traps: fatal error: stack overflow in the interpreter
25,000 13.03 s traps
50,000 13.17 s traps
200,000 15.93 s traps

So the cliff is 20k–22k: below it repr simply completes and the property goes
untested. Shrinking the depth would buy ~3 s and move the case toward passing
while verifying nothing. The depth stays, and the comment says so. The cliff is
self-guarding — a too-small depth yields after 0 and fails the assertion.

Scope

deep_c_recursion is the only outlier, so this stays a one-test change:

containment case cost
deep_c_recursion 15.93 s
descriptor_table 1.61 s
parser_bomb 0.07 s
shell_continues 0.04 s
abort 0.01 s

Risk

  • Low. Test-only; no production code or limits change.
  • The wider budget costs no coverage: if containment regressed, the guest would
    take the host down or return the wrong status/stderr, and the after 1 and
    python3: fatal error: assertions still fail. The budget can only stop the
    shell's wall clock from pre-empting a trap that does work.
  • A genuine hang in this path now takes up to 300 s to surface instead of 30 s.

Checklist

  • Tests added or updated
  • Backward compatibility considered

https://claude.ai/code/session_01BegucbcUgCEsPZiAzt6GRx


Generated by Claude Code

deep_c_recursion_is_contained asserts that a guest C-stack overflow is
contained: the trap ends only that call, exit 1, shell survives. Reaching the
trap is inherently slow. Exhausting the guest's 4 MiB wasm stack through
CPython's C recursion in repr costs ~16 s of interpreted guest execution, and
several times that under AddressSanitizer, against a 30 s default budget. The
shell's own wall clock won the race, so the call ended 124 and the trap was
never reached: nightly run 250 (ASAN) went red, and the same test fails inside
a contended parallel suite run, while passing in 16.28 s in isolation.

The depth is not the waste it looks like. Measured on this path, nesting at or
below 20,000 does not trap at all -- repr succeeds -- and every depth from
22,000 to 200,000 costs the same 12.5-16 s, because the time goes into
exhausting the stack rather than building the list. Shrinking it would buy
~3 s and move the case toward the threshold where it passes while testing
nothing, so the depth stays and the budget is sized for the work instead.

Scope is this one case: the sibling containment tests cost 0.01-1.61 s.

A trap that stops working still fails here, through a host crash or the
stdout/stderr assertions, so the wider budget costs no coverage.

Claude-Session: https://claude.ai/code/session_01BegucbcUgCEsPZiAzt6GRx
@cloudflare-workers-and-pages

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Preview URL Updated (UTC)
✅ Deployment successful!
View logs
bashkit dcd2816 Commit Preview URL

Branch Preview URL
Oct 10 2026, 10:01 AM

@chaliy
chaliy merged commit 03ea9b3 into main Oct 10, 2026
33 checks passed
@chaliy
chaliy deleted the eve/clever-mccarthy-fasrpn branch October 10, 2026 10:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant