Repository navigation
test(cpython): budget the crash-containment case for its real cost - #2666
Merged
Merged
Conversation
deep_c_recursion_is_contained asserts that a guest C-stack overflow is contained: the trap ends only that call, exit 1, shell survives. Reaching the trap is inherently slow. Exhausting the guest's 4 MiB wasm stack through CPython's C recursion in repr costs ~16 s of interpreted guest execution, and several times that under AddressSanitizer, against a 30 s default budget. The shell's own wall clock won the race, so the call ended 124 and the trap was never reached: nightly run 250 (ASAN) went red, and the same test fails inside a contended parallel suite run, while passing in 16.28 s in isolation. The depth is not the waste it looks like. Measured on this path, nesting at or below 20,000 does not trap at all -- repr succeeds -- and every depth from 22,000 to 200,000 costs the same 12.5-16 s, because the time goes into exhausting the stack rather than building the list. Shrinking it would buy ~3 s and move the case toward the threshold where it passes while testing nothing, so the depth stays and the budget is sized for the work instead. Scope is this one case: the sibling containment tests cost 0.01-1.61 s. A trap that stops working still fails here, through a host crash or the stdout/stderr assertions, so the wider budget costs no coverage. Claude-Session: https://claude.ai/code/session_01BegucbcUgCEsPZiAzt6GRx
Deploying with
|
| Status | Name | Latest Commit | Preview URL | Updated (UTC) |
|---|---|---|---|---|
| ✅ Deployment successful! View logs |
bashkit | dcd2816 | Commit Preview URL Branch Preview URL |
Oct 10 2026, 10:01 AM |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
deep_c_recursion_is_containedruns with a budget sized for the work itactually does (300 s, via a
CONTAINMENT_BUDGEThelper that sets both theshell's
ExecutionLimits::timeoutandCPythonLimits::max_duration) instead ofthe 30 s default. The nesting depth is unchanged; the comments now record why.
No production code changes — this is the test's budget only.
Why
The test asserts the TM-PY-CPY-007 property: a guest C-stack overflow is
contained — the trap ends only that call (exit 1, Display-only message) and
the shell survives. Reaching that trap is inherently slow. Exhausting the
guest's 4 MiB wasm stack (
MAX_WASM_STACK) through CPython's C recursion inreprcosts ~16 s of interpreted guest execution, and several times that underAddressSanitizer.
Against a 30 s budget the shell's own wall clock won the race, so the call ended
124 and the trap was never reached. That is nightly run 250 (ASAN) going
red, and it is not ASAN-specific: the same test fails inside a contended
parallel suite run while passing in 16.28 s in isolation.
ci.ymlruns thistest (
-- cpython), so it was a latent timing dependency in required CI thatmain passes by margin, not by design.
Before / After
Why the depth is not the thing to cut
My first instinct was that 200,000 was gratuitous. Measured, it is not — the
cost is flat above the trap threshold, because the time goes into exhausting the
stack rather than building the list:
reprsucceeds, exit 0 — no trapreprsucceeds, exit 0 — no trapreprsucceeds, exit 0 — no trapfatal error: stack overflow in the interpreterSo the cliff is 20k–22k: below it
reprsimply completes and the property goesuntested. Shrinking the depth would buy ~3 s and move the case toward passing
while verifying nothing. The depth stays, and the comment says so. The cliff is
self-guarding — a too-small depth yields
after 0and fails the assertion.Scope
deep_c_recursionis the only outlier, so this stays a one-test change:deep_c_recursiondescriptor_tableparser_bombshell_continuesabortRisk
take the host down or return the wrong status/stderr, and the
after 1andpython3: fatal error:assertions still fail. The budget can only stop theshell's wall clock from pre-empting a trap that does work.
Checklist
https://claude.ai/code/session_01BegucbcUgCEsPZiAzt6GRx
Generated by Claude Code