Skip to content

test(parser): assert full bounds contract in multiline-ARG0 regression - #2680

Merged
chaliy merged 7 commits into
mainfrom
codex/syntax-diagnostic-budget
Oct 11, 2026
Merged

chaliy merged 7 commits into
mainfrom
codex/syntax-diagnostic-budget

Conversation

@chaliy

@chaliy chaliy commented Oct 11, 2026

Copy link
Copy Markdown
Contributor

What changed\nTest-only: the TM-DOS-136 scaled multiline-token x long-ARG0 regression now asserts the full bounded-harness contract at both scales (200/2000 token lines x 8 KiB ARG0): no watchdog timeout, raw child exit 2, no 64 KiB output-cap overshoot, finished reader threads, and effective stack=4096/cpu=10 caps read back from the launcher line. Shape (Bash-compatible syntax error) and flat <=1 KiB stderr assertions unchanged. Reuses the reviewed shared launcher; no argv-specific readers.\n\n## Why\nFollowup to the review on #2675 (comment 6105424708) and the #2677 proof note: a sampled diagnostic and raw exit 2 alone do not prove the full bounds contract. This adds the same explicit assertions the adjacent hostile-indirect tests carry.\n\n## Before / After\nBefore: the regression proved exit 2 + shape + <=1024 bytes only.\nAfter: it additionally proves !oversized, !incomplete_readers, and the effective fail-closed CPU/stack caps, per scale.\n\n\ntest result: ok. 1 passed (syntax_diagnostic_multiline, both scales)\ntest result: ok. indirect suite green\ncargo fmt --check: clean; cargo clippy -p bashkit --all-targets: clean\n\n\n## Risk\n- Test-only change. No production code touched.\n\n## Checklist\n- [x] Tests added or updated\n- [x] Backward compatibility considered (internal test harness)

macOS bash 3.2 reports -c input as line 0, modern GNU bash as line 1.
Compare diagnostic shape with line number normalized.
…gression

A long -c ARG0 stamped on every report line could evict the structural
payload out of the 1 KiB diagnostic budget. Truncate the shell name to 64
chars first so 'syntax error ...' always survives; renumber the threat entry
to TM-DOS-136 (main already owns 135 via #2674); add a bounded-child
regression exercising the exact finding shape (multiline token x long ARG0)
at two scales.
The 64-char cut mangled long script paths (breaking the oracle's path
replacement and hiding which script failed). Widen to 256 chars with a
16-char tail: realistic paths print verbatim with attribution suffixes
(': eval') intact, while a kilobyte-scale ARG0 still cannot evict the
structural payload out of the 1 KiB total. The total stays the enforced
bound.
All five branch commits already reached main via squash merges (#2675, #2677); conflicting files take main's reviewed state.
The scaled multiline-token x long-ARG0 regression now asserts the whole harness contract at both scales (200/2000 lines x 8 KiB ARG0): no watchdog timeout, raw child exit 2, no output-cap overshoot, finished reader threads, and effective stack=4096/cpu=10 caps read back from the launcher. Shape and 1 KiB assertions unchanged.
@cloudflare-workers-and-pages

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Preview URL Updated (UTC)
✅ Deployment successful!
View logs
bashkit c51d511 Commit Preview URL

Branch Preview URL
Oct 11 2026, 07:35 AM

@chaliy
chaliy merged commit fff4240 into main Oct 11, 2026
33 checks passed
@chaliy
chaliy deleted the codex/syntax-diagnostic-budget branch October 11, 2026 08:03
@chaliy

chaliy commented Oct 11, 2026

Copy link
Copy Markdown
Contributor Author

Post-merge audit of fff4240: the added !oversized / !incomplete_readers / effective stack4096 and cpu10 assertions address6105850785 in both exact multiline-token x8KiB ARG0 cases. Actual merged CLI build has producer exit0; remaining test/CI proof is running. This is not finding closure yet.

A source-level gap remains in the reused resource_exhaustion launcher in crates/bashkit/tests/integration/threat_model_tests.rs. run_command_bounded ends with unconditional child.wait() even when deadline expires and status is None, so its claimed deadline-bounded reaping is not implemented. spawn_capped_reader treats Err(_) as normal EOF/result, allowing incomplete_readers=false after a failed capture, and extends before checking samplecap (retained bytes can exceed advertised8192 by a chunk). The two-second drain grace is bounded but outside the stated shared wall deadline; distinguish any documented cleanup grace from the advertised execution deadline rather than claiming one strict total bound. Do not claim all lifecycle bounds proven solely from the new consumer assertions or happy-path resource rejection tests.

Resolve via the current owned syntax branch/worktree and any existing open follow-up first, preserving already merged history; no duplicate PRs, forcepush or reset. Use one clearly specified bounded execution/cleanup/reap/reader lifecycle, preserve pre-kill timeout state, fail closed on real reader errors (retry Interrupted), cap retained samples before extending and signal actual excess, and demonstrate timeout/oversize/read-error/retained-pipe/reap paths with safe bounded self-tests. Keep honest Darwin limits and raw producer versus assertion exits. The syntax product proof must retain actual full required Linux conclusions and exact merged source/regression evidence; --skip misconfig_huge_ast_depth_still_safe is partial local evidence only, not a full-gate waiver. Also correct the new assertion message that says64KiB when BOUNDED_SAMPLE_CAP is8192.

This test-harness audit does not reopen or redundantly resume the independently fixed indirect/arithmetic/function/array/prompt product findings. Their exact product regression receipts remain valid. No cloud closure for syntax until fresh merged proof and full reviews are honestly complete.

chaliy added a commit that referenced this pull request Oct 11, 2026
## What changed

Test-only, one file: threat_model_tests.rs (+211/-33). Reworks the
shared run_command_bounded lifecycle per the review on #2680:

- Retention truncated BEFORE extending: the sample buffer never passes
the 8192 cap (was: extend-then-truncate kept a chunk over).
- Interrupted reads retried; real reader errors end capture with partial
bytes kept and a new read_error flag — never relabeled as clean EOF.
- Exact-cap-then-EOF is not oversize; only real excess sets the flag.
- Reaping bounded inside a finite 5s grace; no unconditional wait() past
the cleanup deadline (targeted allow documents why the lint suggestion
would reintroduce the hang).
- timed_out retained as recorded, never recomputed; drain 2s + reap 5s
graces documented against the execution deadline.
- TM-DOS-136 regression gains !read_error and the corrected 8192-byte
cap text.
- Bounded selftests: exact-cap, cap+1, scripted Interrupted/error pipes,
timeout kill+reap (~2s), overflow flag.

Does not touch #2681 distinct run_bounded_probe helper.

## Before / After

Before: reader errors read as clean EOF; retention could exceed the cap
by a chunk; kill-path wait() unbounded; no lifecycle selftests.
After: threat_model 219 passed (only pre-existing macOS nesting abort
skipped, identical on unmodified code); error 14, partial_parse 11
green; fmt plus clippy -D warnings clean.

## Risk

- Test-only change. No production code touched.

## Checklist

- [x] Tests added or updated
- [x] Backward compatibility considered (internal test harness)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant