[Feat] Add bounded recovery for stalled browser and desktop tasks - #31
Open
prettygirlisnotme wants to merge 2 commits into
Open
prettygirlisnotme wants to merge 2 commits into
prettygirlisnotme wants to merge 2 commits into
Conversation
prettygirlisnotme
marked this pull request as ready for review
September 28, 2026 18:59
prettygirlisnotme
left a comment
Author
There was a problem hiding this comment.
Author self-review of 72bf502 (current head):
- Budget and execution: attempts are charged before the read-only refresh; active refresh/planning time accumulates across the task and progress does not refund it. Plans guide the next ordinary decision rather than executing tools. Browser recovery remains opt-in, the existing tool filtering/permission path remains in force, and the unbounded game branch is preserved.
- Failure semantics: failed observations, empty planner responses, cancellation and exhausted recovery budgets retain their recorded reason. Post-replan BLOCKED remains a failure even with a partial answer. The targeted regressions cover empty/whitespace/null plans, retained observations and attempts, and no renewed browser decision after stopping.
- Existing validation at this head: 640 tests passed, 41 skipped, 106 subtests passed; formatting, lint, type checks, CLI smoke and offline packaging passed. These are the recorded pre-publication results, not a claim that upstream CI has passed.
- Evidence limits: the report and trajectories separate scripted mechanism checks from the 24 real-model fixture trials. Original and exploratory matrices are reported separately. The small sample does not establish general task-success gains, and the final reporting/blank-plan fixes were regression-tested after the measured revision.
The remaining integration check is upstream CI: the latest run is action_required with no jobs executed. Could a maintainer approve the workflow run? The upstream OS/Python matrix is still unverified.
This is an author self-review to make the change easier to assess, not an independent approval.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Browser and desktop policies can repeatedly choose ineffective actions without recovering or explaining how to proceed. This adds bounded replanning for those stalls, following the small repeatable recovery on/off subset agreed in #16.
How
Start with
s1a/recovery.pyfor the cumulative attempt/time budget,s1a/browser/decision_model.pyfor browser triggers and refreshed observations, ands1a/tool/rethink.pyfor desktop integration.A stall spends an attempt, refreshes the observation read-only, and asks the existing chat model for a short plan. The next ordinary decision consumes that plan through the existing tools and permission checks. Progress does not refund attempts or active refresh/planning time. Failed, cancelled or exhausted recovery stops with a reason and a next action. An empty planner response also stops as a planner failure, preserving its charged attempt/time and refreshed observation instead of logging an empty plan as successful.
Failed observations cannot become empty-page DONE results. A partial answer after post-replan BLOCKED remains context and cannot make the task successful. Failed/cancelled calls remain visible in accounting; unknown usage or prices do not become zero cost.
What
--rethink on;--rethink-attemptsand--rethink-timeoutbound recovery. Desktop recovery reusesRethinkRailwhile preserving existing game behavior.Verification
These are local and opt-in system checks. Upstream GitHub Actions currently requires maintainer approval; its checks have not run yet.
.venv/bin/ruff format --check .,.venv/bin/ruff check .,.venv/bin/ty check..venv/bin/pytest -q: 640 passed, 41 skipped, 106 subtests.S1A=.venv/bin/s1a PY=.venv/bin/python scripts/smoke.sh.uv build --offline: source archive and wheel built successfully.CHANGELOG.mdand feature documentation updated.Opt-in system tests:
S1A_BROWSER_TESTS=1 .venv/bin/pytest -q tests/system/test_recovery_browser.pypassed on the cluster (74.34 s);S1A_DESKTOP_TESTS=1 .venv/bin/pytest -q tests/system/test_recovery_desktop.pypassed against a dedicated Windows fixture. Controlled browser/Windows subsets completed 12/6 trials without execution errors. Their scripted decisions and plans validate mechanics, not neural-model performance.With real Laya 0.3.5 and DeepSeek v4.1 Flash, the original 12-trial matrix verified 0/6 in both arms and never triggered recovery. A separately reported exploratory server-validation variant verified 0/6 off and 2/6 on, with six planner calls. Paired mean overhead was 18.194 s, two decision calls, 0.83 chat calls and one unchanged-page action per task. Locked-form trials still failed despite correct plans. These small synthetic results do not establish a general success-rate improvement.
The post-replan BLOCKED correction and blank-plan guard were regression-tested after the frozen experiment; historical outcomes remain unchanged. The added browser/desktop blank-plan regressions cover empty, whitespace-only and null content, preserved budgets/observations, and a clear stop instead of an extra planning attempt. See the report and eval guide for reproduction settings.
Closes #16.