Skip to content

[Feat] Add bounded recovery for stalled browser and desktop tasks - #31

Open
prettygirlisnotme wants to merge 2 commits into
ThinkFlowLab:mainfrom
prettygirlisnotme:feat/bounded-recovery
Open

prettygirlisnotme wants to merge 2 commits into
ThinkFlowLab:mainfrom
prettygirlisnotme:feat/bounded-recovery

Conversation

@prettygirlisnotme

@prettygirlisnotme prettygirlisnotme commented Sep 27, 2026 •

Copy link
Copy Markdown

Why

Browser and desktop policies can repeatedly choose ineffective actions without recovering or explaining how to proceed. This adds bounded replanning for those stalls, following the small repeatable recovery on/off subset agreed in #16.

How

Start with s1a/recovery.py for the cumulative attempt/time budget, s1a/browser/decision_model.py for browser triggers and refreshed observations, and s1a/tool/rethink.py for desktop integration.

A stall spends an attempt, refreshes the observation read-only, and asks the existing chat model for a short plan. The next ordinary decision consumes that plan through the existing tools and permission checks. Progress does not refund attempts or active refresh/planning time. Failed, cancelled or exhausted recovery stops with a reason and a next action. An empty planner response also stops as a planner failure, preserving its charged attempt/time and refreshed observation instead of logging an empty plan as successful.

Failed observations cannot become empty-page DONE results. A partial answer after post-replan BLOCKED remains context and cannot make the task successful. Failed/cancelled calls remain visible in accounting; unknown usage or prices do not become zero cost.

What

Verification

These are local and opt-in system checks. Upstream GitHub Actions currently requires maintainer approval; its checks have not run yet.

  • Formatting, lint and types: .venv/bin/ruff format --check ., .venv/bin/ruff check ., .venv/bin/ty check.
  • .venv/bin/pytest -q: 640 passed, 41 skipped, 106 subtests.
  • S1A=.venv/bin/s1a PY=.venv/bin/python scripts/smoke.sh.
  • uv build --offline: source archive and wheel built successfully.
  • CHANGELOG.md and feature documentation updated.

Opt-in system tests: S1A_BROWSER_TESTS=1 .venv/bin/pytest -q tests/system/test_recovery_browser.py passed on the cluster (74.34 s); S1A_DESKTOP_TESTS=1 .venv/bin/pytest -q tests/system/test_recovery_desktop.py passed against a dedicated Windows fixture. Controlled browser/Windows subsets completed 12/6 trials without execution errors. Their scripted decisions and plans validate mechanics, not neural-model performance.

With real Laya 0.3.5 and DeepSeek v4.1 Flash, the original 12-trial matrix verified 0/6 in both arms and never triggered recovery. A separately reported exploratory server-validation variant verified 0/6 off and 2/6 on, with six planner calls. Paired mean overhead was 18.194 s, two decision calls, 0.83 chat calls and one unchanged-page action per task. Locked-form trials still failed despite correct plans. These small synthetic results do not establish a general success-rate improvement.

The post-replan BLOCKED correction and blank-plan guard were regression-tested after the frozen experiment; historical outcomes remain unchanged. The added browser/desktop blank-plan regressions cover empty, whitespace-only and null content, preserved budgets/observations, and a clear stop instead of an extra planning attempt. See the report and eval guide for reproduction settings.

Closes #16.

@prettygirlisnotme
prettygirlisnotme marked this pull request as ready for review September 28, 2026 18:59

@prettygirlisnotme prettygirlisnotme left a comment

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Author self-review of 72bf502 (current head):

  • Budget and execution: attempts are charged before the read-only refresh; active refresh/planning time accumulates across the task and progress does not refund it. Plans guide the next ordinary decision rather than executing tools. Browser recovery remains opt-in, the existing tool filtering/permission path remains in force, and the unbounded game branch is preserved.
  • Failure semantics: failed observations, empty planner responses, cancellation and exhausted recovery budgets retain their recorded reason. Post-replan BLOCKED remains a failure even with a partial answer. The targeted regressions cover empty/whitespace/null plans, retained observations and attempts, and no renewed browser decision after stopping.
  • Existing validation at this head: 640 tests passed, 41 skipped, 106 subtests passed; formatting, lint, type checks, CLI smoke and offline packaging passed. These are the recorded pre-publication results, not a claim that upstream CI has passed.
  • Evidence limits: the report and trajectories separate scripted mechanism checks from the 24 real-model fixture trials. Original and exploratory matrices are reported separately. The small sample does not establish general task-success gains, and the final reporting/blank-plan fixes were regression-tested after the measured revision.

The remaining integration check is upstream CI: the latest run is action_required with no jobs executed. Could a maintainer approve the workflow run? The upstream OS/Python matrix is still unverified.

This is an author self-review to make the change easier to assess, not an independent approval.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feat] Recover from stalled browser and desktop tasks with bounded replanning

1 participant