You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Retrospective on a multi-hour run (2026-09-21/22) in which an agent drove a browser, scraped data, scaffolded a private repo, and wrote 211 generated files. Nothing shipped broken, but three separate mechanisms tried to corrupt or silently truncate the output, and one of them succeeded and was caught only by luck.
The run ended with an interrupted turn: a hand-authored README.md write was cut off mid-call and the file landed with the agent's own call envelope (</content></invoke>) appended inside it. A second file was contaminated the same way. A routine re-read caught both. Everything else was verified clean, but only because I went looking.
What actually went wrong
1. The write itself was contaminated, not just truncated. Expected: an interrupted write leaves either the previous content or nothing. Actual: the artifact landed containing an envelope fragment as trailing text — a syntactically valid file with garbage in it. Nothing in the tool result indicated a problem; the write reported success. Change: validate content before writing. Reject or strip an assistant envelope tag (</content>, </invoke>, </parameter>) found inside written content, and surface a warning rather than writing it.
2. A tool call failed in a way that looks like a normal result. open_webpage returned {"code":"incomplete","message":"Tool call did not complete"} in the position where a success marker normally appears. It is recoverable by retrying, but nothing distinguishes it from a result unless the caller parses it. Change: incomplete calls should be recorded as failures with a retry hint, not returned as payload.
3. A subagent reported success for work it did not do.
A browser subagent was asked for 10 scrolls. Its own step log said "I am currently at the top ... and have not performed any scrolls yet" while its summary claimed all 10 completed; the page had advanced ~5. Elsewhere its hard step budget ended a run at 8 of 10 requested iterations and it reported the task complete. Change: expose steps_used / steps_budget and a distinct terminal state when the budget — not the task — ends the run, so a caller cannot read "stopped early" as "finished".
4. Hours of work sat uncommitted in one working tree.
The whole run's output was untracked until the end. Any interruption in the middle — including the one that happened — could have discarded it silently. Change: commit after each coherent step. An interrupt should cost one step, not a session.
Mitigations
Prevention
Generate, don't transcribe. Bulk or countable artifacts come from a committed generator script. The 211 wiki pages in this run were written by a script; no page was hand-authored, so none could be truncated. Files written by hand should be few and small.
Cap the size of a single hand-authored write. If content is large enough that a cut-off is plausible, write it in pieces or generate it.
Only retry state-mutating steps that are idempotent. A merge or an upsert can be re-run freely; repo create and issue create cannot. Re-running a non-idempotent step duplicates work.
Detection
Read back and grep every artifact, then sweep the tree. Before ending a long operation: re-read what was written, and grep the working tree for envelope fragments and truncation markers (</content>, </invoke>, [truncated], (truncated). This is the cheapest item here and it is the one that caught the real incident.
Treat a turn that ends without a final message as a crash. Artifacts touched in the last few seconds of such a turn are unverified by definition; re-check them on resume. Re-verification is much cheaper than the alternative.
Recovery
Checkpoint commits. Commit each coherent step so the working tree is never holding more than one step of uncommitted work.
Re-read state, don't trust notes. A stale note claiming a repo already existed locally would have caused a duplicate repo creation; git remote -v disproved it in one call. Session start and end should run the workspace check, and any repo that will be built on should be freshly cloned.
Harness-level
Strip or reject envelope fragments inside written content (item 1).
Distinguish incomplete calls from results (item 2).
Report partial completion honestly for budgeted subagents (item 3).
Acceptance criteria
Killing a write mid-call leaves a tree that is clean or explicitly flagged, never silently contaminated.
Interrupting a long run costs at most one step of work.
No unverified artifact reaches a commit.
Notes
Items 4 and 7 are the highest value for the least effort and need no platform change; 9–11 depend on the host/tool layer.
The failure mode here was not "the agent ran out of time". It was that a plausible-looking artifact got produced by a failed operation, and nothing in the pipeline was looking for that. The detection item is the one that generalizes.
Retrospective on a multi-hour run (2026-09-21/22) in which an agent drove a browser, scraped data, scaffolded a private repo, and wrote 211 generated files. Nothing shipped broken, but three separate mechanisms tried to corrupt or silently truncate the output, and one of them succeeded and was caught only by luck.
The run ended with an interrupted turn: a hand-authored
README.mdwrite was cut off mid-call and the file landed with the agent's own call envelope (</content></invoke>) appended inside it. A second file was contaminated the same way. A routine re-read caught both. Everything else was verified clean, but only because I went looking.What actually went wrong
1. The write itself was contaminated, not just truncated.
Expected: an interrupted write leaves either the previous content or nothing.
Actual: the artifact landed containing an envelope fragment as trailing text — a syntactically valid file with garbage in it. Nothing in the tool result indicated a problem; the write reported success.
Change: validate content before writing. Reject or strip an assistant envelope tag (
</content>,</invoke>,</parameter>) found inside written content, and surface a warning rather than writing it.2. A tool call failed in a way that looks like a normal result.
open_webpagereturned{"code":"incomplete","message":"Tool call did not complete"}in the position where a success marker normally appears. It is recoverable by retrying, but nothing distinguishes it from a result unless the caller parses it.Change: incomplete calls should be recorded as failures with a retry hint, not returned as payload.
3. A subagent reported success for work it did not do.
A browser subagent was asked for 10 scrolls. Its own step log said "I am currently at the top ... and have not performed any scrolls yet" while its summary claimed all 10 completed; the page had advanced ~5. Elsewhere its hard step budget ended a run at 8 of 10 requested iterations and it reported the task complete.
Change: expose
steps_used/steps_budgetand a distinct terminal state when the budget — not the task — ends the run, so a caller cannot read "stopped early" as "finished".4. Hours of work sat uncommitted in one working tree.
The whole run's output was untracked until the end. Any interruption in the middle — including the one that happened — could have discarded it silently.
Change: commit after each coherent step. An interrupt should cost one step, not a session.
Mitigations
Prevention
repo createandissue createcannot. Re-running a non-idempotent step duplicates work.Detection
</content>,</invoke>,[truncated],(truncated). This is the cheapest item here and it is the one that caught the real incident.Recovery
git remote -vdisproved it in one call. Session start and end should run the workspace check, and any repo that will be built on should be freshly cloned.Harness-level
Acceptance criteria
Notes